Vector shuffling method, processor and electronic device

By adding register identifiers and breaking parameters to the instructions, vector breaking operations are directly completed in one instruction, which solves the problems of multiple instructions and complex operations in the prior art, improves execution efficiency and reduces system overshots.

JP7675938B2Active Publication Date: 2025-05-13LOONGSON TECH CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024534616
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-12-10
Filing Date
2022-12-08
Publication Date
2025-05-13
Estimated Expiration
2042-12-08

AI Technical Summary

Technical Problem

In the prior art, implementing vector breaking operation requires multiple instructions, which is complex, resulting in a reduced execution efficiency of specific functions.

Method used

By adding register identifiers and breaking parameters to the instructions, the vector breaking operation is performed, and the vector breaking operation is directly completed in one instruction, reducing the required number of instructions and operation complexity.

Benefits of technology

The implementation of vector breaking operations is simplified, the execution efficiency of specific functions is improved, and the system is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007675938000001
    Figure 0007675938000001
  • Figure 0007675938000002
    Figure 0007675938000002
  • Figure 0007675938000003
    Figure 0007675938000003
Patent Text Reader

Abstract

The present application provides a method, processor and electronic device for shuffling a vector. The method includes the steps of receiving an instruction including a register identifier and a shuffle parameter, the register identifier including a source register identifier and a destination register identifier, the source register identifier is for indicating the source register, the source register is a register for storing a source element to be operated when performing the vector shuffle operation, the destination register identifier is for indicating the destination register, the destination register is a register for storing a target element obtained after performing the vector shuffle operation, and the shuffle parameter is for indicating a parameter to be followed when performing the vector shuffle operation on the source element, executing the instruction to perform the vector shuffle operation on the source element obtained from the source register according to the shuffle parameter, and obtaining a target element after the vector shuffle operation, and writing the target element to the destination register. The present application allows the vector shuffle operation of a specific function to be implemented with a single instruction, improving the execution efficiency of the specific function.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present application relates to the field of computer technology, and in particular to a method, processor and electronic device for shuffling a vector. [Background technology]

[0002] With the development of multimedia applications, more and more processor computing tasks come from the field of digital image processing, and image-based applications have become a non-negligible workload in servers, desktop computers, personal mobile devices, i.e. embedded devices. Updating the instruction set architecture to fit the realities of digital image processing software and adding instruction support for operations commonly used in applications to processors is the main direction of processor development, and is also a simple and effective way to improve the performance of processors for certain applications, so more and more processors are adding Single Instruction Multiple Data (SIMD) architectures that support the same operations on regular data sets.

[0003] Currently, shuffle instructions are widely used in SIMD processors, and different shuffle instructions can meet different needs. However, in the conventional technical solutions, when implementing vector shuffle operations for specific functions, multiple instructions are required to implement a series of operations, which complicates the operation method and reduces the execution efficiency of the specific functions. Summary of the Invention [Problem to be solved by the invention]

[0004] The present application provides a vector shuffling method, a processor and an electronic device to solve the problem in the prior art that multiple instructions are required to realize a series of operations, the operation method is complicated, and the execution efficiency of certain functions is reduced. [Means for solving the problem]

[0005] In order to solve the above problems, the present application discloses a method for shuffling vectors, the method comprising: receiving an instruction comprising register identifiers and shuffle parameters, wherein the register identifiers comprise source register identifiers and destination register identifiers, the source register identifiers for indicating source registers, the source registers being registers for storing source elements to be operated on when performing a vector shuffle operation, the destination register identifiers for indicating destination registers, the destination registers being registers for storing target elements resulting after performing the vector shuffle operation, and the shuffle parameters for indicating parameters according to which the vector shuffle operation is performed on the source elements; Executing the instructions to perform a vector shuffle operation on source elements obtained from the source registers according to the shuffle parameters, and performed getting the next target element; writing the target element to the destination register.

[0006] In order to solve the above problem, the present application discloses a processor, the processor comprising: a source register for storing a data element; and Destination Register A plurality of vector registers, including: a decode unit for decoding a vector shuffle instruction including register identifiers, including source register identifiers and destination register identifiers, and shuffle parameters; an execution unit responsive to the vector shuffle instruction to perform a vector shuffle operation on source elements obtained from the source register in accordance with the shuffle parameters, obtain target elements after the vector shuffle operation, and write the target elements to the destination register.

[0007] In order to solve the above problems, the present application discloses an electronic device including a memory and one or more programs, the one or more programs being stored in the memory and executed by one or more processors. Before The vector shuffling method is configured to perform the vector shuffling method described above. Effect of the Invention

[0008] The present application has the following advantages over the conventional techniques:

[0009] The vector shuffle method, processor and electronic device provided by the embodiments of the present application can add a register identifier and a shuffle parameter to an instruction, and use the shuffle parameter to perform a vector shuffle operation on elements obtained from a source register. This eliminates the need to perform shuffle operations with multiple instructions to achieve a specific function, and allows the vector shuffle operation of the specific function to be achieved with a single instruction, thereby improving the execution efficiency of the specific function. [Brief description of the drawings]

[0010] [Figure 1] 1 is a flowchart of steps of a method for shuffling one vector provided by Example 1 of the present application. [Diagram 2] 1 is a flowchart of steps of a method for shuffling one vector provided by Example 2 of the present application. [Diagram 3] 1 is a flowchart of steps of a method for shuffling one vector provided by Example 3 of the present application. [Figure 4]1 is a flowchart of steps of a method for shuffling one vector provided by Example 4 of the present application. [Diagram 5] 1 is a flowchart of the steps of a method for shuffling one vector provided by Example 5 of the present application. [Figure 6] 1 is a flow chart of steps of a method for shuffling one vector provided by Example 6 of the present application. [Figure 7] FIG. 2 is a structural block diagram of a processor provided by an embodiment of the present application. [Figure 8] FIG. 2 is a structural block diagram of an electronic device provided according to an embodiment of the present application; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] In order to make the above objects, features and advantages of the present application more clear and understandable, the present application will be described in more detail below with reference to the drawings and specific embodiments.

[0012] In this specification and claims, as well as in the drawings, the terms "first," "second," "third," and the like are used to distinguish between similar or like objects or entities and do not necessarily imply a particular order or priority, unless otherwise indicated. It should be understood that the terms used in this manner are interchangeable, for example, so that the embodiments of the present application may be performed in an order other than that shown in the illustrations or descriptions.

[0013] Although the following embodiments are described with reference to a processor, other embodiments are applicable to other types of integrated circuits and logic devices. The above techniques and teachings of the present application can be readily applied to other types of circuits or semiconductor devices that benefit from higher pipeline throughput and improved performance. The embodiments of the present application are applicable to any processor or machine that performs data manipulation. However, the present application is not limited to processors or machines that perform 256-bit, 128-bit, 64-bit, 32-bit, or 16-bit data manipulation, but is applicable to any processor and machine in which complex data needs to be manipulated.

[0014] In the following description, for the purpose of explanation, many specific details are shown to provide a thorough understanding of the present application. However, those skilled in the art should recognize that these specific details are not necessary to implement the present application. In other cases, some well-known electrical structures and circuits are not shown in detail to avoid unnecessarily confusing the present application. Also, in the following description, a number of examples are provided for the purpose of explanation, and various examples are shown in the drawings. However, these examples should not be construed as limiting, as they are intended only to provide some examples of the present application and do not provide an exhaustive list of all possible embodiments of the present application.

[0015] Although the following examples describe instruction processing and distribution in the context of an execution unit, other embodiments of the present application may be implemented in the form of software. In one embodiment, the methods of the present application are expressed as machine-executable instructions. The instructions can be used to cause a general-purpose or special-purpose processor programmed with these instructions to perform the steps of the present application. The steps of the present application can be performed by dedicated hardware components that contain hardwired logic to perform the steps, or by any combination of programmed computer components and custom hardware components. The software can be stored in memory within the system.

[0016] Example 1

[0017] Please refer to Figure 1. Figure 1 shows a flow chart of the steps of one vector shuffling method provided by an embodiment of the present application.

[0018] The vector shuffling method provided by the embodiment of the present application may be executed by a CPU (Central Processing Unit), and includes the following steps 101 to 103.

[0019] In step 101, a command is received that includes a register identifier and a shuffle parameter.

[0020] In the present embodiment, the instruction is an instruction for performing a vector shuffle operation, and the instruction is an instruction for execution by the CPU.

[0021] When performing a vector shuffle operation, an instruction may be received by the CPU to perform the vector shuffle operation, the instruction including a register identifier and a shuffle parameter.

[0022] The register identifier may include a source register identifier and a destination register identifier, the source register identifier is for indicating a source register, which is a register that stores a source element to be operated upon when performing the vector shuffle operation, and the source element to be operated upon when performing the shuffle operation may be all data stored in the source register or may be a part of the data stored in the source register, and the destination register identifier is for indicating a destination register, which is a register that stores a target element obtained after performing the vector shuffle operation.

[0023] In this example, the number of source registers may be one or two, i.e., the source elements are derived from one or two registers; specifically, the number of source registers can be determined according to business needs, and the embodiments of the present application are not limited thereto.

[0024] The shuffle parameters can be used to indicate parameters to follow when performing a vector shuffle operation on the source elements, and in this example, the shuffle parameters can include parameters such as index values ​​and opcodes, and optionally, the index values ​​are represented in the form of immediate values ​​and the opcodes are codes represented in binary form or the opcodes are identifiers that can be converted to binary codes.

[0025] After the command is received, step 102 is executed.

[0026] Step 102 executes the instruction to perform a vector shuffle operation on source elements obtained from the source register according to the shuffle parameters, and performs the vector shuffle operation. After the work Gets the target element of .

[0027] The target element is the element that results after performing a vector shuffle operation on the elements in the source register.

[0028] In an embodiment of the present application, after a CPU receives an instruction to perform a vector shuffle operation, the CPU can execute the instruction to perform the vector shuffle operation on source elements obtained from source registers according to shuffle parameters, and obtain target elements after the vector shuffle operation is performed.

[0029] Step 103 writes the target element to the destination register.

[0030] In the embodiment of the present application, the vector shuffle operation performed After the later target element is obtained, the target element can be written to the destination register.

[0031] Optionally, the method of obtaining source elements according to shuffle parameters, performing a vector shuffle operation, and obtaining target elements includes the steps of selecting source elements from a source register according to position information in a source register of source elements required for the vector shuffle operation and the number of source elements required for the vector shuffle operation, and setting all selected source elements as target elements, which can be specifically described in detail with reference to the specific implementation forms below.

[0032] In one specific implementation of the present application, the above step 102 may include the following sub-steps A1 to A3.

[0033] In sub-step A1, position information of source elements in the source register and the number of source elements required for the vector shuffle operation are determined according to the shuffle parameters, and the selected number of source elements is one or more.

[0034] In the present embodiment, the shuffle parameters include parameters for indicating the location information of the source elements in the source register and the number of source elements.

[0035] After receiving an instruction to perform a vector shuffle operation, the CPU may analyze the instruction to obtain shuffle parameters included in the instruction.

[0036] After analyzing and obtaining the shuffle parameters included in the instruction, the position information of the source elements required for the vector shuffle operation in the source register and the number of source elements required for the vector shuffle operation can be determined according to the shuffle parameters. The number of selected source elements may be one or more. In the following example, the case where there are more than one source elements will be described as an example.

[0037] After determining the position information of the source elements in the source register and the number of the source elements according to the shuffle parameters, sub-step A2 is performed.

[0038] In sub-step A2, a source element is selected from the source register according to the determined position information and the number of source elements.

[0039] When determining the position information of the source elements in the source register and the number of the source elements according to the shuffle parameters, a source element is selected from the source register.

[0040] After selecting a source element from the source register according to the determined position information and the number of source elements, perform sub-step A3.

[0041] In sub-step A3, all the selected source elements are determined as target elements.

[0042] After selecting source elements from the source register according to the determined position information and the number of source elements, all the selected source elements can be written as target elements to the destination register.

[0043] In an embodiment of the present application, the shuffle parameters may include an index value and an opcode, and the index value and the opcode are used to select a source element, as will be described in detail with reference to specific implementations below.

[0044] Optionally, the index value is for indicating position information in the source register of each source element required for the vector shuffle operation, the opcode is for indicating an operation on the source register and the destination register, and the above sub-step A2 may include the following sub-steps B1 to B2.

[0045] In sub-step B1, a selection rule for obtaining a source element is determined according to the index value and the opcode.

[0046] In the present embodiment, a selection rule refers to a restrictive condition for reading a source element from a source register.

[0047] After the shuffle parameters are obtained, the selection rules for obtaining source elements from the source registers can be determined according to the index values ​​and opcodes contained in the shuffle parameters. Specifically, there are two types of situations:

[0048] In the first type of scenario, when the number of index values ​​and the number of source elements are different, a grouping method of the source elements can be determined according to the number of index values, and a selection rule for obtaining the source elements can be determined according to the grouping method and the operation code. That is, first, the source elements in the source register are grouped, for example, N adjacent source elements are grouped into one group, and then a selection rule for the source elements is obtained from the grouped elements according to the index value. Usually, N is 4. Naturally, N can be determined according to a specific application scenario such as the number of bits of the source register, and will not be described again here.

[0049] In the second scenario, when the number of index values ​​and the number of source elements are the same, a selection rule can be determined to obtain the source element according to the opcode.

[0050] After determining the selection rule for obtaining the source element according to the index value and the opcode, execute sub-step B2.

[0051] In sub-step B2, the source elements indicated by the respective index values ​​are obtained from the source register in accordance with the selection rule.

[0052] After determining the selection rule for obtaining the source elements according to the index value and the operation code, the source elements indicated by each index value can be obtained from the source register according to the selection rule.

[0053] In practical application, the CPU uses the opcode of the vector shuffle instruction to determine whether the number of index values ​​is the same as the number of source elements, that is, the CPU can determine the grouping status and selection rule of the source elements according to the opcode.

[0054] Optionally, there is a predetermined correspondence between the index value and the address in the destination register. Optionally, writing the target element to the location of the destination register corresponding to the immediate value means determining the location of the index value from the destination register and sequentially storing the source elements at the determined location. Specifically, each source element is obtained by a determined index value, and the source element is written to an address in the destination register corresponding to the determined index value. For example, the source element A is obtained by an index value ui8[1:0] (ui8 represents an immediate value, the immediate value ui8 is an index value representing one group of data, and ui8[1:0] represents a number consisting of the least significant two bits of the immediate value), but the index value ui8[1:0] corresponds to the lowest address in one group of addresses in the destination register, and the obtained source element A is written to the lowest address of the destination register as one target element. As an example, the immediate value ui8 is a group of data including 8 bits, and four index values ​​are formed using the immediate value ui8, with the number formed by every two bits of ui8 being one index value. Meanwhile, the positions and sequence numbers of these index values ​​in ui8 indicate or imply the element positions where the source elements obtained from the index values ​​should be moved to the destination register. For example, if the index value ui8[7:6] is the fourth index value of ui8, the corresponding Source Element is written to the fourth element position of the destination register. Similarly, if ui8[n:n-1] is the (n+1) / 2th index value in ui8, then the corresponding Source Element is the destination register (n+1) / 2 The immediate value can contain other numbers of index values, and accordingly, the value corresponding to the i-th index value in the immediate value is written to the i-th element position. Source Element is written to the i-th element location of the destination register, where i is a positive integer.

[0055] In the prior art, when implementing a SHUF instruction (a type of shuffle instruction), a shuffle instruction can be set according to a shuffle mode to obtain shuffle effects with different functions, and the shuffle mode is determined according to the needs of the application, and usually, the shuffle mode can be called by at least one other instruction, and the shuffle mode can be transferred to the shuffle instruction, or the shuffle mode can be added to a memory, and the shuffle mode can be obtained by accessing the memory during the execution of the shuffle instruction. As can be seen from this, in the prior art, in order to implement shuffle instructions in different shuffle modes, multiple instructions are required or memory access is required, and either preparing multiple instructions or accessing a memory greatly increases the overhead of the entire CPU system when implementing a shuffle instruction. In view of the technical problems in the prior art, in the embodiments of the present application, by adding shuffle parameters (opcode and index value) to an instruction, shuffle instructions in different shuffle modes can be implemented using different shuffle parameters, and there is no need to realize data shuffling with multiple instructions, and there is no need to obtain the shuffle mode by accessing memory, so that the data shuffle operation can be realized with a single shuffle instruction, effectively reducing system overhead.

[0056] Since the index value can be implemented by an immediate value, the opcode may be implemented by a code shown in binary format, or the opcode may be implemented as an identifier that can be converted into a binary code. Therefore, referring to the implementation method of the vector shuffle instruction including the opcode and the index value in Example 1, specific processing methods of the vector shuffle instruction including different opcodes will be described in detail using the following specific Examples 2 to 6.

[0057] Example 2

[0058] In one specific implementation of the present application, the opcode is a first opcode, and the number of the index values ​​and the number of the source elements are different. As shown in Figure 2, the processing method of the vector shuffle instruction can include the following steps 201 to 205.

[0059] In step 201, a command is received that includes a register identifier and a shuffle parameter.

[0060] In this embodiment, the meaning of the command and the parameters contained in the command are as described in the first embodiment, and will not be repeated here.

[0061] Alternatively, the number of source registers is one, ie the source elements come from one register.

[0062] Optionally, the shuffle parameters include an index value and an opcode, the index value being implemented in the form of an immediate value, the opcode being implemented in the form of an identifier that can be converted to a binary code, and the opcode being the first opcode.

[0063] Optionally, the format of the instruction is "opcode destination register, source register, immediate value". According to the format of the instruction, when specifically implemented, the instruction can be expressed as "[X]VS.{B / H / W} vd,vj,ui8", where [X]VS is the instruction name of the first opcode, [X] is an option for distinguishing registers with different bit numbers, {B / H / W} is the data type of the first opcode, B indicates that the data type is byte, H indicates that the data type is halfword, and W indicates that the data type is word, [X]VS.{B / H / W} is the first opcode in an identifier format, vd indicates the destination register, vj indicates the source register, and ui8 indicates the immediate value. Exemplarily, VS.{B / H / W} is a first opcode that can be converted to a binary format, such as a first opcode that converts [X]VS.B to a 01110011100100 binary format. The immediate value may also be a group of data, and index values ​​at different locations in the register may be indicated by different bits of the immediate value ui8, ui8[1:0], ui8[3:2], ui8[5:4], ui8[7:6].

[0064] In step 202, the instruction is executed to organize adjacent elements in the source register into element groups of N1 elements each according to the opcode and the index value, where the data type of the elements is one of a byte, a halfword, or a word, and N1 is a positive integer greater than 0.

[0065] In the embodiment of the present application, when the number of index values ​​and the number of source elements are different, adjacent elements in the source register can be configured as element groups of one group every N1, and the data type of the adjacent elements can be any one of byte, halfword, and word, for example, adjacent word elements in the source register can be configured as element groups of one group every 4. When the number of index values ​​and the number of source elements are different, adjacent elements in the source register can be configured as element groups of one group every N1, and the data type of the adjacent elements can be any one of byte, halfword, and word, and a source element is selected from the element group, and multiple conditions such as N1 being a positive integer greater than 0 are determined as selection rules. For example, by making N1 the number of index values, even if the number of index values ​​is less than the number of source elements, the source elements are grouped according to the difference between N1 and the number of source elements, so that the number of index values ​​is equal to the number of source elements in each group and each index value corresponds one-to-one to the source element in the group.

[0066] Adjacent elements are elements whose positions are adjacent in sequence in the source register, and the element addresses in adjacent element groups may be partially the same or completely different, and the element addresses are the position information of the elements in the register. When there are elements with the same position information between adjacent element groups, the maximum number of elements with the same position information for each of the two adjacent element groups is N1-1. Furthermore, adjacent elements are cross-adjacent elements in the source register. For example, when the opcode is a first opcode, the data type is a byte, a halfword, or a word, and the source register includes eight elements, namely, element A1, element A2, element A3, element A4, element A5, element A6, element A7, and element A8. The position information of the above elements is adjacent in the order shown, and N1=4. The N1 elements may be adjacent in the order shown, or may be adjacent elements crossing each other. For example, when N1 is 4, the source register includes eight elements, namely, element A1, element A2, element A3, element A4, element A5, element A6, element A7, and element A8. The N1 elements may be adjacent elements A2 to A5, or may be adjacent elements crossing each other, such as element A1, element A3, element A5, and element A7.

[0067] Based on the above embodiment, organizing the adjacent elements in the source register into element groups, one group for every N1 elements, includes two scenarios.

[0068] In the first type of scene, elements A1 to A4 are configured as an element group of one group, and elements A5 to A8 are configured as an element group of the other group. There are no elements with the same position information between the two element groups. In the second type of scenario, elements A1 to A4 are configured as an element group of one group, and elements A2 to A5 are configured as an element group of the other group, and there are three elements with the same position information between the two element groups (i.e., element A2, element A3, and element A4).Otherwise, if it is satisfied that the maximum value of the number of elements with the same position information for each of two adjacent element groups is N1-1, elements A3 to A6 may be selected as an element group of the other group, or elements A4 to A7 may be selected as an element group of the other group, which will not be described again here.

[0069] Alternatively, the data types of the elements included in each of the divided element groups are the same, and the data types of the elements included in different element groups are the same. For example, the divided element groups include element group 1, element group 2, and element group 3, and the data types of the elements included in element group 1, element group 2, and element group 3 are all byte, or the data types of the elements included in element group 1, element group 2, and element group 3 are all halfword, or the data types of the elements included in element group 1, element group 2, and element group 3 are all word.

[0070] Furthermore, different element groups use the same index value, or different element groups use index values ​​that are not completely the same. For example, when different element groups use the same index value, element group 1 to element group 4 all use the same index value UI8, but when different element groups use index values ​​that are not completely the same, element group 1 and element group 2 use UI8a as the index value for source element selection, element group 3 and element group 4 use UI8b as the index value for source element selection, and UI8a and UI8b indicate different positions of UI8 and indicate different index values.

[0071] In another example, multiple element groups may have elements with the same data type for each element in the element group, but different element groups (e.g., 1 and The elements in element group 1) have different data types. Furthermore, the number of elements in each element group can be the same or different. For example, if element group 1 has four elements and element group 2 has two elements, the same immediate value ui8 will provide four index values ​​to element group 1 and two index values ​​to element group 2.

[0072] It should be understood that the above examples are merely examples for better understanding of the technical solutions according to the embodiments of the present application, and do not uniquely limit the embodiments of the present application.

[0073] After arranging adjacent elements in the source register into element groups, one group for every N1 elements, step 203 is performed.

[0074] In step 203, an element in each element group is determined as an initial source element.

[0075] In the embodiment of the present application, adjacent elements in the source register are organized into element groups, one group for every N1 elements, and then an element in each element group can be determined as an initial source element. The initial source element refers to an initial element for selecting a source element.

[0076] After determining the elements of each element group as the initial source elements, step 204 is performed.

[0077] In step 204, the source elements indicated by each index value are obtained from the initial source elements, respectively, and the number of source elements selected from each element group is n1.

[0078] In the embodiment of the present application, after the initial source elements are determined, the source elements indicated by each immediate value can be obtained from the initial source elements, that is, the corresponding source elements can be selected from the element groups according to the immediate values. The number of source elements selected from each element group is n1, where n1 is a positive integer greater than 0.

[0079] Optionally, there is a predetermined correspondence between the immediate value and an element position in the element group of each group, the element position may be an element address or may be a sequence bit in the element group of the element, the sequence bit representing a position number of the element in the element group.

[0080] Optionally, respectively obtaining source elements indicated by each immediate value from the initial source element, i.e. respectively obtaining elements at element positions corresponding to the immediate values ​​from each element group, and determining the obtained elements as source elements. The number of source elements selected from different element groups is the same.

[0081] In a specific implementation, when the opcode is the first opcode, N1=4, and the data type is byte, halfword, or word, the number of initial source elements included in each element group is the same, which is 4, and a source element corresponding to the immediate value is selected from each element group, n1 is 4, and N1=n1. For example, if the immediate value indicates an element address of 3, an element with an address of 3 is selected from each element group, and all selected elements are determined as source elements; for example, if the immediate value indicates that the sequence bit is 3, the third element counting from the first element is selected from each element group, and all selected elements are determined as source elements.

[0082] Alternatively, N1 may or may not be equal to n1, but if N1=n1, step 204 may not be performed and the elements in each element group in step 203 may be directly selected elements.

[0083] Furthermore, the number of source elements selected from each element group is four, and the data type of the source elements can be a byte, a halfword, or a word, but typically the data type of each selected source element is the same.

[0084] In step 205, the selected source element is determined as a target element, and the target element is written to the destination register at a location corresponding to the index value.

[0085] In the embodiment of the present application, there is a predetermined correspondence between an immediate value and an address in a destination register, and optionally writing a target element to a location in the destination register corresponding to the immediate value, i.e., determining from the destination register a location corresponding to the immediate value, and sequentially storing a source element to the determined location.

[0086] Furthermore, in one possible solution, a step of creating an intermediate vector is added between step 201 and step 202, specifically, before selecting source elements from the source register according to the determined position information and the number of source elements, an intermediate vector can be created, the intermediate vector includes at least one intermediate vector parameter, and the number of the intermediate vector parameters is equal to the number of the target elements. Based on the created intermediate vector, after step 204, i.e. after selecting source elements from the source register, each selected source element is stored in a corresponding intermediate vector parameter of the intermediate vector, respectively, and the intermediate vector parameters have a one-to-one corresponding relationship with the selected source element, and step 205, i.e., writing the content of each intermediate vector parameter to a corresponding position of the destination register according to the immediate value.

[0087] Alternatively, the intermediate vector may be created depending on the source register, the type of the source register, etc.

[0088] Alternatively, the number of intermediate vector parameters in the intermediate vector is the same as the number of target elements, and there is a predetermined correspondence between the position of each target element in the destination register and each intermediate vector parameter in the intermediate vector according to an index value. When the source elements are grouped, the content of each intermediate vector parameter is written to the corresponding position of the destination register according to the immediate value, that is, setting a parameter i, where i represents a constant, and the value of i ranges from 0 to n-1, and n is determined according to the number of register bits and the data type, determining a source element in a source register indexed by each intermediate vector parameter of the intermediate vector according to N1 and i, and taking values ​​from 0 to n-1 for i in a cyclic manner, while writing source elements corresponding to different index values ​​in the intermediate vector. Destination Register The intermediate vector is written to the target element position corresponding to the index in. Specifically, [N1i], [N1i+1], [N1i+2] ... [N1i+N1-1] indicate different positions, and the intermediate vector can be expressed as "intermediate vector = {VR[source register].data type [N1i+N1-1], ... VR[source register].data type [N1i]}", where i represents a constant, the value of i ranges from 0 to n-1, and n is determined according to the number of register bits and the data type. For example, if the number of register bits is 128 bits and the data type is byte, n is 4, if the number of register bits is 128 bits and the data type is halfword, n is 2, and if the number of register bits is 128 bits and the data type is word, n is 1.

[0089] Based on the above intermediate vector solution, for example, if the first opcode is VS.B and vj is a source register, create an intermediate vector as follows: vec0={VR[vj].B[4i+3],VR[vj].B[4i+2],VR[vj].B[4i+1],VR[vj].B[4i]}, where VR[vj].B[4i+3],VR[vj].B[4i+2],VR[vj].B[4i+1],VR[vj].B[4i] are intermediate vector parameters, i represents a constant, [4i+0], [4i+1], [4i+2], [4i+3] represent four consecutive positions of a register, and the value of i ranges from 0 to 3. Writing the content of each of the intermediate vector parameters to the corresponding position of the destination register vd can be shown as follows: VR[vd].B[4i+0] = vec0.B[ui8[1:0]] VR[vd].B[4i+1] = vec0.B[ui8[3:2]] VR[vd].B[4i+2] = vec0.B[ui8[5:4]] VR[vd].B[4i+3] = vec0.B[ui8[7:6]]

[0090] ui8[1:0], ui8[3:2], ui8[5:4], and ui8[7:6] are all immediate values ​​that represent index values ​​corresponding to the intermediate vector. Specifically, the least significant two bits (ui8[1:0]) of the immediate value ui8 represent the index in the intermediate vector of the first target element, the third and fourth bits (ui8[3:2]) of the immediate value ui8 represent the index in the intermediate vector of the second target element, the fifth and sixth bits (ui8[5:4]) of the immediate value ui8 represent the index in the intermediate vector of the third target element, and the seventh and eighth bits (ui8[7:6]) of the immediate value ui8 represent the index in the intermediate vector of the fourth target element.

[0091] Similarly, when the data types are halfword and word, the intermediate vector and index scheme is the same as the above example, and a vector shuffle operation needs to be implemented for the two intermediate vectors when the opcode instruction name is XVS.{B / H / W}. Exemplarily, when the first opcode is XVS.B, the intermediate vector is shown as follows: vec0={VR[vj].B[4i+3],VR[vj].B[4i+2],VR[vj].B[4i+1],VR[vj].B[4i]} vec1={VR[vj].B[4i+19],VR[vj].B[4i+18],VR[vj].B[4i+17],VR[vj].B[4i+16]}

[0092] The intermediate vectors are vec0 and vec1, VR[vj].B[4i+3], VR[vj].B[4i+2], VR[vj].B[4i+1], VR[vj].B[4i] are the intermediate vector parameters of intermediate vector vec0, VR[vj].B[4i+19], VR[vj].B[4i+18], VR[vj].B[4i+17], VR[vj].B[4i+16] are the intermediate vector parameters of intermediate vector vec1, B indicates that the data type is byte, i indicates the position of the element in the register, [4i+0], [4i+1], [4i+2], and [4i+3] indicate the elements in four consecutive positions in the register, and [4i+16], [4i+17], [4i+18], and [4i+19] indicate the elements in four consecutive positions in the register.

[0093] For example, if the first opcode is XVS.B, the data type is byte, and N1 is 4, the vector shuffle instruction "XVS.B vd,vj,ui8" reads four adjacent byte elements from vector register vj to form a group of elements and shuffles them, and then writes the result into vector register vd; if the first opcode is VS.H, the data type is halfword, and N1 is 4, the vector shuffle instruction "VS.H vd,vj,ui8" reads four adjacent halfword elements from vector register vj to form a group of elements and shuffles them, and then writes the result into vector register vd; if the first opcode is VS.W, the data type is word, and N1 is 4, the vector shuffle instruction "VS.W vd,vj,ui8" reads four adjacent halfword elements from vector register vj to form a group of elements and shuffles them, and then writes the result into vector register vd; "vd,vj,ui8" reads four adjacent word elements from vector register vj, shuffles them into a group of elements, and writes the result into vector register vd.

[0094] It should be understood that the above examples are merely examples for better understanding of the technical solutions according to the embodiments of the present application, and do not uniquely limit the embodiments of the present application.

[0095] In an embodiment of the present application, a shuffle parameter including an index value and an opcode is added to the vector shuffle instruction, and according to the index value and the opcode, Source Element Since the shuffle operation is implemented when the number of index values ​​is different from that of the vector shuffle instruction, the technical solution of the present application can be implemented as follows: Source Element This implements shuffle operations when the number of index values ​​differs from the number of vectors, and does not require additional instructions to convey the shuffle mode, nor does it require accessing memory to obtain the shuffle mode, effectively reducing system overhead and improving the execution efficiency of vector shuffle operations.

[0096] Example 3

[0097] In one particular implementation of the present application, the opcode is a second opcode, and the number of the index values ​​is the same as the number of the source elements, and as shown in FIG. 3, the processing method of the vector shuffle instruction includes steps 301 to 303.

[0098] In step 301, a command is received that includes a register identifier and a shuffle parameter.

[0099] In the embodiment of the present application, the meaning of the command and the parameters contained in the command are as described in the first and second embodiments, and will not be described again here.

[0100] Alternatively, the number of source registers is two, i.e., the source elements come from two different registers, and if the number of source registers is multiple, each source register identifier of all the source registers is different from the destination register identifier, or if the number of source registers is multiple, all the source registers have one source register identifier that is the same as the destination register identifier.

[0101] Alternatively, the shuffle parameters include an index value and an opcode, the index value is implemented in the form of an immediate value, the opcode is implemented in the form of an identifier that can be converted into a binary code, and the opcode is a second opcode. Illustratively, when the opcode is the second opcode, the source registers include a first source register and a second source register, and the destination register is the first source register or the second source register.

[0102] Optionally, the format of the instruction is "opcode destination register, source register, immediate value". According to the format of the instruction, when specifically implemented, the instruction can be expressed as "[X]VS.D vd,vj,ui8", where [X]VS is the instruction name of the second opcode, D is the data type of the second opcode, D indicates that the data type is double word, [X]VS.D is the second opcode in the identifier format, vd represents the destination register, vj and vd represent the source registers, and ui8 represents the immediate value. Exemplarily, VS.D is a second opcode that can be converted to a binary format, such as a second opcode that converts VS.D to a 01110011100111 binary format. The immediate value may also be a group of data, and the index values ​​may be indicated by different bits of the immediate value ui8, ui8[1:0], ui8[3:2], ui8[5:4], ui8[7:6].

[0103] In step 302, the instruction is executed to generate M bits every N2 bits in the source register according to the opcode and the index value. N2 The data type of each of the elements is double word, and M each of N2 bits is N2 The number of source elements selected from the elements is n2, and N2, M N2 , and n2 are all positive integers greater than 0.

[0104] In the embodiment of the present application, the number of index values ​​and the number of source elements are the same, the opcode is a second opcode, and M N2 From the elements, obtain the source element indicated by each index value, the data type of the element is doubleword, and M N2 The number of source elements selected from the elements is n2, and N2, M N2 , and n1 are both positive integers greater than 0, and so on, are determined as selection rules.

[0105] Optionally, each index value has a predetermined correspondence with an element position in each source register, which may be an element address. N2 The process of obtaining the source elements indicated by the respective index values ​​from the elements is carried out by: N2 1. From the elements, obtain a first source element indicated by each index value, and in the second source register, N2 The second source element indicated by each index value is obtained from each of the elements, and the first source element and the second source element are determined as the finally selected source elements. n2 The elements may be adjacent elements in sequence or adjacent elements crosswise, for example, M n2 Assume that M is 4 and the source register contains eight elements: element A1, element A2, element A3, element A4, element A5, element A6, element A7, and element A8. N4 These elements may be elements A2 to A5, or may be adjacent intersecting elements such as element A1, element A3, element A5, and element A7.

[0106] For example, if the second operand is [X]VS.D, N 2 is 128, M N2 is 4 and n2 is 2.

[0107] In concrete implementation, there are two source registers, namely, a first source register and a second source register, and in the source register, M N2 The acquisition of each of the source elements indicated by the respective index values ​​from the elements includes the step of: N2’ Retrieving a source element indicated by a first index value (e.g., ui8[1:0]) from among the elements; and N2’obtaining a source element from the elements indicated by a second index value (e.g., ui8[3:2]); N2’ is M N2 , the number of source elements selected from the first source register is n2 / 2, and the number of source elements selected from the second source register is n2 / 2. When the number of source registers is multiple, each source register performs vector shuffling according to different bits of the immediate value, that is, the bits of the immediate value corresponding to different source registers are different, and which bits of the immediate value are indexed according to the specific situation is determined according to the specific situation and will not be described again here.

[0108] In step 303, the selected source element is determined as a target element, and the target element is written to the destination register at a location corresponding to the index value.

[0109] In the embodiment of the present application, there is a predetermined correspondence between an immediate value and an address in a destination register, and optionally writing a target element to a location in the destination register corresponding to the immediate value, i.e., determining from the destination register a location corresponding to the immediate value, and sequentially storing a source element to the determined location.

[0110] Further, in one possible solution, a step of creating an intermediate vector is added between step 301 and step 302, specifically, before selecting source elements from the source register according to the determined position information and the number of source elements, an intermediate vector can be created, the intermediate vector includes at least one intermediate vector parameter, and if an element group exists, the number of the intermediate vector parameters is equal to the number of the element group, but if an element group does not exist, the number of the intermediate vector parameters is equal to the number of the source elements. Based on the created intermediate vector, after step 302, i.e., after selecting source elements from the source register, each selected source element is stored in a corresponding intermediate vector parameter of the intermediate vector, respectively, and the intermediate vector parameters have a one-to-one corresponding relationship with the selected source element, and step 303, i.e., writes the content of each intermediate vector parameter to a corresponding position of the destination register according to the immediate value. The method of creating an intermediate vector is the same as the method according to the second embodiment, and will not be repeated here.

[0111] Optionally, writing the content of each of the intermediate vector parameters to a corresponding location of the destination register in response to the immediate value comprises performing an operation of writing, for each of the intermediate vector parameters, the content of the intermediate vector parameter to a location of the destination register indicated by an index value corresponding to the intermediate vector parameter.

[0112] Based on the above intermediate vector solution, for example, if the second opcode is VS.D and the instruction format is VS.D vd,vj,ui8, where vj and vd are source registers, creating an intermediate vector vec0={VR[vj], VR[vd]} and writing the content of each of the intermediate vector parameters to the corresponding position of the destination register vd can be shown as follows: VR[vd].D[0] = vec0.D[ui8[1:0]] VR[vd].D[1] = vec0.D[ui8[3:2]]

[0113] Both ui8[1:0] and ui8[3:2] are immediate values ​​that represent index values ​​corresponding to registers. Specifically, the least significant two bits of the immediate value ui8 (ui8[1:0]) represent the index in the source register of the first target element, and the third and fourth bits of the immediate value ui8 (ui8[3:2]) represent the index in the source register of the second target element.

[0114] If the second opcode is XVS.D, then a vector shuffle operation needs to be implemented on the two intermediate vectors. Illustratively, the intermediate vectors are shown as follows: vec0 = {XR[xj][127:0], XR[xd][127:0]} vec1 = {XR[xj][255:128], XR[xd][255:128]}

[0115] The intermediate vectors are vec0 and vec1, XR[xj][127:0], XR[xd][127:0] represent the intermediate vector parameters of vec0, XR[xj][255:128], XR[xd][255:128] represent the intermediate vector parameters of vec1, and D represents the data type is double word and the width is 64 bits.

[0116] For example, the second opcode is VS.D, the data type is doubleword, N2 is 128, and M N2A vector shuffle instruction "VS.D vd,vj,ui8" indicates that for vector register vj and vector register vd, select two doubleword elements from four doubleword elements in each 128-bit region according to the content of the immediate value, and write the result into the corresponding 128-bit region of vector register vd, where the second opcode is XVS.D, the data type is doubleword, N2 is 128, and M N2 When n is 4 and n2 is 2, the vector shuffle instruction “XVS.D vd,vj,ui8” indicates to read two double word elements from four double word elements in each 128-bit of vector register xj and vector register xd according to the content of the immediate value, and then write the read double word elements into the 128-bit corresponding to xd.

[0117] In the embodiment of the present application, among the two source registers, there is one source register that is the same as the destination register, that is, there is one register that is both a source register and a destination register. By adopting the above technical solution, half of the elements in the destination register can be covered every time a shuffle instruction is executed, and can be applied to software application scenarios that require the execution of corresponding operations.

[0118] In an embodiment of the present application, a shuffle parameter including an index value and an opcode is added to the vector shuffle instruction, and according to the index value and the opcode, Source Element Since the number of index values ​​is the same as that of the vector shuffle operation, the data type is double word, and the register is 128 bits, the technical solution of the present application can be implemented by one vector shuffle instruction. Source ElementThe shuffle operation is implemented when the number of index values ​​is different from the number of index values ​​and the data type is double word, and there is no need to add any other instructions to transmit the shuffle mode, and there is no need to access memory to obtain the shuffle mode, which effectively reduces system overhead and improves the execution efficiency of the vector shuffle operation.

[0119] Example 4

[0120] In one particular implementation of the present application, the opcode is a third opcode, the index values ​​include a first index value, a second index value, a third index value, and a fourth index value, and the first index value, the second index value, the third index value, and the fourth index value each index the same or different positions, and as shown in FIG. 4, the processing method of the vector shuffle instruction can include steps 401 to 405.

[0121] In step 401, a command is received that includes a register identifier and a shuffle parameter.

[0122] In the present embodiment, the meaning of the command and the parameters contained in the command are described in Example 1, Example 2, and Example 3, and will not be repeated here.

[0123] Alternatively, the number of source registers is two, i.e. the source elements come from two different registers, and if the number of source registers is multiple, each source register identifier of all the source registers is different from the destination register identifier, or if the number of source registers is multiple, all the source registers have one source register identifier that is the same as the destination register identifier.

[0124] Alternatively, the shuffle parameters include an index value and an opcode, the index value is implemented in the form of an immediate value, the opcode is implemented in the form of an identifier that can be converted into a binary code, and the opcode is a third opcode. Illustratively, when the opcode is the third opcode, the source registers include a first source register and a second source register, and the destination register is the first source register or the second source register.

[0125] Alternatively, the format of the instruction is "opcode destination register, source register, immediate value". According to the format of the instruction, when specifically implemented, the instruction can be expressed as "[X]VP.W vd / xd,vj / xj,ui8", where [X]VP is the instruction name of the third opcode, W is the data type of the third opcode, W indicates that the data type is word, [X]VP.W represents the third opcode in the identifier format, vd / xd represents the destination register, vj and vd represent the source register (or xj and xd represent the source register), and ui8 represents the immediate value. Exemplarily, VP.W is a third opcode that can be converted to a binary format, such as a third opcode that converts VP.W to a 01110011111001 binary format. Also, the immediate value may be a group of data, and the index values ​​may be indicated by different bits of the immediate value ui8, ui8[1:0], ui8[3:2], ui8[5:4], ui8[7:6].

[0126] In step 402, the instruction is executed to store M every N3 bits in the first source register according to the opcode and the index value. N3 The source elements indicated by the first index value and the second index value are obtained from the elements, and M N3 From this element, source elements indicated by the third and fourth index values ​​are obtained, respectively.

[0127] The data type of the element is word, and M N3 The number of source elements selected from the elements is n3, and N3, M N3 , and n3 are all positive integers greater than 0.

[0128] In the embodiment of the present application, the index value includes four index values, namely, a first index value, a second index value, a third index value, and a fourth index value, and the first index value, the second index value, the third index value, and the fourth index value index different positions. When the number of source registers is more than one, each source register performs vector shuffling according to different bits of the immediate value, that is, the bits of the immediate value corresponding to different source registers are different, and which bits of the immediate value are indexed according to the specific situation is determined according to the specific situation, and will not be described repeatedly here. For example, when the third opcode is [X]VP.W, the first index value is ui8[1:0], the second index value is ui8[3:2], the third index value is ui8[5:4], and the fourth index value is ui8[7:6].

[0129] Also, the number of index values ​​and the number of source elements are the same, and M N3 3 elements, obtain a source element indicated by each index value, the data type of the element being word, and M bits of N3 each are N3 The number of source elements selected from the elements is n3, and N3, M N3 , and n3 are all positive integers greater than 0. N3 The elements may be adjacent elements in sequence or adjacent elements crosswise, for example, M N3 Assume that M is 4 and the source register contains eight elements: element A1, element A2, element A3, element A4, element A5, element A6, element A7, and element A8. N3These elements may be elements A2 to A5, or may be adjacent intersecting elements such as element A1, element A3, element A5, and element A7.

[0130] Alternatively, there is a predetermined correspondence between the index value and an element position in each source register, and the element position may be an element address. N3 The step of obtaining the source elements indicated by the respective index values ​​from the elements includes the steps of: N3 obtaining source elements indicated by a first index value and a second index value from the elements; and storing M N3 The second step is to obtain, from these elements, source elements indicated by the third and fourth index values, respectively.

[0131] For example, if the third opcode is [X]VP.W, N3 is 128, and M N3 is 4 and n3 is 2.

[0132] After obtaining the source elements indicated by the index values ​​from the MN3 elements of the source register for every N3 bits, step 403 and step 404 are executed.

[0133] In step 403, the first index value To The source element indicated by the second index value is determined as the first target element. To The source element indicated by the above is determined as the second target element.

[0134] In the present embodiment, a first index value selected from a first source register To determining the source element indicated by the second index value selected from the first source register as the first target element; ToThe source element indicated by the above is determined as the second target element.

[0135] In step 404, the third index value To The source element indicated by the fourth index value is determined as the third target element. To The indicated source element is determined as the fourth target element.

[0136] In the present embodiment, a third index value selected from the second source register To determining the source element indicated by the fourth index value selected from the second source register as the third target element; To The indicated source element is determined as the fourth target element.

[0137] In the embodiment of the present application, step 403 and step 404 may be steps executed simultaneously or sequentially, and the execution order is not limited. After step 403 and step 404 are all executed, step 405 is executed.

[0138] In step 405, the first target element and the second target element are written to a first location of the destination register, and the third target element and the fourth target element are written to a second location of the destination register.

[0139] In the embodiment of the present application, there is a predetermined correspondence between an immediate value and an address in a destination register, and optionally writing a target element to a location in the destination register corresponding to the immediate value, i.e., determining from the destination register a location corresponding to the immediate value, and sequentially storing a source element to the determined location.

[0140] In an embodiment of the present application, when the opcode is a third opcode, after obtaining a first target element, a second target element, a third target element, and a fourth target element, the first target element and the second target element can be written to a first position of a destination register, and the third target element and the fourth target element can be written to a second position of the destination register.

[0141] For example, the third opcode is VP.W / XVP.W (both are abbreviated as [X]VP.W), the data type is word, N3 is 128, and M N3 When n is 4 and n3 is 2, the vector shuffle instruction "[X]VP.W vd,vj,ui8" indicates that two elements are selected from the four word elements in every 128 bits of vector register vj / xj using the values ​​of ui8[1:0] and ui8[3:2] as index values, and two elements are written to the 0th and 1st word elements of the 128-bit corresponding to vector register vd / xd, and that two elements are selected from the four word elements in every 128 bits of vector register vd / xd using the values ​​of ui8[5:4] and ui8[7:6] as index values, and two elements are selected from the four word elements in every 128 bits of vector register vd / xd, and two elements are written to the second and third word elements of the 128-bit corresponding to vector register vd / xd.

[0142] In an embodiment of the present application, a shuffle parameter including an index value and an opcode is added to the vector shuffle instruction, and according to the index value and the opcode, Source Element Since the number of index values ​​is the same as that of the vector shuffle operation when the data type is word, the technical solution of the present application can be implemented by one vector shuffle instruction. Source ElementThe shuffle operation is implemented when the number of index values ​​is the same as that of the data type of a word, and there is no need to add any additional instructions to transmit the shuffle mode, and there is no need to access memory to obtain the shuffle mode, which effectively reduces system overhead and improves the execution efficiency of the vector shuffle operation.

[0143] Example 5

[0144] In one particular implementation of the present application, the opcode is a fourth opcode, and the number of the index values ​​is the same as the number of the source elements, and as shown in FIG. 5, the processing method of the vector shuffle instruction can include steps 501 to 503.

[0145] In step 501, a command is received that includes a register identifier and a shuffle parameter.

[0146] In the embodiments of the present application, the meaning of the commands and the parameters contained in the commands are as described in the first to fourth embodiments, and will not be described again here.

[0147] Alternatively, the number of source registers is one, ie the source elements come from one register.

[0148] Alternatively, the shuffle parameters include an index value and an opcode, the index value being implemented in the form of an immediate value, the opcode being implemented in the form of an identifier that can be converted into a binary code, and the opcode being a fourth opcode.

[0149] Alternatively, the format of the instruction is "opcode destination register, source register, immediate value". According to the format of the instruction, when specifically implemented, the instruction can be represented as XVP.D xd,xj,ui8, where XVP is the instruction name of the fourth opcode, D is the data type of the fourth opcode, and D indicates that the data type is double word, XVP.D is the fourth opcode in the identifier format, xd represents the destination register, xj represents the source register, and ui8 represents the immediate value. Exemplarily, XVP.D is a fourth opcode that can be converted to a binary format, such as a fourth opcode that converts XVP.D to a 01110111111010 binary format. Also, the immediate value may be a group of data, and the index value can be represented by different bits of the immediate value ui8, ui8[1:0], ui8[3:2], ui8[5:4], ui8[7:6].

[0150] In step 502, the instruction is executed to generate a new value in the source register according to the opcode and the immediate value. N4 The data type of the elements is double word, the number of selected source elements is n4, and the M N4 and n4 are both positive integers greater than 0.

[0151] In an embodiment of the present application, the opcode is a fourth opcode, and the fourth opcode can be used to indicate obtaining an element of a double word data type from a source register. The number of index values ​​and the number of source elements are the same, and M N4 n4 elements, a source element indicated by each index value is obtained from the n4 elements, the data type of the element is double word, the number of selected source elements is n4, N4 and n4 are both positive integers greater than 0. N4 The elements may be adjacent elements in sequence or adjacent elements crosswise, for example, M N4Assume that M is 4 and the source register contains eight elements: element A1, element A2, element A3, element A4, element A5, element A6, element A7, and element A8. N4 These elements may be elements A2 to A5, or may be adjacent intersecting elements such as element A1, element A3, element A5, and element A7.

[0152] Optionally, there is a predetermined correspondence between the index value and an element position in each source register, and the element position may be an element address. After determining a selection rule according to the fourth opcode and the index value, in the source register, N4 The source elements indicated by the respective index values ​​can be obtained from the n4 elements, respectively, the data type of the obtained source elements is double word, the number of selected source elements is n4, and M N4 and n4 are both positive integers greater than 0.

[0153] For example, if the fourth opcode is XVP.D, N4 is 4 and n4 is 4.

[0154] In step 503, the selected source element is determined as a target element, and the target element is written to the destination register at a location corresponding to the index value.

[0155] In the embodiment of the present application, there is a predetermined correspondence between an immediate value and an address in a destination register, and optionally writing a target element to a location in the destination register corresponding to the immediate value, i.e., determining from the destination register a location corresponding to the immediate value, and sequentially storing a source element to the determined location.

[0156] For example, the fourth opcode is XVP.D, the data type is doubleword, and M N4When n is 4 and n4 is 4, the vector shuffle instruction “XVP.D xd,xj,ui8” indicates that the values ​​of ui8[1:0], ui8[3:2], ui8[5:4], and ui8[7:6] are used as index values ​​to select source elements indicated by each index value from the four double word elements of vector register xj, and to sequentially write the source elements to the four double word elements of vector register xd.

[0157] In an embodiment of the present application, a shuffle parameter including an index value and an opcode is added to the vector shuffle instruction, and according to the index value and the opcode, Source Element Since the number of index values ​​is the same as that of the vector shuffle operation, the data type is double word, and the register is 256 bits, the technical solution of the present application can be implemented by one vector shuffle instruction. Source Element The shuffle operation is implemented when the number of index values ​​is the same as that of the data type of a double word and the register is 256 bits. There is no need to add any additional instructions to transmit the shuffle mode, and there is no need to access memory to obtain the shuffle mode, which effectively reduces system overhead and improves the execution efficiency of the vector shuffle operation.

[0158] Example 6

[0159] In one particular implementation of the present application, the opcode is a fifth opcode, the index value includes a first index value and a third index value, the first index value and the third index value respectively index different positions, and the source register includes a first source register and a second source register, and as shown in FIG. 6, the processing method of the vector shuffle instruction may include steps 601 to 603.

[0160] In step 601, a command is received that includes a register identifier and a shuffle parameter.

[0161] In the embodiments of the present application, the meaning of the commands and the parameters contained in the commands are as described in the first to fifth embodiments, and will not be described again here.

[0162] Alternatively, the number of source registers is two, i.e. the source elements come from two different registers, and if the number of source registers is multiple, each source register identifier of all the source registers is different from the destination register identifier, or if the number of source registers is multiple, all the source registers have one source register identifier that is the same as the destination register identifier.

[0163] Alternatively, the shuffle parameters include an index value and an opcode, the index value is implemented in the form of an immediate value, the opcode is implemented in the form of an identifier that can be converted into a binary code, and the opcode is a fifth opcode. Exemplarily, when the opcode is the fifth opcode, the source registers include a first source register and a second source register, and the destination register is the first source register or the second source register.

[0164] Alternatively, the format of the instruction is "opcode destination register, source register, immediate value". According to the format of the instruction, when specifically implemented, the instruction can be represented as XVP.Q vd / xd,vj / xj,ui8, where XVP is the instruction name of the fifth opcode, Q is the data type of the fifth opcode, Q indicates that the data type is 4 words, XVP.Q is the fifth opcode in the identifier format, xd represents the destination register, xj and xd represent the source registers, and ui8 represents the immediate value. Exemplarily, XVP.Q is a fifth opcode that can be converted to a binary format, such as the fifth opcode that converts XVP.Q to 01110111111011 binary format. Also, the immediate value can be a group of data, and the index value can be represented by different bits of the immediate value ui8, ui8[1:0], ui8[5:4].

[0165] In step 602, the instruction is executed to store M in the first source register according to the opcode and the immediate value. N5 1. From the elements, obtain a first source element indicated by a first index value, and in the second source register, N5 From these elements, 3 a second source element indicated by an index value of n5, a data type of the element being quad-word, and the number of selected source elements being n5, where n5 is a positive integer greater than 0.

[0166] In an embodiment of the present application, the opcode may be a fifth opcode, and the fifth opcode may be used to indicate that an element of a data type of quadword is to be obtained from a source register. The index value includes two index values, a first index value and a third index value, and the first index value and the third index value each index different positions, and the first index value and the third index value each represent different bits of the same immediate value, for example, the first index value represents the lower bit of the immediate value ui8, the third index value represents the upper bit of the immediate value ui8, and the first index value further represents the least significant two bits of the immediate value ui8, and the third index value further represents the next least significant two bits of the immediate value ui8. Exemplarily, when the fifth opcode is XVP.Q, the first index value is ui8[1:0], and the third index value is ui8[5:4]. When the number of source registers is multiple, each source register performs vector shuffling according to different bits of the immediate value, that is, the bits of the immediate value corresponding to different source registers are different, and which bits of the immediate value are indexed according to the specific situation is determined according to the specific situation, and will not be described again here.

[0167] In addition, the number of index values ​​and the number of source elements are the same, and in the first source register, M N5 1. From the elements, obtain a first source element indicated by a first index value, and in the second source register, N5 From these elements, 3 The second source element indicated by the index value of M is obtained, and the data type of the element is 4-word, the number of selected source elements is n5, and n5 is a positive integer greater than 0, and so on, are determined as selection rules. N5 The elements may be adjacent elements in sequence or adjacent elements crosswise, for example, M N5Assume that M is 4 and the source register contains eight elements: element A1, element A2, element A3, element A4, element A5, element A6, element A7, and element A8. N5 These elements may be elements A2 to A5, or may be adjacent intersecting elements such as element A1, element A3, element A5, and element A7.

[0168] Optionally, there is a predetermined correspondence between the index value and an element position in each source register, and the element position may be an element address. After determining the selection rule according to the fifth opcode and the index value, in the first source register, N5 1. From the elements, obtain a first source element indicated by a first index value, and in a second source register, N5 The number of source elements selected from the first source register is n. 5 / 2, and the number of source elements selected from the second source register is n 5 When the number of source registers is more than one, each source register performs vector shuffling according to different bits of the immediate value, that is, the bits of the immediate value corresponding to different source registers are different, and which bits of the immediate value are indexed according to the specific situation is determined according to the specific situation, and will not be described again here.

[0169] For example, if the fifth opcode is XVP.Q, then M N5 is 2, and n 5 is 2.

[0170] After obtaining the first source element and the second source element, step 603 is performed.

[0171] In step 603, the first source element and the second source element are determined as target elements and written to the corresponding locations of the destination register.

[0172] In the embodiment of the present application, there is a predetermined correspondence between an immediate value and an address in a destination register, and optionally writing a target element to a location in the destination register corresponding to the immediate value, i.e., determining from the destination register a location corresponding to the immediate value, and sequentially storing a source element to the determined location.

[0173] In an embodiment of the present application, when the opcode is a fifth opcode, after obtaining a first source element and a second source element, the first source element can be determined as a target element and written to a first location of a destination register, and the second source element can be determined as a target element and written to a second location of the destination register, where the first location and the second location are determined by index values, respectively.

[0174] For example, the fifth opcode is XVP.Q, the data type is 4 words, and M N5 is 2, n 5 When ui8[1:0], ui8[5:4] is 2, the vector shuffle instruction “XVP.Q xd,xj,ui8” indicates that one source element is selected from two 4-word elements of vector register xj according to the values ​​of ui8[1:0], ui8[5:4], and one source element is selected from two 4-word elements of vector register xd, and the selected two source elements are written to two 4-word elements of vector register xd according to the index value.

[0175] In an embodiment of the present application, a shuffle parameter including an index value and an opcode is added to the vector shuffle instruction, and according to the index value and the opcode, Source Element Since the number of index values ​​is the same as that of the vector shuffle operation, and the data type is 4-word, the technical solution of the present application can be implemented by one vector shuffle instruction. Source ElementThe shuffle operation is implemented when the number of index values ​​is the same as that of the vector and the data type is 4-word. There is no need to add any additional instructions to transmit the shuffle mode, and there is no need to access memory to obtain the shuffle mode, which effectively reduces system overhead and improves the execution efficiency of the vector shuffle operation.

[0176] Example 7

[0177] Please refer to FIG. 7, which shows a structural schematic diagram of a processor provided according to an embodiment of the present application.

[0178] As shown in FIG. 7, the processor a source register 72 for storing a data element; Destination Register Several vector registers, including 74, a decode unit 71 for decoding a vector shuffle instruction including register identifiers, including source register identifiers and destination register identifiers, and shuffle parameters; and an execution unit 73 responsive to the vector shuffle instruction to perform a vector shuffle operation on source elements obtained from the source register 72 in accordance with the shuffle parameters, obtain target elements after the vector shuffle operation, and write the target elements to the destination register 74.

[0179] Optionally, the instructions are stored in an instruction memory 70 .

[0180] Optionally, the execution unit 73 may allocate the source registers of the source elements according to the shuffle parameters. 72 determine position information and a number of source elements in the source register, the number of selected source elements being one or more, select source elements from the source register according to the determined position information and number of source elements, and determine all the selected source elements as target elements.

[0181] Optionally, the shuffle parameters include an index value and an opcode, the index value being for indicating location information in the source register of each source element required for the vector shuffle operation, and the opcode being for indicating an operation on the source register and the destination register; The execution unit 73 determines a selection rule for obtaining a source element according to the index value and the opcode, and selects a source register. 72 Then, in accordance with the selection rules, the source elements indicated by the respective index values ​​are obtained from the

[0182] Alternatively, the execution unit 73 determines a grouping method for the source elements according to the number of index values ​​when the number of index values ​​is different from the number of source elements, and determines the selection rule according to the grouping method and the opcode, and determines the selection rule according to the opcode when the number of index values ​​is the same as the number of source elements.

[0183] Alternatively, the execution unit 73 may organize adjacent elements in the source register into element groups, one for each N1 element group, where the data type of the elements is one of byte, halfword, or word, where N1 is a positive integer greater than 0, determine an element in each element group as an initial source element, and obtain source elements indicated by respective index values ​​from the initial source elements, where the number of source elements selected from each element group is n1.

[0184] Optionally, the adjacent elements are elements that are sequentially adjacent in position in the source register, and element addresses in adjacent groups of elements may be partially the same or completely different; The data type of the elements included in each element group is the same, and the data types of the elements included in different element groups may be the same or different.

[0185] Alternatively, the opcode is a second opcode, and the number of index values ​​and the number of source elements are the same; The execution unit 73 performs M operations for every N2 bits in the source register. N2 The data type of each of the elements is double word, and M each of N2 bits is N2 The number of source elements selected from the elements is n2, and N2, M N2 , and n2 are all positive integers greater than 0.

[0186] Optionally, the execution unit 73 creates an intermediate vector, the intermediate vector including at least one intermediate vector parameter, where if element groups are present, the number of the intermediate vector parameters is equal to the number of the element groups, but where if element groups are not present, the number of the intermediate vector parameters is equal to the number of the source elements, stores each of the selected source elements in a corresponding intermediate vector parameter of the intermediate vector, the intermediate vector parameters having a one-to-one correspondence with the selected source elements, and writes the content of each of the intermediate vector parameters to a corresponding location of the destination register according to the shuffle parameters.

[0187] Optionally, the opcode is a third opcode, the index values ​​include a first index value, a second index value, a third index value, and a fourth index value, the first index value, the second index value, the third index value, and the fourth index value each index a different position, and the source registers include a first source register and a second source register; The execution unit 73 performs M operations for each N3 bits in the source register 72. N3The source elements indicated by the first index value and the second index value are obtained from the elements, and M N3 From the elements, obtain source elements indicated by a third index value and a fourth index value, respectively, the data type of the elements is word, and M of N3 bits are N3 The number of source elements selected from the elements is n3, and N3, M N3 , and n3 are all positive integers greater than 0, and the first index value To The source element indicated by the second index value is determined as the first target element. To determining the source element indicated by the third index value as a second target element; To The source element indicated by the fourth index value is determined as the third target element. To The source element indicated by the thus specified is determined as a fourth target element, and the first target element and the second target element are written to a first location of the destination register, and the third target element and the fourth target element are written to a second location of the destination register.

[0188] Optionally, the opcode is a fourth opcode; The execution unit 73 stores, in the source register, N4 The data type of the elements is double word, the number of selected source elements is n4, and the M N4 and n4 are both positive integers greater than 0.

[0189] Alternatively, the opcode is a fifth opcode, the index values ​​include a first index value and a third index value, the first index value and the third index value each index a different location, and the source registers include a first source register and a second source register; The execution unit 73 stores in the first source register M N5 1. Obtain a first source element from the elements indicated by a first index value, and store M N5 obtain a second source element indicated by a third index value from the elements, the data type of the element being quad-word, the number of selected source elements being n5, n5 being a positive integer greater than 0, and determining the first source element and the second source element as target elements, respectively, to write them into corresponding positions of the destination register.

[0190] Alternatively, the number of the source registers is one or more, and the number of the destination registers is one; When the number of the source registers is one, the source register identifier and the destination register identifier may be the same or different; When the number of the source registers is multiple, each source register identifier of all the source registers is different from the destination register identifier, or when the number of the source registers is multiple, each source register has one source register identifier that is the same as the destination register identifier.

[0191] The processor provided by the embodiments of the present application can add a register identifier and a shuffle parameter to an instruction, and use the shuffle parameter to perform a vector shuffle operation on elements obtained from a source register. Therefore, there is no need to perform shuffle operations with multiple instructions to realize a specific function, and the vector shuffle operation of the specific function can be realized with a single instruction, thereby improving the execution efficiency of the specific function.

[0192] Example 8

[0193] As shown in FIG. 8, the electronic device may include one or more assemblies including a processing assembly 802, a memory 804, a power supply assembly 806, a multimedia assembly 808, an audio assembly 810, an input / output (I / O) interface 812, a sensor assembly 814, and a communication assembly 816.

[0194] The processing assembly 802 generally controls the overall operation of the electronic device, including operations related to display, data communication, camera operation, and recording operations. Processing Assembly The processing assembly 802 may include one or more processors 820 for processing instructions to complete all or some of the steps of the methods described above. The processing assembly 802 may also include one or more modules for facilitating interaction between the processing assembly 802 and other assemblies. For example, Processing Assembly 802 may include a multimedia module to facilitate interaction between the multimedia assembly 808 and the processing assembly 802 .

[0195] The memory 804 is configured to store various types of data to support operation on the electronic device. Examples of such data include instructions for any application or method operated on the electronic device, contact data, phone book data, messages, images, videos, etc. The memory 804 may be implemented by any type of volatile or non-volatile storage device, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disk, or a combination thereof.

[0196] The power supply assembly 806 provides power to the various assemblies of the electronic device. The power supply assembly 806 may include a power management system, one or more power sources, and other assemblies associated with the generation, management, and distribution of electrical power to the terminal 800.

[0197] The multimedia assembly 808 includes a screen that provides an output interface between the electronic device and a user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen for receiving input signals from a user. The touch panel includes one or more touch sensors for sensing touches, swipes, and gestures on the touch panel. The touch sensors can sense the boundaries of a touch or swipe operation as well as detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia assembly 808 includes one front camera and / or a rear camera. The front camera and / or the rear camera can receive external multimedia data when the electronic device is in an operating mode, such as a photo mode or a video mode. Each front camera and rear camera may be a fixed optical lens system or may have a focal length and optical zoom capability.

[0198] The audio assembly 810 is configured to output and / or input audio signals. For example, the audio assembly 810 includes a microphone (MIC) configured to receive an external audio signal when the terminal is in an operation mode such as a call mode, a recording mode, a voice recognition mode, etc. The received audio signal may be further stored in the memory 804 or transmitted by the communication assembly 816. In some embodiments, the audio assembly 810 further includes a speaker for outputting the audio signal.

[0199] The I / O interface 812 provides an interface between the processing assembly 802 and a peripheral interface module, which may be a keypad, a click wheel, buttons, etc. These buttons may include, but are not limited to, a home button, volume buttons, a start button, and a lock button.

[0200] The sensor assembly 814 includes one or more sensors for providing an assessment of various aspects of the state of the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the electronic device 800, the relative positioning of an assembly, such as the display and keypad of a terminal, and can also detect changes in the position of the terminal or an assembly of the terminal, the presence or absence of a user's contact with the electronic device, the orientation or acceleration / deceleration of the electronic device, and changes in the temperature of the electronic device. The sensor assembly 814 can include a proximity sensor configured to detect the presence of a nearby object without any physical contact. The sensor assembly 814 can also include an optical sensor, such as a CMOS or CCD image sensor for use in imaging applications. In some embodiments, the sensor assembly 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0201] The communication assembly 816 is configured to facilitate wired or wireless communication between the electronic device and other devices. The electronic device can access a wireless network based on a communication standard such as WiFi, 2G / 3G / 4G / 5G, or a combination thereof. In one exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component 816 further comprises a Near Field Communication (NFC) module for facilitating near-field communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0202] In an exemplary embodiment, the electronics may be implemented with one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic elements to perform the vector shuffling methods described above.

[0203] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is further provided, such as a memory 804 including instructions, which are executable by a processor 820 of the electronic device to achieve the above vector shuffling method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0204] The electronic device according to the embodiments of the present application is for implementing the vector shuffling methods corresponding to the embodiments of the methods described above, and has beneficial effects of implementing the corresponding methods, so they will not be described again here.

[0205] Each embodiment in this specification will be described step by step, and each embodiment will focus on the differences with other embodiments, and the same or similar parts between each embodiment can be referred to each other. The device embodiment is basically similar to the method embodiment, so the description will be relatively simple, and the relevant parts can be described with reference to the method embodiment.

[0206] The above is a detailed description of the vector shuffling method, processor and electronic device provided by the present application. In this specification, the principle and embodiment of the present invention are described by applying specific examples, but the above examples are merely described to help understand the method of the present invention and its main concept, and at the same time, those skilled in the art can make changes to the specific embodiment and its application scope based on the concept of the present invention. As such, the description in this specification should not be understood as limiting the present invention.

[0207] The algorithms and displays provided herein are not inherently related to any particular computer, electronic system, or other device. Various general-purpose systems may also be used in conjunction with the teachings herein. The necessary structure for building such a system will be apparent in light of the above description. Furthermore, this application is not directed to any particular programming language. It should be understood that a variety of programming languages ​​can be used to implement the contents of the application as described herein, and the above description of a particular language is provided to disclose the best mode of carrying out the application.

[0208] In the specification provided by the present application, a large amount of specific details are described. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0209] Similarly, in the above description of exemplary embodiments of the present application, it should be understood that various features of the present application may be grouped together in a single embodiment, drawing, or description thereof in order to simplify the disclosure and aid in the understanding of one or more aspects of the various inventions. However, this method of disclosure should not be interpreted as reflecting an intention that the present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, aspects of the present invention include fewer than all features of each of the embodiments disclosed above. Thus, any claims according to a particular embodiment are hereby expressly incorporated into that particular embodiment, with each claim serving as a separate embodiment of the present application in its own right.

[0210] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and arranged in one or more devices different from the embodiments. The modules, units or assemblies in the embodiments can be combined into a single module, unit or assembly, and can be further divided into multiple sub-modules, sub-units or sub-assemblies. All features disclosed in this specification (including the appended claims, abstract and accompanying drawings) and all processes or units of the methods or devices thus disclosed can be combined in any combination, except where at least some of such features and / or processes or units are mutually exclusive. Each feature disclosed in this specification (including the appended claims, abstract and accompanying drawings) may be replaced by an alternative feature serving the same, equivalent or similar purpose, unless otherwise specified.

[0211] In addition, those skilled in the art will appreciate that some embodiments described herein include some features but not other features included in other embodiments, but that combinations of features from different embodiments are within the scope of the present application and are meant to form different embodiments. For example, in the following claims, any one of the claimed embodiments may be used in any combination.

[0212] The embodiments of the various components of the present application may be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or digital signal processor (DSP) may actually be used to implement some or all of the functions of some or all of the components in a browser client device according to the embodiments of the present application. The present application may also be implemented as a device or apparatus program (e.g., computer program and computer program product) for performing some or all of the methods described herein. Such programs implementing the present application may be stored on a computer readable medium or may have the form of one or more signals. Such signals may be downloadable from an Internet site, provided on a carrier signal, or provided in any other form.

[0213] It should be noted that the above examples are illustrative of the present application, but do not limit it, and that a person skilled in the art may design alternative examples without departing from the scope of the appended claims. In the claims, reference signs placed between parentheses shall not be interpreted as limiting the scope of the claims. The term "comprises" does not exclude the presence of an element or step not stated in the claim. The term "a" or "one" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by hardware comprising several different elements, and by a suitably programmed computer. In individual claims enumerating several devices, several of these devices may be embodied by the same hardware. The use of the terms first, second, third, etc. does not indicate any order. These terms may be interpreted as names.

[0214] This application claims priority to a Chinese patent application filed with the China National Intellectual Property Office on December 10, 2021, bearing application number 202111508098.8 and entitled "Vector shuffling method, processor and electronic device," the entire contents of which are incorporated by reference into this application.

Claims

1. receiving an instruction comprising register identifiers and shuffle parameters, the register identifiers comprising source register identifiers and destination register identifiers, the source register identifiers for indicating source registers, the source registers being registers for storing source elements to be operated on when performing a vector shuffle operation, the destination register identifiers for indicating destination registers, the destination registers being registers for storing target elements resulting after performing the vector shuffle operation, and the shuffle parameters for indicating parameters according to which the vector shuffle operation is performed on the source elements; executing the instruction to determine position information of the source elements in the source registers and a number of source elements required for a vector shuffle operation according to the shuffle parameters, the selected number of source elements being one or more, the shuffle parameters including an index value and an opcode, the index value being for indicating position information of each source element in the source registers required for the vector shuffle operation, and the opcode being for indicating an operation on the source registers and destination registers; when the number of the index values ​​is different from the number of the source elements, determining a grouping method of the source elements according to the number of the index values, and determining a selection rule according to the grouping method and the operation code; obtaining, from a source register, a source element indicated by each index value according to the selection rule; determining all the selected source elements as target elements; writing the target element to the destination register.

2. The method of shuffling the vectors is as follows:

2. The method of claim 1, further comprising: if the number of index values ​​and the number of source elements are the same, determining a selection rule according to the opcode.

3. the opcode is a first opcode, and the number of index values ​​and the number of source elements are different; The step of obtaining, from the source register, source elements indicated by each index value according to the selection rule, arranging adjacent elements in the source register into element groups, one group for every N1 elements, where the elements have a data type of one of byte, halfword, or word, and N1 is a positive integer greater than 0; determining an element in each element group as an initial source element; and obtaining, from the initial source elements, source elements indicated by each index value, respectively, wherein the number of source elements selected from each of the element groups is n1.

4. The adjacent elements are elements that are sequentially adjacent in location in the source register, and element addresses in adjacent groups of elements may be partially the same or completely different; 4. The method of shuffling vectors according to claim 3, wherein the data type of elements included in each element group is the same, and the data types of elements included in different element groups are the same or different.

5. the opcode is a second opcode, and the number of index values ​​and the number of source elements are equal; The step of obtaining, from the source register, source elements indicated by each index value according to the selection rule, In the source register, M for every N2 bits N2 The method includes the step of obtaining a source element indicated by each index value from each of the elements, the data type of the elements being a double word, and M N2 The number of source elements selected from the elements is n2, and N2, M N2 3. The method of claim 2, wherein n1, n2, and n3 are positive integers greater than 0.

6. before the step of determining a grouping manner of the source elements according to the number of the index values; The method further includes the step of creating an intermediate vector, the intermediate vector including at least one intermediate vector parameter, the number of the intermediate vector parameters being equal to the number of the element groups if element groups are present, but the number of the intermediate vector parameters being equal to the number of the source elements if element groups are not present; The step of obtaining, from the source register, source elements indicated by each index value according to the selection rule, storing each selected source element in a corresponding intermediate vector parameter of the intermediate vector, the intermediate vector parameters having a one-to-one correspondence with the selected source elements; The step of writing the target element to the destination register comprises: A method for shuffling vectors according to any one of claims 3 to 5, further comprising the step of writing the content of each of said intermediate vector parameters to a corresponding location of said destination register according to said shuffle parameters.

7. the opcode is a third opcode; the index values ​​include a first index value, a second index value, a third index value, and a fourth index value, the first index value, the second index value, the third index value, and the fourth index value each indexing a different location; and the source registers include a first source register and a second source register; The step of obtaining, from the source register, source elements indicated by each index value according to the selection rule, In the first source register, M every N3 bits N3 obtaining, from the elements, source elements indicated by a first index value and a second index value, respectively; In the second source register, M every N3 bits N3 obtaining a source element indicated by a third index value and a fourth index value from the elements, respectively, the data type of the elements being a word, and M every N3 bits; N3 The number of source elements selected from the elements is n3, and N3, M N3 , and n3 are all positive integers greater than 0; The step of writing the target element to the destination register comprises: determining a source element indicated by the first index value as a first target element and determining a source element indicated by a second index value as a second target element; determining a source element indicated by the third index value as a third target element and determining a source element indicated by a fourth index value as a fourth target element; 3. The method of claim 2, further comprising the steps of: writing the first target element and the second target element to a first location of the destination register; and writing the third target element and the fourth target element to a second location of the destination register.

8. the opcode is a fourth opcode, The step of obtaining, from the source register, source elements indicated by each index value according to the selection rule, In the source register, M N4 the data type of the elements is doubleword, the number of selected source elements is n4, and the M N4 3. The method of shuffling vectors according to claim 2, wherein n1 and n2 are both positive integers greater than 0.

9. the opcode is a fifth opcode, the index values ​​include a first index value and a third index value, the first index value and the third index value each indexing a different location, and the source registers include a first source register and a second source register; The step of obtaining, from the source register, source elements indicated by each index value according to the selection rule, In the first source register, M N5 obtaining a first source element from the elements indicated by a first index value; N5 and obtaining a second source element from the elements indicated by a third index value, the data type of the element being quad-word, the number of selected source elements being n5, n5 being a positive integer greater than 0; The step of writing the target element to the destination register comprises:

3. The method of claim 2, further comprising the step of determining a first source element and a second source element as target elements and writing them to corresponding locations in the destination register.

10. the number of the source registers is one or more, the number of the destination registers is one, when the number of the source registers is one, the source register identifier is different from the destination register identifier; The method for shuffling vectors according to any one of claims 1 to 5 or 7 to 9, characterized in that, when the number of source registers is multiple, each source register identifier of all of the source registers is different from the destination register identifier, or, when the number of source registers is multiple, each source register has one source register identifier that is the same as the destination register identifier.

11. a plurality of vector registers including source and destination registers for storing data elements; a decode unit for decoding a vector shuffle instruction including register identifiers, including source register identifiers and destination register identifiers, and shuffle parameters; an execution unit responsive to the vector shuffle instruction to perform a vector shuffle operation on source elements obtained from the source register according to the shuffle parameters, obtain target elements after the vector shuffle operation, and write the target elements to the destination register; The execution unit determines position information of the source elements in the source register and a number of source elements according to the shuffle parameters, the number of the selected source elements being one or more, selects source elements from the source register according to the determined position information and the number of source elements, and determines all the selected source elements as target elements; The shuffle parameters include an index value and an opcode, the index value is for indicating location information in the source register of each source element required for the vector shuffle operation, and the opcode is for indicating an operation on the source register and the destination register; The execution unit determines a selection rule for acquiring a source element according to the index value and the opcode, and acquires, from a source register, each of the source elements indicated by the index value according to the selection rule; The processor, wherein when the number of the index values ​​and the number of the source elements are different, the execution unit determines a grouping method of the source elements according to the number of the index values, and determines the selection rule according to the grouping method and the opcode.

12. 12. The processor of claim 11, wherein the execution unit determines the selection rule according to the opcode when a number of the index values ​​and a number of the source elements are the same.

13. the opcode is a first opcode, and the number of index values ​​and the number of source elements are different; 13. The processor of claim 12, wherein the execution unit organizes adjacent elements in the source register into element groups, one group for every N1 elements, where the data type of the elements is one of byte, halfword, or word, where N1 is a positive integer greater than 0, determines an element in each element group as an initial source element, and obtains from the initial source elements source elements indicated by respective index values, where the number of source elements selected from each element group is n1.

14. The adjacent elements are elements that are sequentially adjacent in location in the source register, and element addresses in adjacent groups of elements may be partially the same or completely different; 14. The processor of claim 13, wherein the data type of elements included in each element group is the same, and the data types of elements included in different element groups are the same or different.

15. the opcode is a second opcode, and the number of index values ​​and the number of source elements are equal; The execution unit performs M operations for each N2 bits in the source register. N2 The data type of each of the elements is a double word, and M each of the N2 bits is N2 The number of source elements selected from the elements is n2, and N2, M N2 13. The processor of claim 12, wherein n1, n2, and n3 are all positive integers greater than 0.

16. 16. The processor of claim 13, wherein the execution unit creates an intermediate vector, the intermediate vector including at least one intermediate vector parameter, a number of the intermediate vector parameters being equal to a number of the element groups if element groups are present, but a number of the intermediate vector parameters being equal to a number of the source elements if element groups are not present, stores each selected one of the source elements in a corresponding intermediate vector parameter of the intermediate vector, the intermediate vector parameters having a one-to-one correspondence with the selected source elements, and writes a content of each of the intermediate vector parameters to a corresponding location of the destination register according to the shuffle parameters.

17. the opcode is a third opcode; the index values ​​include a first index value, a second index value, a third index value, and a fourth index value, the first index value, the second index value, the third index value, and the fourth index value each indexing a different location; and the source registers include a first source register and a second source register; The execution unit performs M operations every N3 bits in the source register. N3 The source elements indicated by the first index value and the second index value are obtained from the elements, and M N3 From the elements, obtain source elements indicated by the third index value and the fourth index value, respectively, the data type of the elements is word, and M every N3 bits are N3 The number of source elements selected from the elements is n3, and N3, M N3 13. The processor of claim 12, wherein n1, n2, and n3 are all positive integers greater than 0, and the execution unit determines a source element indicated by the first index value as a first target element, determines a source element indicated by the second index value as a second target element, determines a source element indicated by the third index value as a third target element, determines a source element indicated by the fourth index value as a fourth target element, writes the first target element and the second target element to a first location of the destination register, and writes the third target element and the fourth target element to a second location of the destination register.

18. the opcode is a fourth opcode, The execution unit stores, in the source register, N4 The data type of the elements is double word, the number of selected source elements is n4, and the M N4 13. The processor of claim 12, wherein n1 and n2 are both positive integers greater than 0.

19. the opcode is a fifth opcode, the index values ​​include a first index value and a third index value, the first index value and the third index value each indexing a different location, and the source registers include a first source register and a second source register; The execution unit stores in the first source register: N5 A first source element indicated by a first index value is obtained from the elements, and in the first source register, N5 13. The processor of claim 12, further comprising: obtaining a second source element from the elements indicated by a third index value, a data type of the element being quad-word, a number of selected source elements being n5, n5 being a positive integer greater than 0, and determining the first source element and the second source element as target elements and writing them to corresponding locations of the destination register.

20. the number of the source registers is one or more, the number of the destination registers is one, when the number of the source registers is one, the source register identifier is different from the destination register identifier; A processor according to any one of claims 11 to 15 or 17 to 19, characterized in that, when the number of the source registers is multiple, each source register identifier of all the source registers is different from the destination register identifier, or, when the number of the source registers is multiple, all the source registers have one source register identifier that is the same as the destination register identifier.

21. An electronic device including a memory and one or more programs, the one or more programs being stored in the memory, and the electronic device being configured to execute the vector shuffling method according to any one of claims 1 to 5 or 7 to 9 by one or more processors.

Citation Information

Patent Citations

  • Instruction for performing overload check

    JP2014182825A

  • Bit shuffle processor, method, system and instructions

    JP2017529601A