Processor, method for data processing, device, and storage medium

The processor architecture addresses the inefficiencies of conventional instruction sets for vector calculations by employing a simple instruction set and inter-memory operations, significantly improving the efficiency of vector operations, particularly in neural network computations.

JP2025519635AActive Publication Date: 2025-06-26BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024573128
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-14
Filing Date
2023-06-06
Publication Date
2025-06-26
Estimated Expiration
2043-06-06

AI Technical Summary

Technical Problem

Conventional instruction set architectures are not well-suited for large-scale vector calculations, such as those encountered in neural network computing, due to their complexity and redundancy, which hinders efficient execution of vector operations.

Method used

A processor architecture with an instruction decoder and arithmetic logic unit is designed to decode and execute target instructions for vector operations, utilizing a simple instruction set that is compatible with inter-memory operations, thereby enhancing the processor's ability to perform vector calculations efficiently.

Benefits of technology

This solution simplifies the processor's operation and improves its efficiency in executing vector calculations by using a simple instruction set, which is particularly beneficial for tasks like neural network operations, leading to enhanced computing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025519635000001_ABST
    Figure 2025519635000001_ABST
Patent Text Reader

Abstract

According to an embodiment of the present invention, a processor, a method for data processing, a device, and a storage medium are provided. The processor includes an instruction decoder configured to decode a target instruction for vector operations. The target instruction includes a target opcode, a source operand, and a target operand. The target opcode indicates a vector operation specified by the target instruction. The source operand specifies a source storage location for reading data to be processed in a memory. The target operand specifies a target storage location for writing a processing result in the memory. The processor further includes an arithmetic logic unit coupled to the instruction decoder and the memory. The arithmetic logic unit is configured to read the data to be processed from the source storage location of the memory, execute an arithmetic logic operation associated with the vector operation specified by the target instruction on the data to be processed, and write the processing result to the target storage location of the memory. Thereby, the efficiency of vector calculation can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application claims the priority of a Chinese patent application for invention with the title "Processor, Method for Data Processing, Device, and Storage Medium" and application number 202210674857.6, filed on June 14, 2022, the entire content of which is incorporated herein by reference.

[0002] Exemplary embodiments of the present invention generally relate to the field of computers, and in particular, to processors, methods for data processing, devices, and computer - readable storage media.

Background Art

[0003] With the development of information technology, various processors have been applied to various scenarios. For various application scenarios, different instruction - set architectures (ISAs) that can be adopted by processors have been proposed currently. These instruction - set architectures often need to be compatible with various usage scenarios. In some vector calculations with a large number of instruction repetitions and a large amount of data, a better instruction - set architecture is needed to enable the processor to process such vector calculations more appropriately.

Summary of the Invention

[0004] In a first aspect of the present invention, a processor is provided. The processor includes an instruction decoder configured to decode a target instruction for vector operations. The target instruction includes a target opcode, a source operand, and a target operand. The target opcode indicates a vector operation specified by the target instruction. The source operand specifies a source storage location for reading data to be processed in memory. The target operand specifies a target storage location for writing a processing result in memory. The processor further includes an arithmetic logic unit coupled to the instruction decoder and the memory. The arithmetic logic unit is configured to read data to be processed from the source storage location in the memory, perform an arithmetic logic operation associated with the vector operation specified by the target instruction on the data to be processed, and write the processing result to the target storage location in the memory.

[0005] In a second aspect of the present invention, a method for data processing is provided. The method includes decoding a target instruction for vector operations. The target instruction includes a target opcode, a source operand, and a target operand. The target opcode indicates a vector operation specified by the target instruction. The source operand specifies a source storage location for reading data to be processed in memory. The target operand specifies a target storage location for writing a processing result in memory. The method further includes reading data to be processed from the source storage location in the memory, performing an arithmetic logic operation associated with the vector operation specified by the target instruction on the data to be processed, and writing the processing result to the target storage location in the memory.

[0006] In a third aspect of the present invention, an electronic device is provided. The electronic device includes at least the processor of the first aspect.

[0007] In a fourth aspect of the present invention, a computer-readable storage medium is provided. A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method of the second aspect is implemented.

[0008] It should be understood that the content described in the summary part of the present invention is not intended to limit the main features or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will be easily understood from the following description.

Brief Description of the Drawings

[0009] With reference to the following detailed description in conjunction with the drawings, the above-described features and other features, advantages, and aspects of each embodiment of the present invention will become more apparent. In the drawings, the same or similar symbols indicate the same or similar elements, where

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Modes for Carrying Out the Invention

[0010] Hereinafter, embodiments of the present invention will be described in more detail with reference to the drawings. Although specific embodiments of the present invention are shown in the drawings, the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Rather, these embodiments are provided for a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes only and are not used to limit the protection scope of the present invention.

[0011] In the description of the embodiments of the present invention, the term "including" and its similar terms are open-ended inclusion meaning "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may be other explicit and implicit definitions below.

[0012] It is understood that the data related to the present technical solution (including but not limited to the data itself, data acquisition, or data use) should comply with the corresponding laws and regulations and related specified requirements.

[0013] As described above, with the development of information technology, various processors are applied to various scenarios. For various application scenarios, different instruction set architectures that can be adopted by processors have been proposed currently. These instruction set architectures often need to be compatible with various usage scenarios. However, the usage scenarios of these conventional instruction set architectures do not match the usage scenarios of vector calculations such as neural network computing. Therefore, for some vector calculations with a large number of instruction repetitions and a large data volume, a better instruction set architecture is required so that the processor can process such vector calculations more appropriately.

[0014] The conventional solution is to adopt a standard processor instruction set such as the Reduced Instruction Set Computer (RISC)-V instruction set. These general-purpose instruction sets can execute various vector calculations such as various neural network operators, but since these general-purpose instruction sets need to be compatible with various usage scenarios, it is difficult to ensure higher execution efficiency. For example, the calculation of neural network operators usually includes large-scale vector calculations, so it is not suitable for general-purpose instruction sets.

[0015] According to research, it has been found that the instruction set architectures of conventional solutions are not suitable for certain large-scale vector calculations. For example, conventional solutions can use digital signal processor (DSP) architectures such as single instruction multiple data (SIMD) architectures, or vector processor architectures. However, the instruction sets of the above-mentioned DSP architectures are usually not publicly available. In contrast, in the case of vector processor architectures such as the vector instruction set based on the RISC-V standard (referred to as the RISC-V vector instruction set), these instruction sets are usually more complex and seem redundant for vector calculations such as neural network operators.

[0016] In summary, for some vector calculations with a large number of instruction repetitions and a large amount of data, in order to improve the computing efficiency of the processor, it is necessary to design an instruction set suitable for vector calculations.

[0017] According to an embodiment of the present invention, an improved solution for a processor is proposed. In this solution, the processor includes an instruction decoder and an arithmetic logic unit. The instruction encoder is used to receive a target instruction for processing vector operations. The target instruction is applicable to a processor architecture of memory-to-memory (MEM to MEM). For example, the target instruction includes a target opcode, a source operand, and a target operand. The target opcode indicates the vector operation specified by the target instruction, the source operand specifies at least a source storage location for reading the data to be processed in the memory, and the target operand specifies at least a target storage location for writing the processing result in the memory.

[0018] The arithmetic logic unit of the processor is coupled to an instruction decoder and a memory. The arithmetic logic unit is configured to perform a vector operation of a target instruction based on the decoding information of the instruction decoder for the target instruction. For example, the arithmetic logic unit is configured to read processing target data from a source storage location in the memory, perform an arithmetic logic operation associated with the vector operation specified by the target instruction on the processing target data, and write the processing result of the processing target data to a target storage location in the memory.

[0019] This solution simplifies the operation of the processor by using a processor suitable for an inter-memory architecture. Thereby, the processor can execute a large number of vector calculations using a simple instruction set. For example, the processor can perform vector calculations of a neural network using a simple instruction set. In this way, this solution can improve the efficiency of the processor executing vector calculations using a simple instruction set.

[0020] FIG. 1 shows a schematic diagram of an exemplary environment 100 in which an embodiment of the present invention can be implemented. In the environment 100, the processor 110 may represent any type of instruction processing device. For example, the processor 110 may be a general-purpose processor or any other suitable processor. The processor 110 is configured to receive an instruction 140 and execute an operation indicated by the instruction 140, such as a vector operation. For example, the processor 110 can receive the instruction 140 from other devices within the environment 100. In some embodiments, the instruction 140 is a SIMD instruction.

[0021] The processor 110 includes an instruction decoder 120 and an arithmetic logic unit 130. Alternatively or additionally, the processor 110 may further include a memory (not shown) or may be communicatively coupled to a memory. For example, the memory may be a data memory (Vector Closely-coupled It may also be (such as Memory, VCCM, etc.). The instruction decoder 120, the arithmetic logic unit 130, and the memory are communicatively coupled. That is, the instruction decoder 120, the arithmetic logic unit 130, and the memory can communicate with each other according to appropriate data transmission protocols and / or standards. During operation, the instruction decoder 120 receives the instruction 140 and decodes the instruction 140. For example, the instruction decoder 120 decodes the instruction 140 into arithmetic operations and / or logical operations that can be processed by the arithmetic logic unit 130. The instruction decoder 120 can be implemented using various different mechanisms. For example, the instruction decoder 120 may be implemented using a hardware circuit or may be at least partially implemented by a software module.

[0022] The arithmetic logic unit 130 is configured to operate based on the information obtained by decoding the instruction 140 by the instruction decoder 120. The arithmetic logic unit 130 can perform various arithmetic operations, logical operations, etc. The arithmetic logic unit 130 can be implemented using various different mechanisms. For example, the arithmetic logic unit 130 may be implemented using a hardware circuit or may be at least partially implemented by a software module.

[0023] It should be understood that the configuration and functions of the environment 100 are described for illustrative purposes only and do not imply any limitation to the scope of the present invention. For example, the processor 110 can be applied in various existing or future computing platforms or computing systems. The processor 110 can be implemented in various embedded applications (such as data processing systems such as mobile network base stations) to provide services such as large-scale vector calculations. The processor 110 may be integrated or embedded in various electronic devices or computing devices to provide various computing services. The application environment and application scenarios of the processor 110 are not limited here.

[0024] In some embodiments, instruction decoder 120 decodes the received instruction 140. In this specification, instruction 140 may also be referred to as a "target instruction", and in this context, the two may be used interchangeably. FIG. 2 shows a schematic diagram of an exemplary instruction 140 according to some embodiments of the present invention. As shown in FIG. 2, instruction 140 includes a target opcode 210, a source operand 220, and a target operand 230. In this specification, target opcode 210 may also be referred to as an "opcode", and in this context, the two may be used interchangeably. Target opcode 210 can indicate the vector operation specified by instruction 140. Source operand 220 is used to specify at least a source storage location for reading the data to be processed in memory. Target operand 230 is used to specify at least a target storage location for writing the processing result in memory.

[0025] In some embodiments, instruction decoder 120 decodes the above-described information indicated by target opcode 210, source operand 220, and target operand 230 so that arithmetic logic unit 130 processes it. For example, arithmetic logic unit 130 is configured to read the data to be processed from the source storage location in memory specified by source operand 220. Arithmetic logic unit 130 executes an arithmetic logic operation associated with the vector operation specified by instruction 140 on the data to be processed. Arithmetic logic unit 130 further writes the processing result of the data to be processed to the target storage location specified by target operand 230.

[0026] In some embodiments, instruction 140 may be encoded using binary or the like. In other embodiments, instruction 140 may be encoded using other encoding formats or other bases. In this specification, unless otherwise specified, the encoding formats and encoding representations of instruction 140 described below are all binary, for example. For example, instruction 140 in binary format may be defined using the format of Table 1 below. Table 1 Instruction Format Definition

[0027] [Table 1] As shown in Table 1, bits 86 to 95 are used to indicate the target opcode 210 of instruction 140. Each operand from bits 22 to 85 is used to indicate the source operand 220 of instruction 140. Each parameter from bits 0 to 21 is used to indicate the target operand 230 of instruction 140. Of course, it should be understood that any specific numerical values or bit numbers appearing here and elsewhere in this specification are exemplary unless otherwise specified. For example, the number of bits where each opcode and / or operand listed above is located is exemplary, not limiting. The target opcode 210, source operand 220, and target operand 230 of instruction 140 may be located at other appropriate numbers of bits.

[0028] For example, the source operand A_vaddr from bit 70 to bit 85 is used to indicate the address index of the data in the A channel (also called the first storage space of the memory) in a memory such as the data memory (VCCM, etc.), that is, the address index of VCCM[A_vaddr]. The address index is in units of one vector word. A vector word can indicate a storage unit of a channel in the memory, and the width of the storage unit is the width of the SIMD. That is, the address index is in units of one SIMD width. In some embodiments, the depth of the memory is, for example, 1024. In this example, only 10 bits within bits 70 to 85 can be used to indicate the address index of the A channel. Of course, it should be understood that the memory may have other appropriate depths, and the address index may have other appropriate numbers of bits.

[0029] Also, for example, the source operand A_index from bit 60 to bit 69 is used to indicate the element index of the vector word in the A channel in the memory. Each vector word may have, for example, 64 elements. The element index can be used to indicate a specific element among the vector words in the A channel. In some embodiments, when the vector word in the A channel is divided into, for example, 64 elements, only 6 bits within bits 60 to 69 can be used to indicate A_index. Further, for example, the source operand A_vm from bit 54 to bit 59 is used to indicate the index of the vector mask (VM) register of the A channel. In some embodiments, the A channel has 16 VM registers. In such an example, only 4 bits within bits 54 to 59 can be used to indicate A_vm.

[0030] Similarly, the source operand B_vaddr from bit 38 to bit 53 is used to indicate the address index of the B channel (also called the second storage space of the memory) in the memory (such as the data memory VCCM), that is, the address index of VCCM[B_vaddr]. The address index is in units of one vector word. That is, the address index is in units of one SIMD width. The source operand B_index from bit 28 to bit 37 is used to indicate the element index of the vector word in the B channel. The source operand B_vm from bit 22 to bit 27 is used to indicate the index of the vector mask register in the B channel.

[0031] The example of the target operand 230 in Table 1 can include C_vaddr from bit 6 to bit 21 and indicate the address index of the C channel in the data memory VCCM, that is, the address index of VCCM[C_vaddr]. The address index is in units of one vector word. That is, the address index is in units of one SIMD width. The example of the target operand 230 further includes C_vm from bit 0 to bit 5 and indicates the index of the vector mask register in the C channel.

[0032] Figure 3 shows a schematic diagram of memory locations corresponding to exemplary source operands according to some embodiments of the present invention. In the example of Figure 3, the memory storage space is divided into a plurality of channels such as channel 310-1, channel 310-2, …, channel 310-N, where N is an integer greater than 1. For ease of explanation, hereinafter, channel 310-1, channel 310-2, …, channel 310-N will be collectively or individually referred to as channel 310. In some embodiments, the value of N may be preset. For example, N may be set to different values such as 1024, 512, etc. Each channel 310 includes, for example, 1024 bits or other appropriate number of bits. The address index 330 (such as source operand A_vaddr, B_vaddr, or target operand C_vaddr) can indicate the address of channel 310-1. The address of channel 310 may be in units of vector words. The address index 330 may be 16 bits. For example, if the address index 330 is "0b0000_0000_0000_0000", the address index 330 can indicate channel 310-1. Also, for example, in some embodiments, in an example where the depth of the memory or the number of channels is 1024, the address index 330 may be 10 bits. For example, channel 310-1 is indicated by the address "0b00_0000_0000". It should be noted that in this specification, all encoded expressions starting with "0b" represent binary expressions and will not be repeatedly explained hereinafter. The "_" appearing in the binary expression is only for the convenience of display and has no actual meaning and does not occupy binary bits.

[0033] In some embodiments, the vector word of each channel 310 may be divided into a plurality of elements such as element 320. Element 320 may include, for example, 64 bits. An element index 340 (such as source operand A_index or B_index) can indicate a specific element, such as the index of element 320. The element index 340 may be 10 bits. For example, if the element index 340 is "0b00_0000_0000", the element index 340 can indicate element 320. Also, for example, in an example where the number of elements in the vector word of each channel 310 is 64, the element index can be 6 bits. For example, the element index "0b00_0000" can indicate element 320.

[0034] It should be understood that any specific numerical values, number of bits, and binary representations appearing herein and in other places in this specification are exemplary unless otherwise specified. For example, in other embodiments, each channel may have a different number of bits, and each vector word may use a different number of bits. Accordingly, the address index and the element index may have different numbers of bits and different encoding representations. The scope of the present invention is not limited in this regard.

[0035] Some examples of the source operand 220 were given above with reference to Table 1. Below, more examples of the source operand 220 will be described with reference to Table 2. Table 2 Instruction Format Definition

[0036] [Table 2] As shown in Table 2, the source operand 220 may include A_imm located from bit 54 to bit 85, indicating one immediate value of the instruction 140. Similarly, the source operand 220 may further include B_imm located from bit 22 to bit 53, indicating another immediate value of the instruction 140. Similar to Table 1, the target operand 230 in Table 2 may include C_vaddr and / or C_vm.

[0037] As described above, it should be understood that each source operand and / or each target operand described in conjunction with Tables 1 and 2 are merely exemplary and not limiting. The source operand 220 and / or the target operand 230 employed in the present invention may include any one of the above or various source operands and / or target operands. In some embodiments, the source operand 220 and / or the target operand 230 may include any other suitable type of operand different from the above source operand and / or target operand.

[0038] Table 3 below illustrates an exemplary encoding method for the opcode of the instruction 140. For example, if bit 0 is 0, it indicates that the instruction 140 is of the variable type. If bit 0 is 1, it indicates that the instruction 140 is of the immediate value type. Bits 1 to 2 indicate the sub-function code of the instruction 140. Bits 3 to 7 indicate the function code of the instruction 140. Bits 8 to 9 indicate the calculation precision of the instruction 140. For example, the binary "00" can indicate that the calculation precision is a single-precision floating-point number. Other binary values can indicate other reserved calculation precisions. Table 3 Exemplary Opcode Encoding Method

[0039] [Table 3] It should be understood that the opcode encoding method of instruction 140 shown in Table 3 is merely exemplary and not limiting. For example, in other embodiments, other encoding methods can be used to encode instruction 140.

[0040] In some embodiments, the vector operation specified by instruction 140 can be determined according to the target opcode 210 of instruction 140. For example, the processor 110 can pre-store the opcode of each instruction. The instruction decoder 120 can determine the vector operation specified by the received instruction 140 based on the target opcode 210 of the instruction 140. For example, when the target opcode 210 of instruction 140 is encoded as "0b00_00110_01_0", the instruction decoder 120 can determine that the instruction 140 is a v2indexr instruction. It should be understood that the opcode and instruction type examples listed above are merely exemplary and not limiting. The instruction encoded as "0b00_00110_01_0" may specify other vector operations.

[0041] Some exemplary execution methods of instruction 140 and execution 140 by the processor 110 will be described below. In some embodiments, the source operand 220 may include two source operands, such as A_vaddr and B_vaddr, A_vaddr and B_imm, etc. The width of each source operand may be the SIMD width. Alternatively or additionally, in some embodiments, the source operand 220 may include only one source operand, such as B_vaddr. The target operand 230, such as C_vaddr, can specify the target storage location where the processing result is written back to memory, that is, VCCM[C_vaddr].

[0042] In some embodiments, the target memory location of instruction 140 includes the processing result vector. The target operand 230 further indicates a target VM register such as C_vm or vm3. The value at each position of the target VM register indicates whether the corresponding processing result is to be written to the corresponding position of the processing result vector. For example, if the target register vm3[i] is 1, it indicates that the i-th element of the processing result vector word is writable and can be written to the corresponding processing result. Conversely, if the target register vm3[i] is 0, the i-th element of the processing result vector word cannot be written to the corresponding processing result.

[0043] Table 4 illustrates some exemplary instructions that can be supported by the processor 110. The instructions in Table 4 are described with reference to the instruction definitions in Table 1 or Table 2 and can be encoded with reference to the exemplary encoding method in Table 3. In the example of Table 4, the target operand 230 includes all of C_vaddr (i.e., &v3) and C_vm (i.e., vm3). The reserved bits in Table 4 indicate one or more reserved bits. These reserved bits are encoded or used later. Table 4 Exemplary Instructions

[0044] [Table 4] As an example, instruction 140 includes a first index determination instruction (e.g., v2indexl or v2indexr in Table 4). In this example, source operand 220 specifies a position in the first storage space of the memory, that is, the address index of Channel A (A_vaddr is &v1). Source operand 220 further specifies a predetermined index value of the data to be processed within the second storage space of the memory, that is, the element index within the vector word of Channel B (B_vaddr is &v2 and B_index is index2). In this example, arithmetic logic unit 130 is configured to determine a first index indicating a storage position in the first storage space of the value of the data to be processed at the position indicated by the predetermined index value.

[0045] For example, the opcode of instruction v2indexl can be encoded as "0b00_00110_00_0", and instruction v2indexl v1,v2,index2,v3,vm3 indicates that v3[i] is assigned as indext, where indext is the index of the leftmost element that can make v1[indext] equal to v2[index2]. If there is no element satisfying the above conditions, indext is set to "-1" indicated by the two's complement code. In some embodiments, target operand 230 further indicates a target vector mask register. The value at each position of the target vector mask register indicates whether the corresponding processing result is written to the corresponding position of the processing result vector. For example, if vm3[i] is equal to 1, v3[i] is writable.

[0046] Also, for example, the opcode of the instruction v2indexr can be encoded as "0b00_00110_01_0". The instruction v2indexr v1,v2,index2,v3,vm3 indicates assigning v3[i] as indext, where indext is the index of the rightmost element that can make v1[indext] equal to v2[index2]. If there is no element satisfying the above conditions, set indext to "-1" indicated by the two's complement code. In this example, if vm3[i] is equal to 1, v3[i] is writable.

[0047] As another example, the instruction 140 includes a second index determination instruction (e.g., v2indexli or v2indexri in Table 4). In this example, the source operand 220 specifies the position in the first storage space of the memory, i.e., the address index of channel A (A_vaddr is &v1). The source operand 220 further specifies the first immediate number, i.e., the immediate number imm2. In this example, the arithmetic logic unit 130 is configured to determine the second index. The second index indicates the storage position of the first immediate number within the first storage space.

[0048] For example, the opcode of the instruction v2indexli can be encoded as "0b00_00110_00_1". The instruction v2indexli v1,imm2,v3,vm3 indicates assigning v3[i] as indext, where indext is the index of the leftmost element that can make v1[indext] equal to imm2. If there is no element satisfying the above conditions, set indext to "-1" indicated by the two's complement code. In this example, if vm3[i] is equal to 1, v3[i] is writable.

[0049] Also, for example, the opcode of the instruction v2indexri can be encoded as "0b00_00110_01_1". The instruction v2indexri v1,imm2,v3,vm3 indicates that v3[i] is assigned as indext, where indext is the index of the leftmost element from the right that can make v1[indext] equal to imm2. If there is no element satisfying the above conditions, indext is set to "-1" indicated by the two's complement code. In this example, if vm3[i] is equal to 1, v3[i] is writable.

[0050] As another example, the instruction 140 may include a first numerical determination instruction such as the instruction sindex2v in Table 4. The source operand 220 specifies the position of the first storage space in the memory, that is, the address index of Channel A (A_vaddr is &v1). The source operand 220 further specifies a predetermined index value of the data to be processed in the second storage space in the memory, that is, the element index within the vector word of Channel B (B_vaddr is &v2 and B_index is index2).

[0051] In this example, the arithmetic logic unit 130 is configured to determine a predetermined value of the data to be processed at the position indicated by the predetermined index value and determine it as the first value at the position using the predetermined value in the first storage space as an index. For example, the instruction sindex2v v1,v2,index2,v3,vm3 has an opcode encoded as "0b00_00110_10_0". This instruction indicates that v3[i] is assigned as v1[v2[index2]]. If vm3[i] is equal to 1, v3[i] is writable.

[0052] In some embodiments, instruction 140 includes a second value determination instruction. In this example, source operand 220 specifies a predetermined index value of the data to be processed within the second storage space of the memory, that is, B_vaddr is &v2 and B_index is index2. The arithmetic logic unit 130 is configured to determine a second value at a position indicated by the predetermined index value of the data to be processed. For example, instruction s2v v2,index2,v3,vm3 has an opcode encoded as "0b00_00110_10_1". This instruction indicates that v3[i] is assigned as v2[index2]. If vm3[i] is equal to 1, v3[i] is writable.

[0053] By using one or more of the first index determination instruction, the second index determination instruction, the first value determination instruction, and the second value determination instruction described above, the processor 110 can more appropriately process some operators such as an operator for obtaining an index of a maximum value (ArgMax), an operator for obtaining an index of a minimum value (ArgMin), an operator for obtaining the top K values of the highest rank (TopK), etc., for obtaining coordinates. Taking ArgMax as an example, it is used to obtain an index at which the value v[index] of the vector v becomes the maximum value. The instructions required for ArgMax of 64 elements are as follows. That is, first, v2smax In the case of v1,vm1,v2,vm2 (this instruction is described in Tables 5 and 6 below), this instruction obtains the maximum element value of v1 and writes it to v2. Here, all bits of the stored values of vm1 and vm2 are 1. Next, v2indexl In the case of v1,v2,0,v3, vm3, this instruction makes the obtained index [index] equal to v2 for v1 and writes the value of index to v3. Here, all bits of the stored value of vm3 are 1.

[0054] In some embodiments, the target instruction includes a vector transpose instruction such as the vtranspose or vstranspose instruction. The source operand 220 specifies a first position within a first storage space of the memory, i.e., A_vaddr is &v1 and A_index is index1. The source operand 220 further specifies a source vector mask register vm1 and an optional vm2. In this example, the arithmetic logic unit 130 is configured to vector transpose the data to be processed at the first position within the first storage space to obtain the transposed data to be processed.

[0055] For example, the vector transpose instruction vtranspose v1, index1, vm1, vm2, v3, vm3 has an opcode encoded as 0b00_00111_11_0 for transposing, for example, a 32*32 vector (or matrix). In this example, the values of vm1 and vm2 are read channel (lane) enabled, and the value of vm3 is write channel enabled. The number of consecutive 1 bits R in vm1 is used to indicate the number of rows of the matrix, and the number of consecutive 1 bits C in vm2 is used to indicate the number of columns of the matrix, where R and C are all arbitrary natural numbers, and R and C may be the same or different. The valid bits of vm1, vm2, and vm3 need to be consecutive. If not, the first 1 of the least significant bit is dominant. The above vector transpose instruction vtranspose can be used to transpose an R*C matrix.

[0056] Also, for example, in some embodiments, the vector transpose instruction vstranspose for transposing a matrix v1, index1, vm1, v3, vm3 can be used. In this example, the value of vm1 is read channel enabled, and the value of vm3 is write channel enabled. The number of consecutive 1 bits R in vm1 is used to indicate the number of rows (or columns) of the matrix. The vector transpose instruction vstranspose can be used to transpose an R*R matrix.

[0057] The vector transpose instruction is a standard RISC type instruction. In the conventional standard RISC type instruction set, the vector transpose function needs to be completed by a plurality of consecutive transpose instructions. In this solution, by using the vector transpose instruction, the computing ability of some networks can be improved. For example, the neural network training process usually includes a large number of matrix or matrix transpose operations. By utilizing the vector transpose instruction of this solution, the computing efficiency of the neural network training process can be improved.

[0058] In some embodiments, the target instruction includes an exponential instruction such as vexp. In this example, the source operand 220 specifies the source storage location, that is, A_vaddr is &v1. The arithmetic logic unit 130 is configured to determine an exponential value with a predetermined numerical value (for example, the natural base e) as the base using the data to be processed at the source storage location as the power. For example, vexp v1,v3,vm3 has an opcode encoded as "0b00_01000_01_0", and this instruction assigns v3[i] as exp(v1[i]). If vm3[i] is equal to 1, v3[i] is writable.

[0059] The above exponential instruction is applied to operators such as the sigmoid operator and hyperbolic functions sinh / cosh / tanh. For example, the sigmoid operator, sinh operator, cosh operator, and tanh operator can be represented by the following formulas (1) to (4).

[0060]

Number

[0061] For example, in neural network activation functions, sigmoid and hyperbolic functions are more common. By using the exponential instruction of this solution, the efficiency of such calculations can be improved.

[0062] In some embodiments, instruction 140 includes a VM register instruction such as a Vm2index instruction. Source operand 220 indicates a source VM register in memory, i.e., vm1. The arithmetic logic unit is configured to store the index at the enabled position of the source VM register in the target memory location. For example, the instruction vm2index vm1,v3,vm3 has an opcode encoded as "0b00_01100_01_0" which represents v3[i]=vm1[i]?i:-1. That is, if the value of vm1[i] is 1, v3[i] is assigned i, and conversely, if the value of vm1[i] is 0, v3[i] is assigned -1 (e.g., "-1" represented by two's complement). If vm3[i] is equal to 1, v3[i] is writable.

[0063] In some embodiments, instruction 140 includes a onehot code conversion instruction such as vindex2vm. This is a VM register operation instruction. In this example, source operand 220 specifies a predetermined index value of the data to be processed in the second memory space of the memory, i.e., B_vaddr is &v2 and B_index is index2. The instruction includes the act of reading the vector mask register rather than being write-enabled for memory writing (all other instructions for writing to memory need to read the vector mask register as write-enabled). Target operand 230 specifies a target VM register, i.e., vm3. The arithmetic logic unit is configured to convert the value of the data to be processed at the predetermined index value into a onehot code and store the onehot code in the target VM register. For example, the instruction vindex2vm v2, index2, and vm3 have an opcode encoded as "0b00_10000_01_0". This instruction indicates that vm3 is assigned as onehot(v2[index2]), where onehot() represents the one-hot encoding function.

[0064] The above one-hot encoding instruction is applicable to index-type instructions and can support the one-hot encoding operator to convert a numerical value into the one-hot encoding format. For example, in actual use, it can be implemented using two instructions: vindex2vm v1, 0, vm1 and vmload vm1, v2, vm2. Here, the first instruction is used to convert the numerical value of v1 into the one-hot encoding format and write it to vm1, and the second instruction (vmload is described in Tables 7 and 8 below) is used to store the numerical value of vm1 in v2, where all bits of the stored value of vm2 are 1.

[0065] As described above, in conjunction with Table 4, examples of various instructions 140 supported by the processor 110 of the present invention have been explained. It should be understood that the processor 110 of the present invention can also support more instructions. Table 5 below shows examples of more conventional instructions 140 supported by the processor 110. The instructions in Table 5 can be described with reference to the instruction definitions in Table 1 or Table 2, and the opcode can be encoded using the exemplary encoding method in Table 3. Table 5 Exemplary Conventional Instructions

[0066]

Table 5

[0067] [Table 6] TIFF2025519635000011.tif246148TIFF2025519635000012.tif251142TIFF2025519635000013.tif224154The MAX() function and MIN() function in Table 6 respectively represent the function of finding the maximum value and the function of finding the minimum value. DW indicates the width of the vector word, LANE_NUM indicates the number of elements in one vector word, the mod() function represents the remainder function, the ceil() function and floor() function respectively represent rounding up and rounding down, and the SUM() function represents the addition function.

[0068] In some embodiments, the instructions supported by the processor 110 further include various vector mask register access and operation instructions. Table 7 shows some examples of vector mask register access and operation instructions. The functions and definitions of each instruction in Table 7 are shown in Table 8. The functions of these instructions include reading, writing, and operating on vector mask registers. These instructions do not include write enable for memory writing and include the act of reading vector mask registers (all other instructions that write to memory read the vector mask register as needs write enable). These instructions will not be described in detail here. Table 7 Exemplary Vector Mask Register Instruction Codes

[0069]

Table 7

[0070]

Table 8

[0071] Table 9 below shows some examples of internal register access and operation instructions. Table 10 shows the functions of each instruction in Table 9. These instructions are not described in detail herein. Note that for the vwcsr instruction in Table 9, the source operand is in channel A, and for the vwcsri instruction, the immediate number is in channel B. Table 9 Exemplary Internal Register Instructions

[0072]

Table 9

[0073]

Table 10

[0074] Also, each instruction in Tables 4 to 10 has been listed with reference to the exemplary instruction definitions and exemplary opcode encoding representations defined in Tables 1 to 3 above, but it should be understood that this is merely exemplary and not limiting. The instruction set supported by the processor of the present invention can be defined and encoded in any suitable manner. For example, each bit of each instruction may have a meaning different from the meaning represented by each bit in Table 1 or Table 2. Also, for example, the encoded representation of the opcode of each instruction may have a different number of bits from Table 3, and each number of bits may have a meaning different from each bit in Table 3. The encoded representation of the opcode of each instruction in Tables 4 to 10 above can be changed or interchanged. Each instruction can also be represented by another name. The scope of the present invention is not limited in this regard.

[0075] The above instructions do not include branch-type instructions and do not include load / store-type instructions. Different from the SIMD processor of the conventional vector register file, the register adopted in the present invention is a SIMD processor architecture between memories. The above instruction set defines a plurality of (for example, 64 or more, or less) vector mask registers to indicate the specific vectors that each SIMD instruction needs to process.

[0076] This solution simplifies the operation of the processor 110 by using a SIMD processor suitable for an inter-memory architecture. As a result, the processor 110 can execute a large number of vector calculations using a simple instruction set. For example, the processor 110 can execute tasks such as vector calculations of neural network operators using a simple instruction set. In this way, this solution improves the efficiency of the processor in executing vector calculations by adopting a simple instruction set. In the case of calculations such as neural network training and / or inference, the solution of the present invention can significantly improve computing efficiency. For example, the processor according to an embodiment of the present invention can support various indexing instructions, thereby improving the efficiency of various vector calculations such as obtaining coordinates. Also, for example, the processor of the present invention can process vector transposition instructions and the like, thereby improving the computing efficiency of the corresponding calculations in the neural network training process. Also, for example, the processor of the present invention can support exponential instructions, thereby improving and optimizing the computing efficiency of sigmoid operators, hyperbolic function operators, and the like.

[0077] FIG. 4 shows a flowchart of a process 400 for data processing according to some embodiments of the present invention. The process 400 may be implemented in the processor 110. For ease of explanation, the process 400 will be described with reference to the environment 100 of FIG. 1.

[0078] In block 410, the processor 110 decodes a target instruction for vector operations, such as instruction 140. For example, the instruction decoder 120 of the processor 110 decodes instruction 140. Instruction 140 includes a target opcode 210, a source operand 220, and a target operand 230. The target opcode 210 indicates the vector operation specified by instruction 140. The source operand 220 specifies at least a source storage location for reading the data to be processed in the memory. The target operand 230 specifies at least a target storage location for writing the processing result in the memory.

[0079] In block 420, the processor 110 reads the data to be processed from the source storage location in the memory. For example, the arithmetic logic unit 130 of the processor 110 may read the data to be processed from the source storage location in the memory. In block 430, the processor 110 executes an arithmetic logic operation associated with the vector operation specified by the target instruction on the data to be processed. For example, the arithmetic logic unit 130 of the processor 110 may execute the above arithmetic logic operation. In block 440, the processor 110 writes the processing result of the data to be processed to the target storage location in the memory. For example, the arithmetic logic unit 130 of the processor 110 may write the processing result to the target storage location.

[0080] In some embodiments, instruction 140 includes an indexing instruction. The indexing instruction may be a first indexing instruction (v2indexl or v2indexr) or a second indexing instruction (v2indexli or v2indexri). Source operand 220 specifies a position in the first storage space of the memory, and source operand 220 further specifies a predetermined index value of the data to be processed within the second storage space of the memory or a first immediate number. In block 430, the arithmetic and logical operation executed by processor 110 includes determining a first index or determining a second index. The bit 1 index indicates the storage position in the first storage space of the value of the data to be processed at the position indicated by the predetermined index value. The bit 2 index indicates the storage position of the first immediate number in the first storage space.

[0081] In some embodiments, instruction 140 includes a first numerical determination instruction (such as instruction sindex2v), source operand 220 specifies a position in the first storage space of the memory, and source operand 220 further specifies a predetermined index value of the data to be processed within the second storage space of the memory. In block 430, the arithmetic and logical operation executed by processor 110 includes determining a predetermined value of the data to be processed at the position indicated by the predetermined index value and determining a first value as the value at the position in the first storage space indexed by the predetermined value.

[0082] In some embodiments, instruction 140 includes a second numerical determination instruction such as instruction s2v. Source operand 220 specifies a predetermined index value of the data to be processed within the second storage space of the memory. In block 430, the arithmetic and logical operation executed by processor 110 includes determining a second value at the position indicated by the predetermined index value of the data to be processed.

[0083] In some embodiments, instruction 140 includes a vector transpose instruction such as instruction vtranspose or vstranspose. Source operand 220 specifies a first location within a first storage space of memory. In block 430, the arithmetic logic operation executed by processor 110 includes vector transposing the data to be processed at the first location within the first storage space to obtain the transposed data to be processed.

[0084] In some embodiments, instruction 140 includes an exponentiation instruction such as instruction vexp. Source operand 220 specifies a source storage location. In block 430, the arithmetic logic operation executed by processor 110 includes determining an exponential value with a predetermined numerical value as the base, taking the data to be processed at the source storage location as a power.

[0085] In some embodiments, instruction 140 includes a VM register instruction such as vm2index. Source operand 220 of instruction 140 indicates a source VM register within memory. In block 430, the arithmetic logic operation executed by processor 110 includes storing an index at an enabled position of the source VM register at a target storage location.

[0086] In some embodiments, the target storage location of each of the above-described instructions 140 includes a processing result vector. Target operand 230 further indicates a target VM register. The value at each position of the target VM register indicates whether the corresponding processing result is to be written to the corresponding position of the processing result vector. For example, if target register vm3[i] is 1, it indicates that the i-th element of the processing result vector word is writable and can be written to the corresponding processing result. Conversely, if target register vm3[i] is 0, the i-th element of the processing result vector word cannot be written to the corresponding processing result.

[0087] In some embodiments, instruction 140 includes a one-hot code conversion instruction such as instruction vindex2vm. Source operand 220 specifies a predetermined index value of the data to be processed within the second storage space of the memory. Target operand 230 specifies a target VM register. In block 430, processor 110 converts the value of the data to be processed at the predetermined index value into a one-hot code. Processor 110 is further configured to store the one-hot code in the target VM register.

[0088] FIG. 5 shows a block diagram of an electronic device 500 including a processor 110 according to one or more embodiments of the present invention. It should be understood that the electronic device 500 shown in FIG. 5 is merely exemplary and should not limit the functions and scope of the embodiments described herein.

[0089] As shown in FIG. 5, the electronic device 500 is in the form of a general-purpose electronic device or a computing device. The components of the electronic device 500 may include, but are not limited to, one or more processors 110, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. In some embodiments, the processor 110 can execute various processes based on a program stored in the memory 520. The processor 110 may be a multi-core processor, and the multi-core processor improves the parallel processing ability of the electronic device 500 by executing computer-executable instructions in parallel.

[0090] The electronic device 500 typically includes a plurality of computer storage media. Such media may be any accessible media that can be accessed by the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 may be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or a specific combination thereof. The storage device 530 may be removable or non-removable media, and may include machine-readable media such as flash memory drives, magnetic disks, or any other media, and can be used to store information and / or data (e.g., training data for training) and can be accessible within the electronic device 500.

[0091] The electronic device 500 can further include another removable / non-removable, volatile / non-volatile storage medium. Although not shown in FIG. 5, a magnetic disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive may be connected to a path (not shown) by one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules, and these program modules are configured to execute various methods or operations of various embodiments of the present invention. For example, these program modules may be configured to implement various functions or operations of the processor 110, such as implementing the functions of the instruction decoder 120 and the arithmetic logic unit 130.

[0092] The communication unit 540 implements communication with other computing devices via a communication medium. Additionally, the functions of the components of the electronic device 500 may be implemented as a single computing cluster or multiple computing machines, which can communicate via a communication connection. Thus, the electronic device 500 can operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or other network nodes.

[0093] The input device 550 may be one or more input devices such as a mouse, keyboard, trackball, etc. The output device 560 may be one or more output devices such as a display, speaker, printer, etc. The electronic device 500 can further communicate with one or more external devices (not shown), such as a storage device, display device, etc., via the communication unit 540 as needed, communicate with one or more devices that enable a user to interact with the electronic device 500, or the electronic device 500 can perform communication with one or more other electronic devices or any device for computing device communication (such as a network card, modem, etc.). Such communication can be carried out via an input / output (I / O) interface (not shown).

[0094] According to an exemplary implementation of the present invention, there is provided a computer-readable storage medium storing one or more computer instructions, and the one or more computer instructions are executed by a processor to implement the above method. According to an exemplary implementation of the present invention, there is further provided a computer program product, and the computer program product is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions executed by a processor to implement the above method.

[0095] Here, each aspect of the present invention has been described with reference to the flowcharts and / or block diagrams of the method, apparatus (system), and computer program product implemented by the present invention. It should be understood that each box in the flowcharts and / or block diagrams and combinations of boxes in the flowcharts and / or block diagrams can all be implemented by computer-readable program instructions.

[0096] These computer-readable program instructions are provided to the processing unit of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to generate a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing apparatus, an apparatus for implementing the functions / operations specified in one or more boxes in the flowchart and / or block diagram can be generated. These computer-readable program instructions may be stored in a computer-readable storage medium, and by operating a computer, programmable data processing apparatus, and / or other devices in a specific manner, the computer-readable medium in which the instructions are stored constitutes a manufactured product including instructions for implementing each aspect of the functions / operations specified in one or more boxes in the flowchart and / or block diagram.

[0097] By loading the computer-readable program instructions into a computer, other programmable data processing apparatus, or other devices, a series of operation steps are executed on the computer, other programmable data processing apparatus, or other devices to generate a process implemented by the computer, whereby the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / operations specified in one or more boxes in the flowchart and / or block diagram.

[0098] According to one or more embodiments of the present invention, Example 1 describes a processor, which includes an instruction decoder configured to decode a target instruction for vector operations. The target instruction includes a target opcode, a source operand, and a target operand. The target opcode indicates the vector operation specified by the target instruction. The source operand specifies at least a source storage location for reading the data to be processed in the memory. The target operand specifies at least a target storage location for writing the processing result in the memory. The processor further includes an arithmetic logic unit coupled to the instruction decoder and the memory. The arithmetic logic unit is configured to read the data to be processed from the source storage location of the memory, execute an arithmetic logic operation associated with the vector operation specified by the target instruction on the data to be processed, and write the processing result of the data to be processed to the target storage location of the memory.

[0099] According to one or more embodiments of the present invention, Example 2 includes the processor described based on Example 1, where the target instruction includes a first index determination instruction, the source operand specifies a position in the first storage space of the memory, and the source operand further specifies a predetermined index value of the data to be processed in the second storage space of the memory. The arithmetic logic unit is configured to determine a first index indicating a storage position in the first storage space of the value of the data to be processed at the position indicated by the predetermined index value in order to execute an arithmetic logic operation associated with the vector operation specified by the target instruction.

[0100] According to one or more embodiments of the present invention, Example 3 includes the processor described based on Example 1, where the target instruction includes a second index determination instruction, the source operand specifies a position in the first storage space of the memory, and the source operand further specifies a first immediate number. The arithmetic logic unit is configured to determine a second index indicating the storage position of the first immediate number in the first storage space in order to execute an arithmetic logic operation associated with the vector operation specified by the target instruction.

[0101] According to one or more embodiments of the present invention, Example 4 includes the processor described based on Example 1, where the target instruction includes a first numerical determination instruction, the source operand specifies a position in the first storage space of the memory, and the source operand further specifies a predetermined index value of the data to be processed in the second storage space of the memory. The arithmetic logic unit is configured to determine a predetermined value of the data to be processed at the position indicated by the predetermined index value and determine it as a first value at a position in the first storage space using the predetermined value as an index, in order to execute an arithmetic logic operation associated with the vector operation specified by the target instruction.

[0102] According to one or more embodiments of the present invention, Example 5 includes the processor described based on Example 1, where the target instruction includes a second numerical determination instruction, and the source operand specifies a predetermined index value of the data to be processed in the second storage space of the memory. The arithmetic logic unit is configured to determine a second value at the position indicated by the predetermined index value of the data to be processed, in order to execute an arithmetic logic operation associated with the vector operation specified by the target instruction.

[0103] According to one or more embodiments of the present invention, Example 6 includes the processor described based on Example 1, where the target instruction includes a vector transposition instruction, and the source operand specifies at least a first position in the first storage space of the memory. The arithmetic logic unit is configured to transpose the data to be processed at the first position in the first storage space to obtain the transposed data to be processed, in order to execute an arithmetic logic operation associated with the vector operation specified by the target instruction.

[0104] According to one or more embodiments of the present invention, Example 7 includes the processor described based on Example 1, where the target instruction includes an exponential instruction, and the source operand specifies a source storage location. The arithmetic logic unit is configured to determine an exponential value with a predetermined numerical value as the base, taking the data to be processed at the source storage location as the power, in order to execute an arithmetic logic operation associated with the vector operation specified by the target instruction.

[0105] According to one or more embodiments of the present invention, Example 8 includes the processor described based on Example 1, where the target instruction includes a vector mask VM register instruction, and the source operand indicates a source VM register in the memory. The arithmetic logic unit is configured to store an index at the enabled position of the source VM register in the target storage location in order to execute an arithmetic logic operation associated with the vector operation specified by the target instruction.

[0106] According to one or more embodiments of the present invention, Example 9 includes the processor described in any one of Examples 2 to 8, where the target storage location includes a processing result vector, the target operand further indicates a target vector mask VM register, and the value at each position of the target VM register indicates whether the corresponding processing result is written to the corresponding position of the processing result vector.

[0107] According to one or more embodiments of the present invention, Example 10 includes the processor described based on Example 1, where the target instruction includes a one-hot code conversion instruction, the source operand specifies a predetermined index value of the data to be processed in the second storage space of the memory, and the target operand specifies a target vector mask VM register. The arithmetic logic unit is configured to convert the value of the data to be processed at the predetermined index value into a one-hot code and store the one-hot code in the target VM register in order to execute an arithmetic logic operation associated with the vector operation specified by the target instruction.

[0108] According to one or more embodiments of the present invention, Example 11 describes a data processing method. The method includes decoding a target instruction for vector operations, where the target instruction includes a target opcode, a source operand, and a target operand. The target opcode indicates the vector operation specified by the target instruction. The source operand at least specifies a source storage location for reading the data to be processed in the memory. The target operand at least specifies a target storage location for writing the processing result in the memory. The method further includes reading the data to be processed from the source storage location of the memory, performing an arithmetic logic operation associated with the vector operation specified by the target instruction on the data to be processed, and writing the processing result of the data to be processed to the target storage location of the memory.

[0109] According to one or more embodiments of the present invention, Example 12 includes the method described based on Example 11, where the target instruction includes an indexing instruction, the source operand specifies a position in the first storage space of the memory, and the source operand further specifies at least one of a predetermined index value of the data to be processed in the second storage space of the memory and a first immediate value number. Performing the arithmetic logic operation associated with the vector operation specified by the target instruction includes at least one of determining a first index indicating a storage position in the first storage space of the value of the data to be processed at the position indicated by the predetermined index value and determining a second index indicating a storage position of the first immediate value number in the first storage space.

[0110] According to one or more embodiments of the present invention, Example 13 includes the method described based on Example 11, where the target instruction includes a first numerical determination instruction, the source operand designates a position in the first storage space of the memory, and the source operand further designates a predetermined index value of the data to be processed within the second storage space of the memory. Executing the arithmetic logic operation associated with the vector operation specified by the target instruction includes determining a predetermined value of the data to be processed at the position indicated by the predetermined index value and determining a first value at the position in the first storage space using the predetermined value as an index.

[0111] According to one or more embodiments of the present invention, Example 14 includes the method described based on Example 11, where the target instruction includes a second numerical determination instruction and the source operand designates a predetermined index value of the data to be processed within the second storage space of the memory. Executing the arithmetic logic operation associated with the vector operation specified by the target instruction includes determining a second value at the position indicated by the predetermined index value of the data to be processed.

[0112] According to one or more embodiments of the present invention, Example 15 includes the method described based on Example 11, where the target instruction includes a vector transposition instruction, and here the source operand designates at least a first position within the first storage space of the memory. Executing the arithmetic logic operation associated with the vector operation specified by the target instruction includes vector transposing the data to be processed at the first position within the first storage space to obtain the transposed data to be processed.

[0113] According to one or more embodiments of the present invention, Example 16 includes the method described based on Example 11, where the target instruction includes an exponent instruction and the source operand designates a source storage position. Executing the arithmetic logic operation associated with the vector operation specified by the target instruction includes determining an exponent value with a predetermined numerical value as the base, using the data to be processed at the source storage position as a power.

[0114] According to one or more embodiments of the present invention, Example 17 includes the method described based on Example 11, where the target instruction includes a vector mask VM register instruction, and the source operand indicates a source VM register in memory. Executing the arithmetic logic operation associated with the vector operation specified by the target instruction includes storing an index at the enabled position of the source VM register in the target memory location.

[0115] According to one or more embodiments of the present invention, Example 18 includes the method described based on Example 11, where the target instruction includes a one-hot code conversion instruction, the source operand specifies a predetermined index value of the data to be processed in the second storage space of the memory, and the target operand specifies a target vector mask VM register. Executing the arithmetic logic operation associated with the vector operation specified by the target instruction includes converting the value of the data to be processed at the predetermined index value into a one-hot code and storing the one-hot code in the target VM register.

[0116] According to one or more embodiments of the present invention, Example 19 describes an electronic device, which includes at least the processor described in any one of Examples 1 to 10.

[0117] According to one or more embodiments of the present invention, Example 20 describes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any one of Examples 11 to 18 is implemented.

[0118] The flowcharts and block diagrams in the drawings illustrate the implementable architectures, functions, and operations of multiple implementable systems, methods, and computer program products according to the present invention. In this regard, each box in the flowchart or block diagram can represent one module, program segment, or part of an instruction, and the module, program segment, or part of an instruction includes one or more executable instructions for implementing the specified logical function. In some implementations as an alternative, the functions represented in the boxes may occur in an order different from that shown in the drawings. For example, two consecutive boxes may actually be executed substantially in parallel, or depending on the functions involved, may be executed in the reverse order. It should also be noted that each box in the block diagram and / or flowchart, and combinations of boxes in the block diagram and / or flowchart, may be implemented by a special-purpose hardware-based system that executes the specified function or operation, or may be implemented by a combination of special-purpose hardware and computer instructions.

[0119] As described above for each implementation of the present invention, the above description is illustrative and not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best interpret the principles of each implementation, its practical application, or improvements to the technology in the market, or to enable those skilled in the art to understand each implementation form disclosed herein.

Claims

1. An instruction decoder configured to decode a target instruction for vector operations, wherein the target instruction includes a target opcode, a source operand, and a target operand, the target opcode indicates a vector operation specified by the target instruction, the source operand specifies at least a source storage location for reading data to be processed in a memory, and the target operand specifies at least a target storage location for writing a processing result in the memory; an instruction decoder, An arithmetic logic unit coupled to the instruction decoder and the memory, configured to read the data to be processed from the source storage location of the memory, execute an arithmetic logic operation associated with the vector operation specified by the target instruction on the data to be processed, and write a processing result of the data to be processed to the target storage location of the memory; a processor comprising: Processor.

2. The target instruction includes a first index determination instruction, the source operand specifies a location in a first storage space of the memory, and the source operand further specifies a predetermined index value of the data to be processed in a second storage space of the memory, The arithmetic logic unit is configured to determine a first index indicating a storage location in the first storage space of a value of the data to be processed at a location indicated by the predetermined index value in order to execute an arithmetic logic operation associated with the vector operation specified by the target instruction; a processor according to claim 1. The processor according to claim 1.

3. The target instruction includes a second index determination instruction, the source operand specifies a location in a first storage space of the memory, and the source operand further specifies a first immediate number, The arithmetic logic unit is configured to determine a second index indicating a storage location of the first immediate number in the first storage space in order to execute an arithmetic logic operation associated with the vector operation specified by the target instruction; a processor according to claim 1. The processor according to claim 1.

4. The target instruction includes a first numerical value determination instruction, the source operand specifies a location in a first storage space of the memory, and the source operand further specifies a predetermined index value of the data to be processed in a second storage space of the memory, The arithmetic logic unit is configured to determine a predetermined value of the data to be processed at a position indicated by the predetermined index value in order to execute an arithmetic logic operation associated with the vector operation specified by the target instruction, and determine the value as a first value at a position in the first storage space using the predetermined value as an index. The processor according to claim 1.

5. The target instruction includes a second numerical determination instruction, and the source operand specifies a predetermined index value of the data to be processed in the second storage space of the memory. The arithmetic logic unit is configured to determine a second value at a position indicated by the predetermined index value of the data to be processed in order to execute an arithmetic logic operation associated with the vector operation specified by the target instruction. The processor according to claim 1.

6. The target instruction includes a vector transposition instruction, and the source operand specifies at least a first position in the first storage space of the memory. The arithmetic logic unit is configured to transpose the data to be processed at the first position in the first storage space to obtain the transposed data to be processed in order to execute an arithmetic logic operation associated with the vector operation specified by the target instruction. The processor according to claim 1.

7. The target instruction includes an exponentiation instruction, and the source operand specifies the source storage position. The arithmetic logic unit is configured to determine an exponent value with a predetermined numerical value as the base by using the data to be processed at the source storage position as a power in order to execute an arithmetic logic operation associated with the vector operation specified by the target instruction. The processor according to claim 1.

8. The target instruction includes a vector mask VM register instruction, and the source operand indicates a source VM register in the memory. The arithmetic logic unit is configured to store the index at the enabled position of the source VM register at the target storage position in order to execute an arithmetic logic operation associated with the vector operation specified by the target instruction. The processor according to claim 1.

9. The target memory location includes a processing result vector, the target operand further indicates a target vector mask VM register, and the value at each position of the target VM register indicates whether a corresponding processing result is to be written to the corresponding position of the processing result vector. The processor according to any one of claims 2 to 8.

10. The target instruction includes a one-hot code conversion instruction, the source operand designates a predetermined index value of the data to be processed within the second storage space of the memory, and the target operand designates a target vector mask VM register. The arithmetic logic unit is configured to convert the value of the data to be processed at a predetermined index value into a one-hot code and store the one-hot code in the target VM register in order to execute an arithmetic logic operation associated with the vector operation designated by the target instruction. The processor according to claim 1.

11. Decoding a target instruction for a vector operation, the target instruction including a target opcode, a source operand, and a target operand, the target opcode indicating the vector operation designated by the target instruction, the source operand at least designating a source storage location for reading data to be processed within the memory, and the target operand at least designating a target storage location for writing a processing result within the memory. Reading the data to be processed from the source storage location of the memory. Executing an arithmetic logic operation associated with the vector operation designated by the target instruction on the data to be processed. Writing the processing result of the data to be processed to the target storage location of the memory. A data processing method.

12. The target instruction includes an index determination instruction, the source operand designates a position in the first storage space of the memory, and the source operand further designates at least one of a predetermined index value of the data to be processed within the second storage space of the memory and a first immediate number. Executing the arithmetic logic operation associated with the vector operation designated by the target instruction is Determining a first index indicating a storage position within the first storage space of the value of the data to be processed at the position indicated by the predetermined index value. Determining at least one of: determining a second index indicating a storage position of the first immediate value number in the first storage space; The data processing method according to claim 11.

13. The target instruction includes a first numerical determination instruction, the source operand designates a position in a first storage space of the memory, and the source operand further designates a predetermined index value of the data to be processed in a second storage space of the memory. Executing an arithmetic logic operation associated with the vector operation specified by the target instruction includes: Determining a predetermined value of the data to be processed at a position indicated by the predetermined index value; Determining, as a first value, a value at a position in the first storage space indexed by the predetermined value; The data processing method according to claim 11.

14. The target instruction includes a second numerical determination instruction, and the source operand designates a predetermined index value of the data to be processed in the second storage space of the memory. Executing an arithmetic logic operation associated with the vector operation specified by the target instruction includes determining a second value at a position indicated by the predetermined index value of the data to be processed. The data processing method according to claim 11.

15. The target instruction includes a vector transposition instruction, and the source operand designates at least a first position in the first storage space of the memory. Executing an arithmetic logic operation associated with the vector operation specified by the target instruction includes vector transposing the data to be processed at the first position in the first storage space to obtain transposed data to be processed. The data processing method according to claim 11.

16. The target instruction includes an exponent instruction, and the source operand designates the source storage position. Executing an arithmetic logic operation associated with the vector operation specified by the target instruction includes determining an exponent value with a predetermined numerical value as the base, taking the data to be processed at the source storage position as a power. The data processing method according to claim 11.

17. The target instruction includes a vector mask VM register instruction, and the source operand indicates a source VM register in the memory. Executing an arithmetic logic operation associated with the vector operation specified by the target instruction includes storing an index at an enabled position of the source VM register in the target memory location. The data processing method according to claim 11.

18. The target instruction includes a one-hot code conversion instruction, the source operand specifies a predetermined index value of the data to be processed within the second storage space of the memory, and the target operand specifies a target vector mask VM register. Executing an arithmetic logic operation associated with the vector operation specified by the target instruction converting the value of the data to be processed at the predetermined index value into a one-hot code and storing the one-hot code in the target VM register. The data processing method according to claim 11.

19. An electronic device comprising at least the processor according to any one of claims 1 to 10. Electronic device.

20. A computer program is stored, and the computer program is executed by a processor to implement the data processing method according to any one of claims 11 to 18. Computer-readable storage medium.

Citation Information

Patent Citations

  • Information processing apparatus, arithmetic processing device and memory access control method

    JP2006285683A

  • Processing in neural networks

    JP2019079524A

  • Systems and methods for executing a fused multiply-add instruction for complex numbers

    US20180095758A1