Tensor processing unit and method, and computer-readable storage medium

Through a unified macro instruction, different computing modules are called for in-memory or non-in-memory calculations, the problems of programming complexity and computing efficiency in the prior art are solved, and efficient calculation of larger shape matrices is achieved, and power consumption is reduced.

WO2025124578A1PCT designated stage expired Publication Date: 2025-06-19SUZHOU YIZHU INTELLIGENT TECH CO LTD

Patent Information

Application Number
PCT/CN2024/139363
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-12-13
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

In the prior art, although the storage and computing integrated architecture can effectively overcome the bottleneck of the von Neumann architecture, it cannot support general matrix calculations with dynamic changes in weights, resulting in an increase in the complexity of programming interfaces.

Method used

A unified macro instruction is used to call different calculation modules for in-memory calculation or non-in-memory calculation. The instruction sequence module generates an appropriate instruction sequence based on tensor operation instructions and preset configuration information, and automatically calculates the source address and target address of the sub-operation data.

Benefits of technology

It reduces programming complexity and improves computing efficiency. It realizes large-shaped matrix multiplication through multiple matrix multiplication operations, hides hardware information, reduces instruction transmission and decoding overhead, and reduces power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024139363_19062025_PF_FP_ABST
    Figure CN2024139363_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application is a tensor processing unit. The tensor processing unit comprises: an instruction sequence module, which is configured to generate a first instruction sequence and / or a second instruction sequence on the basis of a tensor operation instruction and preset configuration information, wherein the first instruction sequence comprises a plurality of first operation instructions, the second instruction sequence comprises a plurality of second operation instructions, and the preset configuration information comprises first configuration information and / or second configuration information; a data loading module, which is configured to sequentially load, on the basis of the first instruction sequence or the second instruction sequence, into corresponding registers data to be operated; a first computing module, which is configured to sequentially read, on the basis of the first instruction sequence, data from the registers, so as to perform in-memory computing; and a second computing module, which is configured to sequentially read, on the basis of the second instruction sequence, data from the registers, so as to perform non-in-memory computing. Further provided in the present application is a tensor processing method. A unified macro instruction is used to call different computing modules for in-memory computing or non-in-memory computing, thereby reducing the programming complexity and improving the computing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Tensor processing device, method, and computer-readable storage medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 13, 2023, with application number 202311714060.5 and application name “Tensor processing device, method and computer-readable storage medium”, the entire contents of which are incorporated into this application by reference. Technical Field

[0002] The present invention relates to the field of computing, and more particularly, to a tensor processing device, method, and computer-readable storage medium. Background Art

[0003] Existing AI chip solutions for detection and computing tasks fall into two main categories: general-purpose processing devices such as CPUs (Central Processing Units), FPGAs (Field Programmable Gate Arrays), and GPUs (Graphics Processing Units); and specialized accelerators for artificial neural networks, such as Google's TPUs (Tensor Processing Units). These chips all utilize the von Neumann architecture, separating computation and storage, which work together to perform data access and computation. However, because processor design primarily focuses on improving computational speed, while storage prioritizes capacity expansion and cost optimization, a performance mismatch between "memory" and "computing" exists, leading to problems such as low memory access bandwidth, extended latency, and high power consumption—commonly known as the "memory wall" and "power wall." The more intensive the memory access, the more severe the "wall" problem, and the more difficult it is to increase computing power.

[0004] As a new computing architecture, integrated storage and computing fully integrates storage and computing, effectively overcoming the bottleneck of the von Neumann architecture and achieving orders of magnitude improvements in computing energy efficiency. However, in-memory computing is only suitable for general-purpose matrix calculations with fixed weights and does not support general-purpose matrix calculations with dynamically changing weights. Therefore, two different computing engines are required to perform general-purpose matrix calculations with fixed weights and general-purpose matrix calculations with dynamically changing weights, respectively. However, these two different computing engines have two different hardware interfaces and different matrix multiplication computing power, which increases the complexity of the programming interface. Summary of the Invention

[0005] In view of the above problems, the purpose of the present invention is to provide a tensor processing device, method and computer-readable storage medium, which use a unified macro instruction to call different computing modules for in-memory or non-memory calculations to reduce programming complexity.

[0006] According to a first aspect of the present invention, a tensor processing device is provided, comprising: an instruction sequence module for generating a first instruction sequence and / or a second instruction sequence based on tensor operation instructions and preset configuration information, wherein the first instruction sequence includes multiple first operation instructions, the second instruction sequence includes multiple second operation instructions, and the preset configuration information includes first configuration information and / or second configuration information; a data loading module for loading the data to be operated into the corresponding register in sequence according to the first instruction sequence or the second instruction sequence; a first computing module for reading data from the register in sequence according to the first instruction sequence for in-memory calculation; and a second computing module for reading data from the register in sequence according to the second instruction sequence for non-in-memory calculation.

[0007] Preferably, the first configuration information is used to characterize the data size, shape, starting address and addressing method of each calculation performed by the first computing module; the second configuration information is used to characterize the data size, shape, starting address and addressing method of each calculation performed by the second computing module.

[0008] Preferably, the instruction sequence module is further used to generate a selection signal according to the tensor operation instruction, and send the first instruction sequence to the first computing module or send the second instruction sequence to the second computing module according to the selection signal.

[0009] Preferably, the tensor operation instruction includes an operation code and an operation field, wherein the operation code includes an instruction type and an operation code, and the operation field includes a source address and a target address of the operation data.

[0010] Preferably, the instruction sequence module includes: an instruction parsing unit, which is used to parse the tensor operation instruction to obtain the operation code and the source address and target address of the operation data; an instruction sequence unit, which is used to automatically calculate the source address and target address of multiple sub-operation data based on preset configuration information and the source address of the operation data, and generate multiple first operation instructions and / or multiple second operation instructions based on the operation code and the source address and target address of the multiple sub-operation data.

[0011] Preferably, the instruction sequence unit is also used to automatically calculate the addresses of multiple first sub-operation data based on the first configuration information and the source address of the operation data to obtain multiple first operation instructions; and / or automatically calculate the addresses of multiple second sub-operation data based on the second configuration information and the source address of the operation data to obtain multiple second operation instructions.

[0012] Preferably, the instruction sequence module further includes: a control unit for generating a selection signal according to a source address of the operation data, and sending the first instruction sequence to the first computing module or sending the second instruction sequence to the second computing module according to the selection signal.

[0013] Preferably, the preset configuration information is set in a tensor operation instruction or a configuration register.

[0014] Preferably, the instruction sequence module further includes: a configuration information unit, configured to obtain the first configuration information or the second configuration information from the instruction parsing unit or the configuration register according to the selection signal.

[0015] Preferably, the source address of the operation data is any one of an on-chip memory, an off-chip memory, and a register.

[0016] Preferably, the register includes an input register, a weight register and an accumulation register, wherein the input register is used to store activation data; the weight register is used to store weight data; and the accumulation register is used to store the intermediate result of multiplication and addition of activation data and weight data.

[0017] According to a second aspect of the present invention, a tensor processing method is provided, comprising: generating a first instruction sequence and / or a second instruction sequence according to a tensor operation instruction and preset configuration information, wherein the first instruction sequence includes multiple first operation instructions, the second instruction sequence includes multiple second operation instructions, and the preset configuration information includes first configuration information and / or second configuration information; generating a selection signal according to the tensor operation instruction, and executing the first instruction sequence or the second instruction sequence according to the selection signal; loading the data to be operated into the corresponding register in sequence according to the first instruction sequence or the second instruction sequence; reading data from the register in sequence according to the first instruction sequence for in-memory calculation or reading data from the register in sequence according to the second instruction sequence for non-memory calculation.

[0018] Preferably, the first configuration information is used to characterize the data size, shape, starting address and addressing method of each calculation performed by the first computing module; the second configuration information is used to characterize the data size, shape, starting address and addressing method of each calculation performed by the second computing module.

[0019] Preferably, the tensor operation instruction includes an operation code and an operation field, wherein the operation code includes an instruction type and an operation code, and the operation field includes a source address and a target address of the operation data.

[0020] Preferably, generating a first instruction sequence and / or a second instruction sequence according to tensor operation instructions and preset configuration information includes: parsing the tensor operation instructions to obtain the opcode and the source address and target address of the operation data; automatically calculating the source address and target address of multiple sub-operation data according to the preset configuration information and the source address of the operation data, and generating multiple first operation instructions and / or multiple second operation instructions according to the opcode and the source address and target address of the multiple sub-operation data.

[0021] Preferably, the preset configuration information is set in a tensor operation instruction or a configuration register.

[0022] Preferably, the tensor processing method further includes: acquiring the first configuration information or the second configuration information according to a selection signal.

[0023] According to a third aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any one of the above methods is implemented.

[0024] The tensor processing device, method, and computer-readable storage medium provided by the present invention use a unified macro instruction to call different computing modules to perform in-memory or non-memory calculations, thereby reducing programming complexity and improving computing efficiency.

[0025] Furthermore, the tensor processing device, i.e., the hardware itself, can automatically calculate the source address and target address of multiple sub-operation data based on preset configuration information and the source address and target address of the operation data, and realize matrix multiplication operations of larger shapes through multiple matrix multiplication operations. In this way, the hardware information can be hidden and only one macro instruction is needed to realize matrix multiplication operations of larger shapes, which greatly reduces the complexity of programming, reduces the overhead of instruction issuance and decoding, and reduces power consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:

[0027] FIG1 is a schematic structural diagram of a tensor processing device provided by an embodiment of the present invention;

[0028] FIG2 is a schematic diagram showing the structure of an instruction sequence module provided by an embodiment of the present invention;

[0029] FIG3 is a schematic diagram showing a matrix multiplication according to an embodiment of the present invention;

[0030] FIG4 shows a block diagram of a computing system according to an embodiment of the present invention;

[0031] FIG5 shows a structural block diagram of another computing system provided by an embodiment of the present invention;

[0032] FIG6 shows a flowchart of a tensor processing method provided by an embodiment of the present invention;

[0033] FIG7 shows a flowchart of step S610 provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0034] Various embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. In each of the accompanying drawings, identical elements are represented by identical or similar reference numerals. For the sake of clarity, the various parts in the accompanying drawings are not drawn to scale.

[0035] The specific implementation of the present invention is further described in detail below with reference to the accompanying drawings and examples.

[0036] As mentioned above, in-memory computation is only suitable for general-purpose matrix calculations with fixed weights and does not support general-purpose matrix calculations with dynamically changing weights. Therefore, two different computation engines are required to perform general-purpose matrix calculations with fixed weights and with dynamically changing weights, respectively. However, these two different computation engines have two different sets of hardware interfaces and differ in their matrix multiplication capabilities, increasing the complexity of the programming interface.

[0037] In response to the above technical problems, the basic concept of this application is to use the same macro instruction to call different computing modules to perform in-memory or non-memory calculations, thereby reducing programming complexity.

[0038] Figure 1 shows a schematic diagram of the structure of a tensor processing device provided by an embodiment of the present invention. As shown in Figure 1, the tensor processing device 100 includes an instruction sequence module 110, a data loading module 120, a first calculation module 130, a second calculation module 140, an input register 150, a weight register 160, and an accumulation register 170.

[0039] The instruction sequence module 110 is configured to generate a first instruction sequence and / or a second instruction sequence according to the tensor operation instruction and preset configuration information, wherein the first instruction sequence includes a plurality of first operation instructions and the second instruction sequence includes a plurality of second operation instructions.

[0040] In this embodiment, the tensor operation instruction includes at least one of a matrix multiplication operation (MM), a matrix multiply-add operation (MAC), and a convolution operation (Conv).

[0041] The instruction format of a tensor operation instruction includes an opcode and an operation field, wherein the opcode includes an instruction type and an operation code, and the operation field includes at least the source address and target address of the operation data. Specifically, the instruction type is used to describe the type of operation involved, and the instruction type is a tensor operation. The operation code is used to describe the operation to be performed by the calculation instruction (for example, addition, subtraction, multiplication, division, special functions, etc.), which specifically describes the nature and function of the operation.

[0042] The operation domain includes the source address and destination address of the operation data. The source address and destination address can be memory addresses or register addresses (i.e., register numbers). The memory or register can be off-chip memory, or in actual applications, on-chip memory, used to store data. The data can specifically be n-dimensional data (tensor), where n is an integer greater than or equal to 2. For example, if n = 2, it is a two-dimensional data, i.e., a matrix. If n = 3 or more, it is a multi-dimensional tensor.

[0043] In this embodiment, the opcode can be the part of the instruction or field (usually represented by a code) specified in the computer program to perform the operation, which is an instruction sequence number used to inform the device executing the instruction which specific instruction needs to be executed. The operation domain can be the source of all data required to execute the corresponding instruction, such as the corresponding address, etc. All data required to execute the corresponding instruction include the data to be operated and the corresponding instruction processing method, etc. For an operation instruction, it must include an opcode and an operation domain, wherein the operation domain includes at least the source address and the target address of the operation data. It should be understood that those skilled in the art can set the instruction format of the operation instruction and the opcode and operation domain contained therein as needed, and the present disclosure does not limit this.

[0044] In one possible implementation, the instruction format of the operation instruction may be as shown in the following table:

[0045] Table 1:

[0046] Wherein, type, opcode is the operation code of the computing instruction, dst, scr is the operation domain of the computing instruction, and dst is the target address. src is the source address of the data to be operated on. When there are multiple data to be operated on, src can include multiple data to be operated on src0, src1, ..., srcn, and dst can also include multiple target addresses dst0, dst1, ..., dstn. This disclosure does not limit this.

[0047] In some specific embodiments, the tensor operation instruction may be as shown in the following table:

[0048] Table 2:

[0049] When executing the operation instructions shown in the above table, the tensor processing device 100 reads Num1 tensor data from the address specified in the register Reg1, reads Num2 weight data from the address specified in the register Reg2, and then performs a convolution operation or a multiplication and addition operation, and then stores the calculation result in the address space specified in the register Reg3.

[0050] In a preferred embodiment, the source address of the operation data can be the starting address of the storage space where the operation data is located. The tensor processing device 100 can obtain instructions and data through a data input and output unit, which can be one or more data I / O interfaces or I / O pins. Furthermore, the tensor processing device 100 can determine the operation data based on the source address of the operation data and obtain the operation data. Of course, in other embodiments, the tensor processing device 100 can also determine the operation data required for the operation based on the opcode of the tensor operation instruction.

[0051] In this embodiment, the preset configuration information includes first configuration information and / or second configuration information. The first configuration information is used to characterize the data size, shape, and related information required for address calculation each time calculated by the first calculation module; the second configuration information is used to characterize the data size, shape, and related information required for address calculation each time calculated by the second calculation module. The data size calculated each time by the first calculation module and the data size calculated each time by the second calculation module can be the same or different, and can be set arbitrarily as needed. Among them, the relevant information required for address calculation includes the starting address and addressing method of the data, etc. The addressing method can be an interleave addressing method, but is not limited to this.

[0052] The instruction sequence module 110 is further configured to generate a selection signal according to the tensor operation instruction, and send the first instruction sequence to the first computing module or the second instruction sequence to the second computing module according to the selection signal.

[0053] In this embodiment, the instruction sequence module 110 converts the tensor operation instruction into multiple first operation instructions according to the first configuration information, or converts the tensor operation instruction into multiple second operation instructions according to the second configuration information. The operation data of each operation instruction includes first operation data A and second operation data B. The first operation data A is divided into multiple first sub-operation data, and the second operation data B is divided into multiple second sub-operation data. Each sub-operation data is a data block. The address of the first sub-operation data in each row is moved in sequence along the row direction, and the address of the second sub-operation data in each column is moved in sequence along the column direction. The first sub-operation data in each row is multiplied by the second sub-operation data in each column to obtain a corresponding sub-matrix block, and finally a matrix C is obtained.

[0054] In this embodiment, the instruction sequence module 110 includes an instruction parsing unit 111, an instruction sequence unit 112, a control unit 113, and a configuration information unit 114. The instruction parsing unit 111 is used to parse the tensor operation instruction to obtain the opcode and the source address and target address of the operation data; the instruction sequence unit 112 is used to automatically calculate the source address and target address of multiple sub-operation data according to preset configuration information and the source address and target address of the operation data, and generate multiple first operation instructions and / or multiple second operation instructions according to the opcode and the source address and target address of the multiple sub-operation data; the control unit 113 is used to generate a selection signal according to the source address of the operation data, and send the first instruction sequence to the first computing module or send the second instruction sequence to the second computing module according to the selection signal; the configuration information unit 114 is used to obtain the first configuration information or the second configuration information according to the selection signal, or to parse the tensor operation instruction to obtain the first configuration information or the second configuration information.

[0055] In this embodiment, the preset configuration information can be set in the tensor operation instruction, that is, the tensor operation instruction also includes the preset configuration information, or can be placed in a separate configuration register. When the tensor operation instruction also includes the preset configuration information, the instruction parsing unit parses the tensor operation instruction to obtain the preset configuration information.

[0056] In this embodiment, the operation data includes activation data and weight data; if the source address of the weight data is a storage-computing integrated unit, the weight data is static data, and the control unit 113 sends a first instruction sequence to the first computing module 130 according to the selection signal, and the configuration information unit 114 obtains the first configuration information from the instruction parsing unit or the configuration register according to the selection signal. If the source address of the weight data is a memory, the weight data is dynamic data, and the control unit 113 sends a second instruction sequence to the second computing module 140 according to the selection signal, and the configuration information unit 114 obtains the second configuration information from the instruction parsing unit or the configuration register according to the selection signal.

[0057] Since the hardware itself can only calculate matrix multiplication operations of a certain size, if you want to implement matrix multiplication operations of larger shapes, you usually need to obtain hardware information, that is, the size and shape of the data calculated by the hardware each time, during compilation. Then, you need to split a single instruction into multiple operation instructions. This increases the complexity of programming and increases the overhead of instruction issuance and decoding. However, the tensor processing device of the present application, that is, the hardware itself, can automatically calculate the source and target addresses of multiple sub-operation data based on preset configuration information and the source and target addresses of the operation data, and implement matrix multiplication operations of larger shapes through multiple matrix multiplication operations. In this way, the hardware information can be hidden and only a single macro instruction is needed to implement matrix multiplication operations of larger shapes, greatly reducing the complexity of programming, reducing the overhead of instruction issuance and decoding, and reducing power consumption.

[0058] The data loading module 120 is used to load the data to be operated into the corresponding registers in sequence according to the first instruction sequence or the second instruction sequence.

[0059] In this embodiment, the first operation data is sequentially loaded into the input register 150 according to a plurality of first operation instructions in the first instruction sequence. The first operation data is sequentially loaded into the input register 150 and the second operation data is sequentially loaded into the weight register 160 according to a plurality of second operation instructions in the second instruction sequence.

[0060] The first calculation module 130 is used to read data from the register in sequence according to the first instruction sequence to perform in-memory calculations.

[0061] In this embodiment, the first computing module 122 can perform computation in memory (CIM) and is composed of an SRAM, ReRAM, or other storage medium integrated computing unit.

[0062] The second calculation module 140 is used to read data from the register in sequence according to the second instruction sequence to perform non-memory calculations.

[0063] In this embodiment, the non-in-memory calculation includes near-memory calculation or other calculation, such as general matrix multiplication (GEMM).

[0064] The input register 150 is used to store activation data; the weight register 160 is used to store weight data; and the accumulation register 170 is used to store an intermediate result of multiplication and addition of the activation data and the weight data.

[0065] Referring to FIG3 , this embodiment uses a matrix A of 512*256 and a matrix B of 256*512, with a minimum calculation size of 8*64, as an example for explanation, but is not limited thereto. Matrix A and matrix B are divided into at least 8*64 data blocks. The minimum data block in the first column of matrix A is shifted by 4 along the column direction, and the minimum data block in the first row of matrix B is shifted by 4 along the row direction. The intermediate results of the matrix multiplication calculation are then placed in accumulation register 170 for cumulative calculation. Correspondingly, the minimum data block in the next column of matrix A is shifted by 4 along the column direction, and the minimum data block in the next row of matrix B is shifted by 4 along the row direction. The intermediate results of the matrix multiplication calculation are then placed in accumulation register 170 for cumulative calculation. Similarly, the submatrix within the dotted box of matrix A and the submatrix within the dotted box of matrix B complete the matrix multiplication calculation to obtain submatrix block 0 of matrix C. Repeated calculations in this manner can complete the matrix multiplication operation of matrix A and matrix B. FIG3 schematically illustrates a data block division method and a calculation sequence, and the data block division method and the calculation sequence can be configured according to actual needs, and the present invention does not limit this.

[0066] The tensor processing device provided by the present invention uses a unified macro instruction to call different computing modules to perform in-memory or non-memory computing, thereby reducing programming complexity and improving computing efficiency.

[0067] Furthermore, the tensor processing device, i.e., the hardware itself, automatically calculates the source address and target address of multiple sub-operation data according to preset configuration information and the source address and target address of the operation data, and realizes matrix multiplication operations of larger shapes through multiple matrix multiplication operations. In this way, the hardware information can be hidden and only one macro instruction is needed to realize matrix multiplication operations of larger shapes, which greatly reduces the complexity of programming, reduces the overhead of instruction issuance and decoding, and reduces power consumption.

[0068] 4 shows a block diagram of a computing system according to an embodiment of the present invention. The computing system includes a tensor processing device 100 and a general processing device 200. The general processing device 200 is used to send tensor operation instructions to the tensor processing device 100.

[0069] In this embodiment, the first computing module 130 and the second computing module 140 are connected to the general processing device 200 using a unified interface, which can be one of a PCIE interface and a UCIE interface. In this embodiment, the activation data and output data of the first computing module 130 and the second computing module 140 have the same data structure. A single instruction can be used to drive the first computing module 130 and the second computing module 140 to perform calculations, and an instruction can also be used to switch between the first computing module 130 and the second computing module 140 for calculations.

[0070] In this embodiment, the tensor processing device 100 and the general processing device 200 can be packaged in the same core. In a preferred embodiment, referring to FIG4 , the tensor processing device 100 and the general processing device 200 can also be packaged into different cores and integrated into the same chip or deployed on different chips. The positional relationship between the general processing device 110 and the tensor processing device 120 can be set according to actual applications and is not limited thereto.

[0071] FIG5 is a schematic diagram showing the structure of a computing system provided by another embodiment of the present invention. As shown in FIG5 , the computing system includes a task distribution module 510 and a plurality of computing devices 520 .

[0072] This embodiment is described by taking two computing devices 520A and 520B as an example, but is not limited thereto.

[0073] In this embodiment, the task distribution module 510 is used to distribute computing instructions to multiple computing devices 520. The computing devices 520 receive the computing instructions and perform corresponding calculations according to the computing instructions. The computing devices 520A and 520B include a tensor processing device 521 and a general processing device 522. The tensor processing device 521 is the same as that described in the above embodiment and will not be repeated here.

[0074] In other embodiments, the computing devices 520A and 520B only include the general processing device 522 , the tensor processing device 521 is located outside the computing devices, and multiple computing devices 520 share the same tensor processing device.

[0075] Figure 6 shows a flow chart of a tensor processing method provided by an embodiment of the present invention. Referring to Figure 6 , the calculation method is executed by the tensor processing device 100 provided by the above embodiment, and includes the following steps.

[0076] In step S610, a first instruction sequence and / or a second instruction sequence is generated based on the tensor operation instruction and preset configuration information, wherein the first instruction sequence includes multiple first operation instructions, the second instruction sequence includes multiple second operation instructions, and the preset configuration information includes first configuration information and / or second configuration information.

[0077] In this embodiment, referring to FIG. 7 , step S610 includes steps S611 to S613 .

[0078] In step S611, the tensor operation instruction is parsed to obtain the operation code, the source address and the target address of the operation data, and the preset configuration information.

[0079] The tensor operation instruction includes at least one of a matrix multiplication operation (MM), a matrix multiplication-addition operation (MAC), and a convolution operation (Conv).

[0080] The instruction format of a tensor operation instruction includes an opcode and an operation field, wherein the opcode includes an instruction type and an operation code, and the operation field includes at least the source address and target address of the operation data. Specifically, the instruction type is used to describe the type of operation involved, and the instruction type is a tensor operation. The operation code is used to describe the operation to be performed by the calculation instruction (for example, addition, subtraction, multiplication, division, special functions, etc.), which specifically describes the nature and function of the operation.

[0081] The operation domain includes the source address and destination address of the operation data. The source address and destination address can be memory addresses or register addresses (i.e., register numbers). The memory or register can be off-chip memory, or in actual applications, on-chip memory, used to store data. The data can specifically be n-dimensional data (tensor), where n is an integer greater than or equal to 2. For example, if n = 2, it is a two-dimensional data, i.e., a matrix. If n = 3 or more, it is a multi-dimensional tensor.

[0082] In this embodiment, the opcode can be the part of the instruction or field (usually represented by a code) specified in the computer program to perform the operation, which is an instruction sequence number used to inform the device executing the instruction which specific instruction needs to be executed. The operation domain can be the source of all data required to execute the corresponding instruction, such as the corresponding address, etc. All data required to execute the corresponding instruction include the data to be operated and the corresponding instruction processing method, etc. For an operation instruction, it must include an opcode and an operation domain, wherein the operation domain includes at least the source address and the target address of the operation data. It should be understood that those skilled in the art can set the instruction format of the operation instruction and the opcode and operation domain contained therein as needed, and the present disclosure does not limit this.

[0083] In a preferred embodiment, the source address of the operation data can be the starting address of the storage space where the operation data is located. The tensor processing device 100 can obtain instructions and data through a data input and output unit, which can be one or more data I / O interfaces or I / O pins. Furthermore, the tensor processing device 100 can determine the operation data based on the source address of the operation data and obtain the operation data. Of course, in other embodiments, the tensor processing device 100 can also determine the operation data required for the operation based on the opcode of the tensor operation instruction.

[0084] In step S612, preset configuration information is acquired according to the selection signal, wherein the preset configuration signal includes the first configuration information and / or the second configuration information.

[0085] In this embodiment, the operation data includes activation data and weight data. If the source address of the weight data is the storage-computing integrated unit, the weight data is static data, and the first configuration information is obtained and the first instruction sequence is executed according to the selection signal. If the source address of the weight data is the memory, the weight data is dynamic data, and the second configuration information is obtained and the second instruction sequence is executed according to the selection signal.

[0086] In step S613, the source addresses and target addresses of multiple sub-operation data are automatically calculated based on the preset configuration information and the source address and target address of the operation data, and multiple first operation instructions and / or multiple second operation instructions are generated based on the operation code and the source address and target address of the multiple sub-operation data.

[0087] In this embodiment, the preset configuration information includes first configuration information and / or second configuration information. The first configuration information is used to characterize the data size, shape, and related information required for address calculation each time calculated by the first calculation module; the second configuration information is used to characterize the data size, shape, and related information required for address calculation each time calculated by the second calculation module. The data size calculated each time by the first calculation module and the data size calculated each time by the second calculation module can be the same or different, and can be set arbitrarily as needed. Among them, the relevant information required for address calculation includes the starting address and addressing method of the data, etc. The addressing method can be an interleave addressing method, but is not limited to this.

[0088] Specifically, the tensor operation instruction is converted into a plurality of first operation instructions according to the first configuration information, or the tensor operation instruction is converted into a plurality of second operation instructions according to the second configuration information. The operation data of each operation instruction includes first operation data A and second operation data B, the first operation data A is divided into a plurality of first sub-operation data, the second operation data B is divided into a plurality of second sub-operation data, each sub-operation data is a data block, the address of the first sub-operation data of each row is sequentially moved along the row direction, and the address of the second sub-operation data of each column is sequentially moved along the column direction, the first sub-operation data of each row is multiplied by the second sub-operation data of each column to obtain a corresponding sub-matrix block, and finally a matrix C is obtained.

[0089] In step S620, a selection signal is generated according to the tensor operation instruction, and the first instruction sequence or the second instruction sequence is executed according to the selection signal. In this embodiment, the preset configuration information can be set in the tensor operation instruction, that is, the tensor operation instruction also includes the preset configuration information, or it can be placed in a separate configuration register. When the tensor operation instruction also includes the preset configuration information, the instruction parsing unit parses the tensor operation instruction to obtain the preset configuration information.

[0090] In step S630 , the data to be operated is loaded into the corresponding registers in sequence according to the first instruction sequence or the second instruction sequence.

[0091] In this embodiment, the first operation data is sequentially loaded into the input register 150 according to a plurality of first operation instructions in the first instruction sequence. The first operation data is sequentially loaded into the input register 150 and the second operation data is sequentially loaded into the weight register 160 according to a plurality of second operation instructions in the second instruction sequence.

[0092] In step S640, data is sequentially read from registers according to the first instruction sequence to perform in-memory calculations, or data is sequentially read from registers according to the second instruction sequence to perform non-memory calculations.

[0093] Referring to FIG3 , this embodiment uses a matrix A of 512*256 and a matrix B of 256*512, with a minimum calculation size of 8*64, as an example for explanation, but is not limited thereto. Matrix A and matrix B are divided into at least 8*64 data blocks. The minimum data block in the first column of matrix A is shifted by 4 along the column direction, and the minimum data block in the first row of matrix B is shifted by 4 along the row direction. The intermediate results of the matrix multiplication calculation are then placed in accumulation register 170 for cumulative calculation. Correspondingly, the minimum data block in the next column of matrix A is shifted by 4 along the column direction, and the minimum data block in the next row of matrix B is shifted by 4 along the row direction. The intermediate results of the matrix multiplication calculation are then placed in accumulation register 170 for cumulative calculation. Similarly, the submatrix within the dotted box of matrix A and the submatrix within the dotted box of matrix B complete the matrix multiplication calculation to obtain submatrix block 0 of matrix C. Repeated calculations in this manner can complete the matrix multiplication operation of matrix A and matrix B.

[0094] The tensor processing method provided by the present invention uses a unified macro instruction to call different computing modules to perform in-memory or non-memory calculations, thereby reducing programming complexity and improving computing efficiency.

[0095] Furthermore, the hardware itself automatically calculates the source address and target address of multiple sub-operation data based on preset configuration information and the source address and target address of the operation data, and realizes matrix multiplication operations of larger shapes through multiple matrix multiplication operations. In this way, the hardware information can be hidden and only one macro instruction is needed to realize matrix multiplication operations of larger shapes, which greatly reduces the complexity of programming, reduces the overhead of instruction issuance and decoding, and reduces power consumption.

[0096] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0097] An embodiment of the present application provides a computer program product. When the computer program product is run on an electronic device, the electronic device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0098] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry computer program code to a device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal and a software distribution medium. For example, a USB flash drive, a mobile hard drive, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0099] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0100] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0101] In the embodiments provided in this application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely schematic. For example, the block division of the modules or units is merely a logical function block division. In actual implementation, there may be other block division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0102] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0103] While embodiments of the present invention have been described above, these embodiments do not exhaustively describe all details and do not limit the invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the above description. These embodiments are selected and described in detail in this specification in order to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better utilize the present invention and its modifications. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A tensor processing device, characterized in that: include: An instruction sequence module, configured to generate a first instruction sequence and / or a second instruction sequence according to a tensor operation instruction and preset configuration information, wherein the first instruction sequence includes a plurality of first operation instructions, the second instruction sequence includes a plurality of second operation instructions, and the preset configuration information includes first configuration information and / or second configuration information; A data loading module, used for sequentially loading the data to be operated into the corresponding register according to the first instruction sequence or the second instruction sequence; A first calculation module, used to read data from the register in sequence according to a first instruction sequence to perform in-memory calculation; The second computing module is used to read data from the register in sequence according to the second instruction sequence to perform non-memory computing.

2. The tensor processing device according to claim 1, characterized in that: The first configuration information is used to characterize the data size, shape, starting address and addressing method of each calculation by the first calculation module; the second configuration information is used to characterize the data size, shape, starting address and addressing method of each calculation by the second calculation module.

3. The tensor processing device according to claim 1, characterized in that: The instruction sequence module is also used to generate a selection signal according to the tensor operation instruction, and send the first instruction sequence to the first computing module or send the second instruction sequence to the second computing module according to the selection signal.

4. The tensor processing device according to claim 1 or 3, characterized in that: The tensor operation instruction includes an operation code and an operation field, wherein the operation code includes an instruction type and an operation code, and the operation field includes a source address and a target address of the operation data.

5. The tensor processing device according to claim 4, characterized in that: The instruction sequence module comprises: An instruction parsing unit, used for parsing the tensor operation instruction to obtain an operation code and a source address and a target address of operation data; An instruction sequence unit is used to automatically calculate the source address and target address of multiple sub-operation data according to preset configuration information and the source address and target address of the operation data, and generate multiple first operation instructions and / or multiple second operation instructions according to the operation code and the source address and target address of the multiple sub-operation data.

6. The tensor processing device according to claim 5, characterized in that: The instruction sequence unit is also used to automatically calculate the addresses of multiple first sub-operation data based on the first configuration information and the source address and destination address of the operation data to obtain multiple first operation instructions; and / or automatically calculate the addresses of multiple second sub-operation data based on the second configuration information and the source address of the operation data to obtain multiple second operation instructions.

7. The tensor processing device according to claim 5, characterized in that: The instruction sequence module also includes: The control unit is used to generate a selection signal according to a source address of the operation data, and send the first instruction sequence to the first computing module or send the second instruction sequence to the second computing module according to the selection signal.

8. The tensor processing device according to claim 5, characterized in that: The preset configuration information is set in a tensor operation instruction or a configuration register.

9. The tensor processing device according to claim 8, characterized in that: The instruction sequence module also includes: A configuration information unit is used to obtain the first configuration information or the second configuration information from the instruction parsing unit or the configuration register according to the selection signal.

10. The tensor processing device according to claim 1, characterized in that: The source address of the operation data is any one of an on-chip memory, an off-chip memory, and a register.

11. The tensor processing device according to claim 1, characterized in that: The registers include an input register, a weight register and an accumulation register, The input register is used to store activation data; the weight register is used to store weight data; and the accumulation register is used to store the intermediate result of multiplication and addition of activation data and weight data.

12. A tensor processing method, characterized in that: include: Generate a first instruction sequence and / or a second instruction sequence according to the tensor operation instruction and preset configuration information, wherein the first instruction sequence includes a plurality of first operation instructions, the second instruction sequence includes a plurality of second operation instructions, and the preset configuration information includes first configuration information and / or second configuration information; generating a selection signal according to the tensor operation instruction, and executing a first instruction sequence or a second instruction sequence according to the selection signal; Loading the data to be operated into corresponding registers in sequence according to the first instruction sequence or the second instruction sequence; Data is sequentially read from the register according to the first instruction sequence to perform in-memory calculations, or data is sequentially read from the register according to the second instruction sequence to perform non-in-memory calculations.

13. The tensor processing method according to claim 12, characterized in that: The first configuration information is used to characterize the data size, shape, starting address and addressing method of each calculation by the first calculation module; the second configuration information is used to characterize the data size, shape, starting address and addressing method of each calculation by the second calculation module.

14. The tensor processing method according to claim 12, characterized in that: The tensor operation instruction includes an operation code and an operation field, wherein the operation code includes an instruction type and an operation code, and the operation field includes a source address and a target address of the operation data.

15. The tensor processing method according to claim 14, characterized in that: Generating the first instruction sequence and / or the second instruction sequence according to the tensor operation instruction and the preset configuration information includes: Parsing the tensor operation instruction to obtain an operation code and a source address and a target address of operation data; The source addresses and target addresses of multiple sub-operation data are automatically calculated according to preset configuration information and the source address and target address of the operation data, and multiple first operation instructions and / or multiple second operation instructions are generated according to the operation code and the source address and target address of the multiple sub-operation data.

16. The tensor processing method according to claim 15, characterized in that: Also includes: A selection signal is generated according to a source address of the operation data, and the first configuration information or the second configuration information is acquired according to the selection signal.

17. The tensor processing method according to claim 16, characterized in that: The preset configuration information is set in a tensor operation instruction or a configuration register.

18. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 12 to 17 is implemented.

Citation Information

Patent Citations

  • Tensor operation method and device

    CN110941789A

  • Calculation task scheduling device, calculation device, calculation task scheduling method and calculation method

    CN116897581A

  • Tensor processing device and method and computer readable storage medium

    CN119088752A

  • Neural network weight matrix adjustment method, writing control method, and related device

    WO2021163866A1

Cited By

  • Tensor memory accelerator instruction synchronization system and method, computer equipment, readable storage medium and program product

    CN121166571A