An apparatus and method for performing matrix multiplication operations
By designing a matrix computing device containing customized hardware circuits and efficient instruction sets, the problem of the bottleneck of matrix multiplication operation performance in the prior art is solved, and efficient matrix computing performance is achieved.
Patent Information
- Application Number
- CN201911203825.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2016-04-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2036-04-26
AI Technical Summary
Existing computer technologies have performance bottlenecks when performing matrix multiplication operations, including low computing performance of a single general-purpose processor, communication problems when multiple processors are executed in parallel, and bandwidth bottlenecks caused by small on-chip cache of GPU.
A special device for performing matrix multiplication operations is designed, including a control unit, a matrix operation unit, a storage queue module and a high-speed temporary memory. The matrix multiplication operation is realized through customized hardware circuits and an efficient instruction set.
This device can efficiently support matrix operations of different scales, improve the performance of matrix multiplication operations, and avoid communication bottlenecks and cache restrictions in traditional technologies.
Smart Images

Figure CN111090467B_ABST
Abstract
Description
[0001] This disclosure is a divisional application of the following patent application: Application No. CN201610266627.0; Application Date: April 26, 2016; Invention Title: "Device and Method for Performing Matrix Multiplication Operations". Technical Field
[0002] This disclosure relates to the field of computers, and particularly to a device and method for performing matrix multiplication operations. Background Art
[0003] In the current field of computers, with the maturity of emerging technologies such as big data and machine learning, more and more tasks involve various matrix multiplication operations, especially the multiplication of large matrices, which often become the bottleneck for improving the speed and effectiveness of algorithms. Taking the currently popular deep learning as an example, it contains a large number of matrix multiplication operations. In the fully connected layer of the artificial neural network in deep learning, the operation expression of the output neuron is y = f(wx + b), where w is the weight matrix, x is the input vector, and b is the bias vector. The process of calculating the output matrix y is to multiply the matrix w by the vector x, add the vector b, and then perform an activation function operation on the obtained vector (i.e., perform an activation function operation on each element in the matrix). In this process, the complexity of the matrix-vector multiplication operation is much higher than the subsequent operations of adding the bias and performing activation. Efficiently implementing the former has the most important impact on the entire operation process. Thus, it can be seen that efficiently implementing matrix multiplication operations is an effective method for improving many computer algorithms.
[0004] In the prior art, a known solution for matrix operations is to use a general-purpose processor. This method executes general instructions through a general register file and general functional units to perform matrix multiplication operations. However, one of the disadvantages of this method is that a single general-purpose processor is mostly used for scalar calculations and has low operation performance when performing matrix operations. When multiple general-purpose processors are used to execute in parallel, the improvement effect is not significant enough when the number of processors is small; when the number of processors is large, the communication between them may become a performance bottleneck.
[0005] In another prior art, a graphics processing unit (GPU) is used to perform a series of matrix multiplication calculations. Among them, the operations are performed by using a general register file and general stream processing units to execute general SIMD instructions. However, in the above solution, the on-chip cache of the GPU is too small, and off-chip data transfer needs to be continuously performed when performing large-scale matrix operations, and the off-chip bandwidth has become the main performance bottleneck.
[0006] In another prior art, a specially customized matrix operation device is used to perform matrix multiplication operations, where a customized register file and a customized processing unit are used for matrix operations. However, according to this method, the existing dedicated matrix operation devices are limited by the design of the register file and cannot flexibly support matrix operations of different lengths.
[0007] In summary, the existing on-chip multi-core general-purpose processors, inter-chip interconnected general-purpose processors (single-core or multi-core), or inter-chip interconnected graphics processors cannot perform efficient matrix multiplication operations, and these prior arts have problems such as a large amount of code, being limited by inter-chip communication, insufficient on-chip cache, and inflexible support for matrix sizes when dealing with matrix multiplication operation problems. Summary of the Invention
[0008] Based on this, the present disclosure provides a device and method for performing matrix multiplication operations.
[0009] According to one aspect of the present disclosure, there is provided a device for performing matrix multiplication operations, including: a control unit for decoding matrix operation instructions and controlling the operation process of matrix operation instructions; a matrix operation unit connected to the control unit; the matrix operation unit is configured to receive the decoded matrix operation instructions and input matrices, and perform matrix multiplication operation operations on the input matrices according to the decoded matrix operation instructions; wherein, the matrix operation unit is a customized hardware circuit.
[0010] According to another aspect of the present disclosure, there is provided an apparatus for performing matrix multiplication operations, including: an instruction fetching module for fetching the next matrix operation instruction to be executed from an instruction sequence and transmitting the matrix operation instruction to a decoding module; a decoding module for decoding the matrix operation instruction and transmitting the decoded matrix operation instruction to an instruction queue module; an instruction queue module for temporarily storing the decoded matrix operation instruction and obtaining scalar data related to the operation of the matrix operation instruction from the matrix operation instruction or scalar registers; after obtaining the scalar data, sending the matrix operation instruction to a dependency processing unit; a scalar register file including a plurality of scalar registers for storing scalar data related to matrix operation instructions; a dependency processing unit for determining whether there is a dependency relationship between the matrix operation instruction and previously uncompleted operation instructions; if there is a dependency relationship, sending the matrix operation instruction to a storage queue module, and if there is no dependency relationship, sending the matrix operation instruction to a matrix operation unit; a storage queue module for storing matrix operation instructions having a dependency relationship with previously executed operation instructions and, after the dependency relationship is released, sending the matrix operation instruction to the matrix operation unit; a matrix operation unit for performing matrix multiplication operations on an input matrix according to the received matrix operation instruction; a cache memory for storing the input matrix and the output matrix; an input / output access module for directly accessing the cache memory and responsible for reading the output matrix from the cache memory and writing the input matrix.
[0011] The present disclosure also provides methods for performing matrix-vector multiplication and matrix-scalar multiplication.
[0012] The present disclosure can be applied to the following (including but not limited to) scenarios: various electronic products such as data processing, robots, computers, printers, scanners, telephones, tablet computers, smart terminals, mobile phones, dash cams, navigators, sensors, cameras, cloud servers, cameras, video cameras, projectors, watches, earphones, mobile storage, wearable devices, etc.; various means of transportation such as airplanes, ships, vehicles, etc.; various household appliances such as televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, range hoods, etc.; and various medical devices including nuclear magnetic resonance instruments, B-ultrasound machines, electrocardiogram machines, etc. Description of the Drawings
[0013] Figure 1 is a schematic structural diagram of an apparatus for performing matrix multiplication operations according to an embodiment of the present disclosure.
[0014] Figure 2 is an operation schematic diagram of a matrix operation unit according to an embodiment of the present disclosure.
[0015] Figure 3 is a format schematic diagram of an instruction set according to an embodiment of the present disclosure.
[0016] Figure 4 It is a schematic structural diagram of a matrix operation device according to an embodiment of the present disclosure.
[0017] Figure 5 It is a flowchart of a matrix operation device according to an embodiment of the present disclosure for executing a matrix-vector multiplication instruction.
[0018] Figure 6 It is a flowchart of a matrix operation device according to an embodiment of the present disclosure for executing a matrix-scalar multiplication instruction. Detailed implementation manners
[0019] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the following further describes the present disclosure in detail with reference to specific embodiments and the accompanying drawings.
[0020] The present disclosure provides a matrix multiplication operation device, including: a storage unit, a register unit, a control unit, and a matrix operation unit;
[0021] The storage unit stores matrices;
[0022] The register unit stores a vector address, a vector length, a matrix address, a matrix length, and scalar data required in other calculation processes;
[0023] The control unit is used to perform decoding operations on matrix operation instructions and control each module according to the matrix operation instructions to control the execution process of matrix multiplication operations;
[0024] The matrix operation unit obtains a vector address, a vector length, a matrix address, a matrix length, and scalar data required in other calculations in the instruction or the register unit. Then, it obtains the corresponding matrix from the storage unit according to the input matrix address, and then performs matrix multiplication operations based on the obtained matrix to obtain an operation result.
[0025] The present disclosure temporarily stores the matrix data participating in the calculation in a storage unit (for example, a cache memory), so that different-width data can be supported more flexibly and effectively during the matrix operation process, and the execution performance of tasks including a large number of matrix multiplication operations can be improved.
[0026] In the present disclosure, the matrix multiplication operation unit can be implemented as a customized hardware circuit (for example, including but not limited to FPGA, CGRA, application-specific integrated circuit ASIC, analog circuit, and memristor, etc.).
[0027] Figure 1 It is a schematic structural diagram of a device for performing matrix multiplication operations provided by the present disclosure, as Figure 1 shown, the device includes:
[0028] A storage unit for storing matrices. In one implementation, the storage unit can be a scratchpad memory that can support matrices of different sizes. The present disclosure temporarily stores the necessary calculation data in the scratchpad memory, enabling the computing device to more flexibly and effectively support data of different widths during matrix operations. The storage unit can be implemented by various different storage devices (such as SRAM, eDRAM, DRAM, memristors, 3D-DRAM, or non-volatile storage, etc.).
[0029] A register unit for storing matrix addresses, where the matrix address is the address at which the matrix is stored in the storage unit. In one implementation, the register unit can be a scalar register file that provides the scalar registers required during the operation. The scalar registers store the input matrix address, the input matrix length, and the output matrix address. When it comes to matrix-scalar operations, the matrix operation unit not only retrieves the matrix address from the register unit but also retrieves the corresponding scalar from the register unit.
[0030] A control unit for controlling the behavior of each module in the device. In one implementation, the control unit reads the prepared instructions, decodes them to generate control signals, and sends them to other modules in the device. The other modules perform corresponding operations according to the obtained control signals.
[0031] A matrix operation unit for obtaining various matrix multiplication operation instructions, retrieving the matrix address from the register unit according to the instructions, then retrieving the corresponding matrix from the storage unit according to the matrix address, then performing operations on the retrieved matrix to obtain a matrix operation result, and storing the matrix operation result in the storage unit. The matrix operation unit is responsible for all matrix multiplication operations of the device, including but not limited to matrix-vector operations, vector-matrix operations, matrix multiplication operations, and matrix-scalar operations. The matrix multiplication operation instructions are sent to this operation unit for execution.
[0032] Figure 2A schematic block diagram of a matrix operation unit according to an embodiment of the present disclosure is shown. As shown in the figure, the matrix multiplication operation unit consists of a main operation module and multiple slave operation modules. Among them, each slave operation module can implement the operation of vector dot product, including the pairwise multiplication of two vectors and the summation of the multiplication results. Each slave operation module consists of three parts, namely, a pairwise multiplication module, an adder tree module, and an accumulation module. The pairwise multiplication module completes the pairwise multiplication of two vectors. The adder tree module adds the result vectors of the pairwise multiplication into a number. The accumulation module accumulates the result of the adder tree on the previous partial sum. During the actual calculation of the multiplication operation of two matrices A and B, the data of matrix B is stored column by column in each slave operation module. That is, the first column is stored in slave operation module 1, the second column is stored in slave operation module 2,..., the (N + 1)-th column is stored in slave operation module 1, and so on. And the data of matrix A is stored in the main operation module. During the calculation process, the main operation module fetches a row of data of A each time and broadcasts it to all slave operation modules. Each slave operation module completes the dot product operation of the row data of A and the column data it stores, and returns the result to the main operation module. The main operation module obtains the data returned by all slave operation modules and finally gets a row in the result matrix. Specifically, in the calculation process of each slave operation module, each time a part of two vectors is fetched, the width of which is equal to the calculation bit width of the slave operation module, that is, the number of pairwise multiplications that the pairwise multiplier of the slave operation module can execute simultaneously. The result after the pairwise multiplication gets a value through the adder tree, and this value is temporarily stored in the accumulation module. When the result of the adder tree of the next segment of the vector is sent to the accumulation module, it is accumulated on this value. The operation of the vector inner product is completed through this segmented operation method.
[0033] The operation of multiplying a matrix by a scalar is completed in the main operation module, that is, there is also the same pairwise multiplier in the main operation module, and one of its inputs comes from the matrix data, and the other input comes from the vector expanded from the scalar.
[0034] The pairwise multiplier described above is multiple parallel scalar multiplication units, and these multiplication units simultaneously read the data at different positions in the vector and calculate the corresponding products.
[0035] According to an embodiment of the present disclosure, the matrix multiplication operation device further includes: an instruction cache unit for storing matrix operation instructions to be executed. During the execution of the instructions, the instructions are also cached in the instruction cache unit. When an instruction is executed, the instruction will be submitted.
[0036] According to an embodiment of the present disclosure, the control unit in the device further includes: an instruction queue module for sequentially storing the decoded matrix operation instructions, and after obtaining the scalar data required for the matrix operation instructions, sending the matrix operation instructions and the scalar data to the dependency processing module.
[0037] According to an embodiment of the present disclosure, the control unit in the device further includes: a dependency processing unit, configured to determine whether there is a dependency relationship between the operation instruction and the previous uncompleted operation instruction before the matrix operation unit obtains the instruction, such as whether the same matrix storage address is accessed. If so, the operation instruction is sent to the storage queue module, and after the previous operation instruction is executed, the operation instruction in the storage queue is provided to the matrix operation unit; otherwise, the operation instruction is directly provided to the matrix operation unit. Specifically, when a matrix operation instruction needs to access the cache memory, the previous and subsequent instructions may access the same storage space. To ensure the correctness of the instruction execution result, if the current instruction is detected to have a data dependency relationship with the previous instruction, this instruction must wait in the storage queue until the dependency relationship is eliminated.
[0038] According to an embodiment of the present disclosure, the control unit in the device further includes: a storage queue module, which includes an ordered queue. Instructions that have a data dependency relationship with the previous instruction are stored in this ordered queue until the dependency relationship is eliminated, and after the dependency relationship is eliminated, it provides the operation instruction to the matrix operation unit.
[0039] According to an embodiment of the present disclosure, the device further includes: an input / output unit, configured to store a matrix in the storage unit or obtain an operation result from the storage unit. Among them, the input / output unit can directly access the storage unit and is responsible for reading matrix data from the memory to the storage unit or writing matrix data from the storage unit to the memory.
[0040] During the process of the device performing matrix operations, the device fetches an instruction for decoding and then stores it in the instruction queue. According to the decoding result, each parameter in the instruction is obtained. These parameters can be directly written in the operation field of the instruction or read from a specified register according to the register number in the instruction operation field. The advantage of using registers to store parameters is that there is no need to change the instruction itself. As long as the value in the register is changed by the instruction, most loops can be implemented, thus greatly saving the number of instructions required to solve certain practical problems. After all the operands, the dependency processing unit will determine whether there is a dependency relationship between the data actually required by the instruction and the previous instruction, which determines whether this instruction can be immediately sent to the matrix operation unit for execution. Once a dependency relationship with the previous data is found, this instruction must wait until the instruction it depends on is executed before it can be sent to the matrix operation unit for execution. In the customized matrix operation unit, this instruction will be quickly executed, and the result, that is, the generated result matrix, will be written back to the address provided by the instruction, and this instruction is executed.
[0041] Figure 3It is a schematic diagram of the format of the matrix multiplication operation instruction provided by the present disclosure. As Figure 3 shown, the matrix multiplication operation instruction includes an opcode and at least one operand field. Among them, the opcode is used to indicate the function of the matrix operation instruction. The matrix operation unit can perform different matrix operations by identifying the opcode. The operand field is used to indicate the data information of the matrix operation instruction. Among them, the data information can be an immediate number or a register number. For example, when obtaining a matrix, according to the register number, the starting address and length of the matrix can be obtained in the corresponding register, and then the matrix stored at the corresponding address can be obtained in the storage unit according to the starting address and length of the matrix.
[0042] There are the following several matrix multiplication operation instructions:
[0043] Matrix multiply vector instruction (MMV). According to this instruction, the device fetches matrix data and vector data of a specified size from a specified address in the cache memory, performs matrix-vector multiplication in the matrix operation unit, and writes the calculation result back to the specified address in the cache memory. It should be noted that a vector can be stored in the cache memory as a special form of matrix (a matrix with only one row of elements).
[0044] Vector multiply matrix instruction (VMM). According to this instruction, the device fetches vector data and matrix data of a specified length from a specified address in the cache memory, performs vector-matrix multiplication in the matrix operation unit, and writes the calculation result back to the specified address in the cache memory. It should be noted that a vector can be stored in the cache memory as a special form of matrix (a matrix with only one row of elements).
[0045] Matrix multiplication instruction (MM). According to this instruction, the device fetches matrix data of a specified size from a specified address in the cache memory, performs matrix multiplication in the matrix operation unit, and writes the calculation result back to the specified address in the cache memory.
[0046] Matrix multiply scalar instruction (MMS). According to this instruction, the device fetches matrix data of a specified size from a specified address in the cache memory, fetches scalar data from a specified address in the scalar register file, performs matrix-scalar multiplication in the matrix operation unit, and writes the calculation result back to the specified address in the cache memory. It should be noted that the scalar register file stores not only the address of the matrix but also scalar data.
[0047] Figure 4 It is a schematic diagram of the structure of the matrix operation device provided by an embodiment of the present disclosure. As Figure 4As shown, the device includes an instruction fetch module, a decoding module, an instruction queue module, a scalar register file, a dependency handling unit, a store queue module, a matrix operation unit, a cache, and an I / O memory access module;
[0048] The instruction fetch module is responsible for fetching the next instruction to be executed from the instruction sequence and passing the instruction to the decoding module;
[0049] The decoding module is responsible for decoding the instruction and passing the decoded instruction to the instruction queue;
[0050] The instruction queue is used to temporarily store the decoded matrix operation instructions and obtain scalar data related to the matrix operation instructions from the matrix operation instructions or scalar registers; after obtaining the scalar data, the matrix operation instructions are sent to the dependency handling unit;
[0051] The scalar register file provides the scalar registers required during the operation of the device; the scalar register file includes multiple scalar registers for storing scalar data related to matrix operation instructions;
[0052] The dependency handling unit processes the possible memory dependencies between the instruction and the previous instruction. Matrix operation instructions access the cache, and consecutive instructions may access the same memory space. That is, this unit detects whether there is an overlap between the memory range of the input data of the current instruction and the memory range of the output data of the previous uncompleted instruction. If there is an overlap, it means that this instruction logically needs to use the calculation result of the previous instruction. Therefore, it must wait until the dependent instructions before it are executed before it can start execution. During this process, the instruction is actually temporarily stored in the following store queue. To ensure the correctness of the instruction execution result, if the current instruction is detected to have a data dependency with the previous instruction, the instruction must wait in the store queue until the dependency is eliminated.
[0053] The store queue module is an ordered queue, and instructions that have a data dependency with the previous instruction are stored in this queue until the storage relationship is eliminated;
[0054] The matrix operation unit is responsible for performing matrix multiplication operations;
[0055] The cache is a dedicated cache storage device for matrix data and can support matrix data of different sizes; it is mainly used to store input matrix data and output matrix data;
[0056] The I / O memory access module is used to directly access the cache and is responsible for reading data from or writing data to the cache.
[0057] The matrix operation device provided by the present disclosure temporarily stores the matrix data participating in the calculation in a storage unit (for example, a scratchpad memory), so that different-width data can be supported more flexibly and effectively during the matrix operation process. At the same time, the customized matrix operation module can implement various matrix multiplication operations more efficiently, improving the execution performance of tasks involving a large number of matrix calculations. The instruction set adopted by the present disclosure is easy to use and supports flexible matrix lengths.
[0058] Figure 5 is a flowchart of the operation device provided by an embodiment of the present disclosure for executing a matrix multiply vector instruction. As Figure 5 shown, the process of executing the matrix multiply vector instruction includes:
[0059] S1, The instruction fetch module fetches this matrix multiply vector instruction and sends the instruction to the decoding module.
[0060] S2, The decoding module decodes the instruction and sends the instruction to the instruction queue.
[0061] S3, In the instruction queue, this matrix multiply vector instruction needs to obtain the data in the scalar registers corresponding to the five operation fields in the instruction from the scalar register file, including the input vector address, input vector length, input matrix address, output vector address, and output vector length.
[0062] S4, After obtaining the required scalar data, the instruction is sent to the dependency processing unit. The dependency processing unit analyzes whether there is a data dependency between this instruction and the previously unexecuted instructions. This instruction needs to wait in the storage queue until there is no longer a data dependency with the previously unexecuted instructions.
[0063] S5, After the dependency does not exist, this matrix multiply vector instruction is sent to the matrix operation unit.
[0064] S6, The matrix operation unit fetches the required matrix and vector data from the cache according to the address and length of the required data, and then completes the matrix multiply vector operation in the matrix operation unit.
[0065] S7, After the operation is completed, the result is written back to the specified address in the scratchpad memory.
[0066] Figure 6 is a flowchart of the operation device provided by an embodiment of the present disclosure for executing a matrix multiply scalar instruction. As Figure 6 shown, the process of executing the matrix multiply scalar instruction includes:
[0067] S1’, The instruction fetch module fetches this matrix multiply scalar instruction and sends the instruction to the decoding module.
[0068] S2’, the decoding module decodes the instruction and sends the instruction to the instruction queue.
[0069] S3’, in the instruction queue, this matrix multiply scalar instruction needs to obtain the data in the scalar registers corresponding to the four operation fields in the instruction from the scalar register file, including the input matrix address, the input matrix size, the input scalar, and the output matrix address.
[0070] S4’, after obtaining the required scalar data, the instruction is sent to the dependency processing unit. The dependency processing unit analyzes whether there is a data dependency between this instruction and the previously unexecuted instructions. This instruction needs to wait in the storage queue until there is no longer a data dependency between it and the previously unexecuted instructions.
[0071] S5’, after the dependency does not exist, this matrix multiply scalar instruction is sent to the matrix operation unit.
[0072] S6’, the matrix operation unit fetches the required matrix data from the high-speed cache according to the address and length of the required data, and then completes the matrix multiply scalar operation in the matrix operation unit.
[0073] S7’, after the operation is completed, the result matrix is written back to the specified address in the high-speed cache memory.
[0074] In summary, the present disclosure provides a matrix operation device, which, in cooperation with corresponding instructions, can well solve the problem that more and more algorithms in the current computer field involve a large number of matrix multiplication operations. Compared with the existing traditional solutions, the present disclosure can have the advantages of convenient use, flexible supported matrix scale, sufficient on-chip cache, etc. The present disclosure can be used for various computing tasks involving a large number of matrix multiplication operations, including the backward training and forward prediction of the artificial neural network algorithm, which currently performs very well, and traditional numerical calculation methods such as the power method for solving the largest eigenvalue of an irreducible matrix.
[0075] Each functional unit / module / sub-module in the present disclosure can be hardware. For example, the hardware can be a circuit, including digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes but is not limited to physical devices, and physical devices include but are not limited to transistors, memristors, etc. The computing module in the computing device can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. The storage unit can be any suitable magnetic storage medium or magneto-optical storage medium, such as RRAM, DRAM, SRAM, EDRAM, HBM, HMC, etc.
[0076] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0077] The specific embodiments described above have further elaborated on the objectives, technical solutions, and beneficial effects of the present disclosure. It should be understood that the above are only specific embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A matrix operation unit, characterized in that, it is applied in an application specific integrated circuit (ASIC) to identify matrix operation instructions for multiplication operations. The matrix operation unit includes: a main operation module; and slave operation modules; wherein, when the matrix operation instruction is matrix multiplication by matrix, matrix multiplication by vector or vector multiplication by matrix, the main operation module cooperates with the slave operation modules to perform bitwise multiplication to complete the multiplication operation; when the matrix operation instruction is matrix multiplication by scalar, the main operation module performs bitwise multiplication to complete the multiplication operation; wherein, the main operation module stores data of a first matrix, data of a second matrix is stored column by column in each slave operation module, and each slave operation module includes a bitwise multiplication module and an addition tree module. During the calculation process of each slave operation module: the bitwise multiplication module is used to complete the bitwise multiplication operation of a row of data in the first matrix and the column data stored by itself to obtain the result after bitwise multiplication; the addition tree module is used to add up the results after bitwise multiplication into one number.
2. The matrix operation unit according to claim 1, wherein the main operation module stores the first matrix, and the slave operation modules include: a first slave operation module that stores the first column data of multiple columns of data of the second matrix; and a second slave operation module that stores the second column data of multiple columns of data of the second matrix; wherein the first matrix performs bitwise multiplication with the first column data and the second column data respectively.
3. The matrix operation unit according to claim 2, the first matrix includes first row data and second row data, wherein: the main operation module broadcasts the first row data to the first slave operation module and the second slave operation module, the first slave operation module performs a dot product operation on the first row data and the first column data, and the second slave operation module performs a dot product operation on the first row data and the second column data; the main operation module broadcasts the second row data to the first slave operation module and the second slave operation module, the first slave operation module performs a dot product operation on the second row data and the first column data, and the second slave operation module performs a dot product operation on the second row data and the second column data.
4. The matrix operation unit according to claim 3, wherein the main operation module obtains a result matrix according to the dot product operation results returned by the first slave operation module and the second slave operation module.
5. The matrix operation unit according to claim 1, wherein the slave operation module further includes: an accumulation module for accumulating the number of the addition tree module on the previous partial sum to complete the operation of vector inner product based on the method of segmented operation.
6. The matrix operation unit according to claim 1, wherein the bitwise multiplication module takes out a specific bit width for calculation each time, and the specific bit width is equal to the calculation bit width of the bitwise multiplication module.
7. The matrix operation unit according to claim 1, wherein the main operation module includes a bitwise multiplier. When the matrix operation instruction indicates matrix multiplication by a scalar, the matrix operation unit expands the scalar into a vector, and the bitwise multiplier receives the vector for operation.
8. A matrix operation unit characterized in that it is applied to an application-specific integrated circuit (ASIC) and includes: a main operation module; and slave operation modules, including a first slave operation module and a second slave operation module. The first slave operation module and the second slave operation module respectively include: a bitwise multiplication module for performing bitwise multiplication to generate bitwise multiplication result data; and an adder tree module for adding the bitwise multiplication result data into one number; wherein the vectors stored in the main operation module are respectively bitwise multiplied with the vectors stored in the first slave operation module and the second slave operation module.
9. The matrix operation unit according to claim 8, wherein the first slave operation module and the second slave operation module further include an accumulation module for accumulating the number from the adder tree module on a previous partial sum to complete the operation of the inner product of vectors based on a segmented operation method.
10. The matrix operation unit according to claim 8, wherein the bitwise multiplication module extracts a specific bit width for calculation each time, and the specific bit width is equal to the calculation bit width of the bitwise multiplication module.
11. The matrix operation unit according to claim 8, wherein the main operation module stores a first matrix, and the first matrix includes first row data and second row data; the first slave operation module stores the first column data of multiple columns of a second matrix, and the second slave operation module stores the second column data of multiple columns of the second matrix; wherein: the main operation module broadcasts the first row data to the first slave operation module and the second slave operation module. The first slave operation module performs a dot product operation on the first row data and the first column data, and the second slave operation module performs a dot product operation on the first row data and the second column data; the main operation module broadcasts the second row data to the first slave operation module and the second slave operation module. The first slave operation module performs a dot product operation on the second row data and the first column data, and the second slave operation module performs a dot product operation on the second row data and the second column data.
12. The matrix operation unit according to claim 11, wherein the main operation module obtains a result matrix based on the dot product operation results returned by the first slave operation module and the second slave operation module.
13. The matrix operation unit according to claim 8, wherein the main operation module includes a bitwise multiplier. When the matrix operation is matrix multiplication by a scalar, the matrix operation unit expands the scalar into a vector, and the bitwise multiplier receives the vector for operation.
Citation Information
Patent Citations
Parallel vector processing system for individual and broadcast distribution of operands and control information
US5226171A