Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

558 results about "Matrix multiplication" patented technology

In mathematics, matrix multiplication or matrix product is a binary operation that produces a matrix from two matrices with entries in a field, or, more generally, in a ring or even a semiring. The matrix product is designed for representing the composition of linear maps that are represented by matrices. Matrix multiplication is thus a basic tool of linear algebra, and as such has numerous applications in many areas of mathematics, as well as in applied mathematics, statistics, physics, economics, and engineering.

Processor and electronic equipment

The invention discloses a processor and electronic equipment. The processor includes a computing unit including a tensor core configured to perform a matrix multiplication operation using a scaling factor, and a memory, the computing unit further including a scaling factor processing module configured to determine and cache a scaling factor for each tensor associated with the matrix multiplication operation, the computing unit further comprises at least one storage module arranged on a data path between the tensor core and the memory, and the at least one storage module is exclusively occupied by the tensor core when the tensor core executes tensor related operation. And the scaling factor processing module is arranged on the at least one storage module. At present, scaling factors and floating-point number quantization are completed by a vector calculation core, so that performance is reduced, and delay becomes high, and a scaling factor processing module arranged on a storage module in a calculation unit can improve the overall execution efficiency of low-precision matrix multiplication using the scaling factors.
Owner:SHANGHAI BIREN TECH CO LTD

Parameter-efficient large-language fine-tuning federated learning framework

Provided in the present invention is a parameter-efficient large-language fine-tuning federated learning framework, comprising the following steps: performing modeling on LoRA adapters of different edge clouds; since different weights exhibit different average performances on the LoRA adapters, using singular values to quantify the importance of the weights, and therefore, before each round of independent training of the LoRA adapters using N edge clouds, using a matrix singular value to decompose a BA matrix in the LoRA adapter for each trainable weight; configuring heterogeneous LoRA adapters on the basis of the importance of the weights; and using different numbers of quantization bits to quantize a pre-trained model, and performing high-precision inverse quantization on the pre-trained model only when matrix multiplication is executed, wherein the pre-trained model is quantized to the maximum number of quantization bits on the basis of the memory budget of the edge clouds. The present invention has the following beneficial effects: the present invention determines the optimal fine-tuning model structure, thereby improving the performance of LLM fine-tuning, and adapts to heterogeneous and resource-constrained edge clouds.
Owner:FUDAN UNIVERSITY

Tensor core matrix multiplication and accumulation with hardware-based statistics collection and outlier suppression

An apparatus providing tensor core matrix multiplication and accumulation (MMA) with hardware-based statistics collection and outlier suppression is disclosed. The apparatus includes processor circuitry comprising at least one processor core comprising matrix multiplication circuitry to: execute a matrix multiplication operation on first input data from a first set of registers and on second input data from a second set of registers; collect, as part of executing the matrix multiplication operation via statistics collection hardware circuitry of the matrix multiplication circuitry, output statistics data corresponding to the matrix multiplication operation; and output the output statistics data along with a result of the matrix multiplication operation; and output statistics storage to store the output statistics data.
Owner:INTEL CORP

Mixed data precision matrix multiplication and addition unit and calculation method

The invention provides a mixed data precision matrix multiplication and addition unit and a calculation method, the matrix multiplication and addition unit comprises a calculation unit, and the calculation unit comprises a format division module, a multiplication array module, an addition tree module, an accumulator module, a normalization module and a shift register module. The calculation unit converts the first input matrix and the second input matrix into input data in a middle floating point format; executing parallel multiplication operation on the input data to generate an intermediate product result; performing index alignment and accumulation on the intermediate product result to generate an intermediate accumulated value; accumulating the intermediate product result and the value of the third input matrix in a form of accumulating an intermediate accumulated value, and outputting an accumulated result; and converting an accumulation result into a normalized result and outputting the normalized result. The format division module supports various precisions and converts data with different widths into an intermediate floating point format, so that other hardware units can be reused, and the problems that hardware resources are complex and different model reasoning scenes are difficult to meet are solved.
Owner:NANJING UNIV

Tensor core, processor, data processing method, electronic device and storage medium

The invention discloses a tensor core, a processor, a data processing method, electronic equipment and a storage medium, and is applied to the field of tensor processing. The tensor kernel comprises a first dot multiplication unit and a scaling factor matrix multiplication processing module, the scaling factor matrix multiplication processing module comprises a second dot multiplication unit, and the first dot multiplication unit and the second dot multiplication unit support dot multiplication operations of different floating-point number precisions. The tensor core is configured to receive a first tensor, a second tensor, a scaling factor of the first tensor, a scaling factor of the second tensor, an offset term of the first tensor, and an offset term of the second tensor, perform a matrix multiplication operation using the scaling factor and the offset term using a first dot multiplication unit and a scaling factor matrix multiplication processing module, and obtaining a matrix multiplication operation result of the first tensor and the second tensor. Matrix multiplication operation using scaling factors is executed by multiplexing dot multiplication units with different precisions in a tensor kernel, extra hardware area cost is reduced, and existing hardware resources are fully multiplexed.
Owner:SHANGHAI BIREN TECH CO LTD

AI compiler and compiling method based on multistage intermediate representation framework

PendingCN120560627ABiological modelsIntelligent editorsActivation functionComposite operator
The invention relates to the technical field of artificial intelligence compilers, in particular to an AI compiler and compiling method based on a multi-level intermediate representation framework, and the compiler comprises a high-level semantic retention layer which converts models of different AI frameworks into Lalg-on-Tensor IR intermediate representations; the hardware perception optimization layer comprises a tensor packaging and propagation module which is used for performing block packaging, layout propagation and folding of redundant packaging / unpackaging operation on the input tensor; the dynamic partitioning module is used for automatically selecting the partitioning size based on the cache capacity and the core number of the target hardware; the microkernel fusion module is used for fusing matrix multiplication, bias addition and an activation function into a single composite operator; and the microkernel collaboration layer is in butt joint with the hardware acceleration library through the XSMM dialect to generate a target hardware code. The hardware perception optimization layer can perform optimization according to different hardware characteristics, so that codes generated by the compiler can better adapt to target hardware, and the hardware utilization rate is improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Calculation acceleration method and device for model reasoning, medium and program product

The embodiment of the invention discloses a model reasoning calculation acceleration method and device, a medium and a program product, and the method comprises the steps: dividing the number of heads in multi-head attention processing into each processor core group according to the number of processor core groups in a processor in the multi-head attention processing of model reasoning; when the processor core group performs operation between a query vector and a key vector in multi-head attention processing, according to the total amount of the operation tasks and the number of the slave cores in the processor core group, the operation tasks are divided into the slave cores with uniform data volume; when an operation task is jointly processed by at least two slave cores, matrix segmentation is carried out on a query vector and a key vector in the operation task, and each slave core reads a corresponding vector according to a matrix segmentation result to carry out matrix multiplication operation; and caching the first matrix multiplication operation sub-result obtained by calculation of each slave core to a high-speed cache of a specified first target slave core to generate a first matrix multiplication operation result. The calculation efficiency can be improved by balancing the data volume of each slave core.
Owner:太初(无锡)电子科技有限公司

Optimization method of hybrid expert system, computer equipment, readable storage medium and program product

The invention relates to an optimization method of a hybrid expert system, computer equipment, a readable storage medium and a program product. A plurality of experts contained in the MOE are deployed in a plurality of artificial intelligence chips in groups, and the method comprises the following steps: carrying out routing calculation on an original input tensor to obtain a routing calculation result; determining input element grouping information corresponding to each expert based on expert index information and expert weight information in the routing calculation result; taking the expert dimension as a parallel dimension, executing rearrangement operation for the original input tensor in parallel based on the input element grouping information, taking a rearrangement result as input information of an expert, executing matrix multiplication and accumulation operation, and taking the expert dimension as the parallel dimension, and based on the input element grouping information, executing anti-rearrangement operation of the matrix multiplication and accumulation operation result in parallel to obtain a final operation result. By adopting the method, the MOE reasoning performance can be improved.
Owner:SHANGHAI BIREN TECH CO LTD

Attention mechanism calculation method and device, storage medium and product

The invention discloses an attention mechanism calculation method and device, a storage medium and a product, and the method comprises the steps: carrying out matrix multiplication operation through employing a query matrix block of a first register block and a key matrix block of a shared memory, obtaining a first product matrix block, and writing the first product matrix block into a second register block; performing exponential operation by using the first product matrix blocks to obtain sub-matrix blocks, writing the sub-matrix blocks into a second register block in a covering manner, and writing the sub-matrix blocks into a third register block in the form of a target precision type; performing matrix multiplication operation by using the sub-matrix blocks of the third register block and the value matrix blocks of the shared memory to obtain second product matrix blocks, and writing the second product matrix blocks into a second register block; performing softmax operation by using the second product matrix blocks to obtain attention result matrix blocks, and writing the attention result matrix blocks into a fourth register block; and writing the attention result matrix of the fourth register group into the shared memory in blocks. According to the embodiment of the invention, overflow of the register can be avoided, and the utilization of hardware resources is maximized.
Owner:SHANGHAI BIREN TECH CO LTD

Attention mechanism calculation method and device, medium and product

The invention discloses an attention mechanism calculation method and device, a medium and a product. The method comprises the steps that a second thread bundle group is controlled to load an ith query matrix block; controlling the second thread bundle group to perform matrix multiplication operation and exponential operation by using the ith query matrix block and the transposed jth key matrix block to obtain a jth attention score matrix block; and controlling the first thread bundle group and the second thread bundle group to alternately use different sub-blocks of the jth value matrix block to perform attention mechanism operation of the jth attention score matrix block until the last sub-block of the jth attention result matrix block is obtained. By adopting the embodiment of the invention, sufficient register resources can be provided for the calculation of the attention mechanism, so that the calculation efficiency is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Matrix multiplication and accumulation operation unit and operation method, hardware accelerator and electronic equipment

The embodiment of the invention provides a matrix multiplication and accumulation operation unit and method, a hardware accelerator and electronic equipment, and the matrix multiplication and accumulation operation unit comprises a data loading storage engine, a tensor register file and a matrix multiplication engine. The data loading and storage engine is used for loading data of a plurality of matrixes to be subjected to matrix multiplication and accumulation calculation; the tensor register file is used for storing data of a plurality of matrixes acquired from the data loading storage engine; the tensor register file comprises at least three tensor register groups, each tensor register group comprises a plurality of tensor registers, and different tensor register groups are used for storing data of different matrixes in the plurality of matrixes; and the matrix multiplication engine is used for carrying out matrix multiplication accumulation calculation based on the data of the plurality of matrixes stored in the tensor register file. According to the embodiment of the invention, more efficient MMA calculation is realized under the conditions of low cost, low power consumption and less occupied space.
Owner:ALIBABA (CHINA) CO LTD

Multi-precision matrix calculation unit and use method thereof

The invention discloses a multi-precision matrix calculation unit and a use method thereof, and relates to the technical field of integrated circuits, the multi-precision matrix calculation unit comprises a control logic unit, a cache module and a calculation array; wherein the control logic unit is used for configuring a computing array and controlling the cache module to read in and read out data; the cache module comprises three first cache sub-modules and one second cache sub-module; wherein the three first cache sub-modules are respectively used for caching a first to-be-processed matrix, a second to-be-processed matrix and a third to-be-processed matrix, and the second cache sub-module is used for caching a result matrix; and the calculation array is used for carrying out multiplication and addition operation on different types of to-be-processed matrixes to obtain a result matrix and writing the result matrix back to the cache module. The method can serve as a basic operation core for integrated processing of large-scale matrix multiplication and can also be integrated in processors such as RISC-V for matrix operation, universality is higher, and the method can adapt to fast and efficient calculation scenes.
Owner:SUN YAT SEN UNIV

Dynamic pressure test scene generation method and system driven by multi-modal data

The invention discloses a multi-modal data driven dynamic pressure test scene generation method and system, and the method comprises the steps: carrying out the event extraction and matching of a user operation log and a full-link API call sequence, and carrying out the feature enhancement of each alignment event pair, and obtaining a local feature vector; statistical features are extracted from each user session, a graph attention network is constructed, and a group feature matrix is obtained; obtaining a global incidence matrix according to the mapping between the service load and the resource consumption; and splicing the user session feature matrix and the group feature matrix to generate group enhancement features, performing matrix multiplication on the group enhancement features and the global incidence matrix to obtain system-level risk features, splicing local feature vectors and the encoded user behavior logs, splicing the system-level risk features and the encoded performance indexes, and obtaining the user behavior log. And finally, carrying out multi-source fusion to generate a pressure measurement scene. The simulation precision, the dynamic adaptive capacity, the abnormal reproduction capacity, the resource utilization rate and the like are remarkably improved.
Owner:HAIER CONSUMER FINANCE CO LTD

Input data sharing and cache optimization method and system in matrix multiplication calculation and application

The invention discloses an input data sharing and cache optimization method in matrix multiplication calculation. The method comprises the steps of 1, segmenting and distributing an input data matrix and a weight matrix according to the number N of calculation cores on a chip; 2, sequentially connecting the plurality of calculation cores end to end to form a data transmission annular structure; step 3, calculating the distributed matrix multiplication by each calculation core, and transmitting the current input sub-matrix of the calculation core to the next calculation core; step 4, performing matrix multiplication operation on the transmitted input sub-matrix and the weight sub-matrix in the next calculation kernel; and 5, iterating transmission and calculation of the input sub-matrixes, and carrying out N rounds of matrix multiplication of the input sub-matrixes and the weight sub-matrixes to complete the whole operation process. The invention further discloses a system for implementing the method, and the system has wide application value.
Owner:SHANGHAI QUSU CHAOWEI TECHNOLOGY CO LTD

Matrix multiplication task execution method and device, equipment, medium and program product

The invention discloses a matrix multiplication task execution method and device, equipment, a medium and a program product, the method is applied to a processor unit in a graphics processor, and the method comprises the steps that in the process that the processor unit executes a matrix multiplication task, the processor unit executes the matrix multiplication task; distributing a target register for the matrix multiplication task in the processor unit, wherein the target register comprises a shared register; generating at least two types of second thread bundles for executing the matrix multiplication task; the at least two types of second thread bundles asynchronously execute the matrix multiplication task based on the common register; and the common register is used for storing an intermediate result of the matrix multiplication task.
Owner:MOORE THREADS TECH CO LTD

Matrix multiplication operation method and device and storage medium

According to the matrix multiplication operation method and device and the storage medium, an original input matrix of a client side is packaged and encrypted through homomorphic encryption, a server side is allowed to operate in a ciphertext mode, and therefore the data privacy of the client side is protected. Afterwards, verification is carried out on the result of matrix multiplication of the server side by using homomorphism of encryption operation through homomorphic hash, and the matrix multiplication and the result of matrix multiplication form a closed loop through a polynomial packaging technology, so that the data security of the client side and the verifiability of a reasoning result are effectively ensured. Moreover, under dual verification of commitment verification and result verification, the accuracy of the target operation result is ensured.
Owner:ZHEJIANG LAB

CNN-oriented batch matrix multiplication parallel optimization method and system on SW architecture

The invention provides a CNN-oriented batch matrix multiplication parallel optimization method and system on a SW architecture, and belongs to the technical field of artificial intelligence parallel optimization. Comprising the following steps: respectively converting an input feature map and a convolution kernel in a convolution layer into an input matrix and a weight matrix, and processing the input matrix and the weight matrix into a plurality of groups of independent matrix multiplication tasks in batches; the main core encapsulates a matrix multiplication task into a parameter structure array, the parameter structure array is transmitted to the slave core through single DMA, and the slave core divides rows of an input matrix into row block tasks by adopting a dynamic row block division algorithm according to the total number of threads and the height of the matrix; and executing sub-matrix multiplication calculation on the distributed independent row blocks, asynchronously prefetching matrix sub-blocks by adopting a double-buffer DMA (Direct Memory Access), and executing matrix multiply-accumulate calculation. The parallel processing efficiency of batch matrix multiplication between the master core and the slave core of the SW processor can be improved, and the algorithm performance is optimized.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Matrix multiplication implementation method and device, electronic equipment, storage medium and program product

The invention relates to the technical field of artificial intelligence, and provides a matrix multiplication implementation method and device, electronic equipment, a storage medium and a program product.The method comprises the steps that calculation cores on a chip are divided into a plurality of calculation groups based on a plurality of matrix multiplication operations needing to be executed at the same time, each calculation group at least comprises two calculation cores, each calculation group corresponds to one matrix multiplication operation; and controlling each calculation group to execute the corresponding matrix multiplication operation, and obtaining a result matrix of each matrix multiplication operation. According to the method, the calculation cores are grouped, and each calculation group executes one matrix multiplication operation in parallel, so that the function of executing a plurality of matrix multiplication operations in parallel on the chip is realized; the number of calculation cores in each calculation group is reduced relative to the whole chip, one matrix in each matrix multiplication operation is divided into a smaller number of sub-matrixes, and the number of rows of the sub-matrixes is large, so that the number of rows of the sub-matrixes can cover the minimum calculation granularity of the calculation cores, and the utilization rate of the calculation cores is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Split weights for deep neural network inference with non-volatile memory arrays

To reduce programming noise for matrix values stored in a memory array for use in an in-array vector-matrix multiplication, such as for a neural network, the matrix is partitioned into a linear combination of matrices that will preserve the output after the combination, with the small value component matrices being normalize to the lager, full range of values before being programmed into the memory arrays. After multiplying each matrix of the combination with the vector by applying a set of bias values, the outputs are rescaled to undo the normalization before adding the individual outputs back together for the final output. This re-scaling causes the effective noise of small weights to be reduced, providing large noise tolerance for these small weight values.
Owner:SANDISK TECHNOLOGIES LLC

Method and apparatus for obtaining lower triangular matrix for matrix multiplication result value

A computer implementation method for obtaining a lower triangular matrix for a matrix multiplication result value, the method comprising: a process of allocating elements of each of a first operand matrix pair and a second operand matrix pair input into a memory to calculation units of an accelerator; and a process of the calculation units of the accelerator performing a matrix multiplication calculation on the elements of the first operand matrix pair and the second operand matrix pair, wherein the process of allocating the elements thereof to the calculation units of the accelerator comprises: allocating elements of the first operand matrix pair to first calculation units corresponding to elements of the lower triangular matrix for the matrix multiplication result value of the first operand matrix pair; and allocating elements of the second operand matrix pair to second calculation units corresponding to elements of the lower triangular matrix for the matrix multiplication result value of the second operand matrix pair.
Owner:ELECTRONICS & TELECOMM RES INST

Matrix multiplication calculation task processing system, method and equipment, storage medium, program product and chip

The invention belongs to the technical field of data processing. The invention discloses a matrix multiplication calculation task processing system, method and device, a storage medium, a program product and a chip, and the system comprises a storage module which is used for storing a preset weight matrix in advance; the linear transformation control module comprises a finite-state machine used for matrix multiplication operation scheduling and is used for generating an input request signal so that the matrix multiplication calculation task processing system can receive an input matrix; the operation module is used for executing matrix multiplication operation on the input matrix and a preset weight matrix to obtain a matrix multiplication operation result and generate an output matrix; and the data flow control module is used for controlling a plurality of groups of data flows formed by the input data in the input matrix according to rows to sequentially enter the operation module according to the input request signal until all the data in the input matrix execute and complete the matrix multiplication operation. The problems that in the prior art, a matrix multiplication accelerator lacks flexibility in a linear layer implementation task of artificial intelligence hardware, the data migration cost is high, and the energy efficiency ratio is low are solved, and the method is suitable for self-defined extension based on an RISC-V instruction set architecture, can be used as a coprocessor of an RISC-V processor, and can be used as a coprocessor of the RISC-V processor. And linear layer calculation in the neural network is accelerated, especially in an artificial intelligence task, the calculation efficiency can be remarkably improved, the power consumption can be reduced, and the real-time processing requirement can be met.
Owner:SUZHOU CHUNYA GERMINATION SEMICONDUCTOR TECHNOLOGY CO LTD

Method for calculating matrix multiplication, artificial intelligence chip, calculation device, medium and program product

The invention relates to a method for calculating matrix multiplication, an artificial intelligence chip, a calculation device, a medium and a program product. The method comprises the following steps of: accumulating a calculation result of a current cycle calculation and a calculation result of a previous cycle calculation of matrix multiplication executed in a thread bundle group granularity by utilizing a buffer which is configured in a calculation core and is used for accumulation operation; determining whether the last cycle calculation of the matrix multiplication performed at the thread bundle group granularity is completed; and in response to determining that the last cycle computation of the matrix multiplication performed at the thread bundle group granularity is completed, writing the computation results accumulated via the buffer to the register file. According to the invention, the write bandwidth of the register and the occupation of the register space can be obviously reduced.
Owner:SHANGHAI BIREN TECH CO LTD

Attention mechanism calculation method and device, storage medium and product

The invention discloses an attention mechanism calculation method and device, a storage medium and a product, and the method comprises the steps: carrying out the matrix multiplication operation of a query matrix block and a key matrix block, obtaining a first product matrix block of a first precision type, writing the first product matrix block into a second register group, and then carrying out the index operation, and obtaining a molecular matrix block; writing the quantized sub-matrix blocks into a third register block according to a second precision type; writing the sub-matrix blocks into a fourth register block according to a third precision type, and performing layout conversion through a shared memory; performing softmax operation by using the molecular matrix blocks and the value matrix blocks after layout conversion to obtain attention result matrix blocks of a first precision type, and writing the attention result matrix blocks into a fifth register block; and writing the attention result matrix blocks into the shared memory according to the third precision type. The embodiment of the invention can reduce the accumulative error of the intermediate operation result.
Owner:SHANGHAI BIREN TECH CO LTD

Method for performing matrix multiplication operation by processor comprising plurality of computing units

The invention relates to the technical field of computers, and provides a method for executing matrix multiplication by a processor comprising a plurality of computing units. In the method, a quantization parameter matrix is divided into a plurality of parameter blocks, and the parameter blocks are sequentially and respectively distributed to a plurality of calculation units in a non-repeated manner, so that at most one calculation unit is distributed with less than a first number of continuous parameter blocks in the plurality of parameter blocks, each computing unit is allocated a first number of consecutive parameter blocks in the plurality of parameter blocks; each calculation unit is used for carrying out dequantization on the distributed parameter blocks; each of the plurality of calculation units obtains a portion of the input matrix that should be multiplied by the allocated parameter block, and performs matrix multiplication on the dequantized allocated parameter block and the portion of the input matrix to complete matrix multiplication of the input matrix and the model parameter. Therefore, only one-time solution quantization needs to be carried out on the quantized parameter matrix globally, calculation resources are saved, and the reasoning efficiency is improved.
Owner:SHANGHAI INFINIGENCE AI INTELLIGENT TECHNOLOGY CO LTD

Multi-modal systolic array for matrix multiplication

A system and method for matrix multiplication using a systolic array configurable between a plurality of operating modes. A pulsation processor may receive a data type indicator for matrix multiplication. For a first data type, the systolic processor may load right-hand side data from a right-hand matrix register into data processing units of the systolic array between rows 0 and M-1, and pass respective rows of left-hand side data through corresponding rows of the systolic array between rows 0 and M-1. For the second data type, the pulsation processor may segment each element of the left-hand-side data and the right-hand-side data into a respective first element half and a respective second element half, and move each element half through a corresponding row of the pulsation array between rows 0 and 2M-1.
Owner:GOOGLE LLC

Acceleration method and accelerator for video generation model using 3D attention mechanism

The invention discloses an acceleration method and accelerator for a video generation model using a 3D attention mechanism, and the method comprises the following steps: employing a protocol speculation mode to check an important part in attention calculation, employing an FP-FP mode to calculate a matrix multiplication for the important part, and employing an FP-FP mode to calculate a matrix multiplication for the important part; for the non-important part, a matrix multiplication is calculated in an FP-INT mode; in the process of calculating the matrix multiplication by adopting an FP-INT mode, the hybrid calculation engine obtains the product of the mantissa and the shaping of the floating-point number in a table look-up mode. Through a speculation-based similarity detection method and a cache lookup table architecture, redundant operation of attention calculation is remarkably reduced, higher than 65% of high-overhead FP-FP calculation in the original attention calculation process is replaced with low-overhead FP-INT calculation, efficiency is approximately improved, energy consumption in the video generation process is reduced, and the video generation efficiency is improved. And the large-scale video generation task is more economical and efficient.
Owner:SHANGHAI JIAOTONG UNIV

Matrix multiplication optimization method and system for matrix acceleration unit

The invention discloses a matrix multiplication optimization method and system for a matrix acceleration unit, and the method comprises the steps: jointly constructing a microkernel generation frame through a matrix multiplication instruction set provided by the matrix acceleration unit by adopting a thread rearrangement and address mapping strategy and a multi-stage pipeline mechanism, and generating a candidate microkernel set; feature extraction is carried out on an input matrix, a search space formed by a dynamic block and a scheduling strategy is constructed according to the features of the input matrix, an optimal microkernel is dynamically selected, and the scheduling strategy is dynamically adjusted; and constructing a performance prediction model, performing modeling and pruning on the candidate microkernel set, and preferentially selecting an optimal scheduling strategy combination in a search space to realize optimization of matrix multiplication oriented to the matrix acceleration unit. According to the method, the automation degree and adaptability of matrix multiplication optimization are remarkably improved, and the method is a universal high-performance optimization scheme which can be applied to various matrix multiplication architectures and supports a dynamic input scene.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Acceleration processing method and device for sparse matrix vector multiplication

The invention provides an acceleration processing method and device for sparse matrix vector multiplication. The method comprises the following steps: acquiring a sparse matrix; dividing the sparse matrix into segments, and distributing threads for the segments; a matrix multiplication-accumulation instruction in a preset instruction set architecture is called, a tensor calculation core Tensor Core is used for carrying out matrix multiplication calculation of small blocks on the fragments, result data are obtained, and the result data are used for representing vector data objects of the linear equation set solution vectors. According to the method, the obtained sparse matrix is divided into the fragments and the threads are distributed, so that Tensor Core concurrent calculation is realized, the calculation efficiency is greatly improved, and the calculation time is remarkably shortened especially for solving a large-scale linear equation set.
Owner:CHINA UNIV OF PETROLEUM (BEIJING)