Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

61 results about "Sparse matrix multiplication" patented technology

Image super-resolution method and system based on semantic perception token

The invention discloses an image super-resolution method and system based on semantic perception tokens, and relates to the technical field of computer vision, and the method comprises the steps: generating semantic confidence and grouping information through the aggregation of content perception tokens, and decoupling a basic residual error into a texture enhancement and degradation inhibition guidance graph; in combination with a static semantic constraint mask and a sparse matrix multiplication mechanism, progressive focusing of attention is realized; a diffusion time step embedding and cooperative modulator is introduced, semantic guidance information is dynamically injected into a multi-step denoising process, adaptive attention features and diffusion reconstruction features are fused, and finally a high-fidelity and high-resolution image is output. According to the method, content-adaptive high-resolution image reconstruction is realized through collaborative modulation of a sparse attention mechanism guided by semantic grouping and diffusion denoising guided by semantic decoupling.
Owner:HUAQIAO UNIVERSITY

Combined MX and sparsity representation

One embodiment provides a graphics processor comprising a memory interface and a processing cluster array including a plurality of processing resources interconnected via a switched interconnect network, at least one of the plurality of processing resources including a matrix accelerator configured to execute an instruction to perform a multi-dimensional sparse matrix multiply and accumulate operation having input in a sparse microscaling format including merged sparsity and scaling metadata.
Owner:INTEL CORP

Hardware accelerator facing triple sparse matrix multiplication, equipment and application method thereof

The invention discloses a hardware accelerator and equipment oriented to triple sparse matrix multiplication and an application method thereof.The hardware accelerator comprises a high-bandwidth memory HBM, a crossbar switch network and an on-chip processing unit which are connected in sequence, and the on-chip processing unit comprises a hierarchical cache module, a global controller and a plurality of computing chips; each calculation piece comprises an RA calculation array, a TP calculation array and a local controller, wherein the RA calculation array and the TP calculation array are respectively used for executing front-end operation T = R * A and rear-end operation C = T * P in triple sparse matrix multiplication. The method aims at solving the problem that when a traditional universal processor processes triple sparse matrix multiplication, due to irregular memory access, uneven calculation load and sharp increase of middle parts and results, huge off-chip data carrying is confronted with serious performance and energy efficiency bottlenecks, and the calculation performance and energy efficiency of triple sparse matrix multiplication are improved.
Owner:NAT UNIV OF DEFENSE TECH

GPU acceleration system and method based on medium-sparseness large language model sparse matrix multiplication

The invention provides a GPU acceleration system and method based on medium sparseness large language model sparse matrix multiplication, and belongs to the technical field of computers. The system cuts a large model weight matrix through an unstructured cutting module to obtain a sparse weight matrix; a sparse matrix is converted into an MLSF format through a sparse tensor kernel layer, a slot filling layer and a residual element layer in the sparse matrix preprocessing module so as to minimize storage and decoding overhead; a sparse matrix multiplication GPU operation module is used, an MLSF format and a dense matrix are used as input, a fragmentation strategy and asynchronous loading are combined, a pipeline mechanism of data loading, decoding and calculation is constructed, and high overlapping of calculation and memory access operation is achieved. According to the method, dense matrix calculation with sparse matrix multiplication performance exceeding highly optimized performance is realized under medium sparseness, performance improvement of 5.29 times at most is obtained, and technical support is provided for efficient deployment of a large language model.
Owner:ZHEJIANG UNIV

Matrix multiplication optimization method and system based on RISC-V architecture

The invention discloses a matrix multiplication optimization method and system based on an RISC-V architecture, belongs to the technical field of machine learning, and aims to solve the technical problem of how to effectively improve the performance of sparse matrix multiplication. Comprising the following steps: constructing a sparse matrix adaptive to a machine learning model; the sparse matrix is converted into a structured sparse matrix, non-zero elements in the structured sparse matrix are stored in a vector register file, and column indexes corresponding to the non-zero elements are stored in a scalar register file; determining expansion factors of the internal circulation and the external circulation, and adjusting the expansion factors of the internal circulation; executing instructions in different loop iterations in a staggered manner in a staggered manner; designing and realizing a user-defined vector index multiply-add instruction; a custom vector index multiply-add instruction is executed.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Combining MX and sparsity representations

The name of the invention is combining MX and sparsity representation. One embodiment provides a graphics processor comprising a memory interface and a processing cluster array comprising a plurality of processing resources interconnected via a switched interconnect network, at least one of the plurality of processing resources comprising a matrix accelerator, the matrix accelerator is configured to execute instructions to perform multi-dimensional sparse matrix multiplication and accumulation operations with inputs in a sparse microscaling format that includes merged sparsity and scaling metadata.
Owner:INTEL CORP

Sparse matrix multiplication acceleration method, system and equipment based on tensor processor

The invention provides a sparse matrix multiplication acceleration method based on a tensor processor, which relates to the field of high-performance computing, and comprises the following steps: carrying out block processing on a sparse matrix to obtain sparse matrix blocks matched with the register granularity of the tensor processor; performing block processing on the dense matrix to obtain dense matrix blocks; generating a matrix multiplication task set based on the sparse matrix block and the dense matrix block; and distributing the task set to a plurality of threads, controlling the plurality of threads to respectively call the tensor processors to execute matrix multiplication operation, and aggregating calculation results of the threads to obtain an output matrix. Compared with the prior art, the sparse matrix blocks are processed into the sparse matrix blocks matched with the register granularity of the tensor processor, the index coding and decoding overhead, the reconstruction overhead and the communication overhead of the sparse matrix are reduced, and therefore the operation performance of sparse matrix multiplication based on the tensor processor is improved. The system has the same beneficial effects.
Owner:NAT UNIV OF DEFENSE TECH

AI data processing method and system based on edge computing

The invention discloses an AI data processing method and system based on edge computing, and relates to the field of edge AI processing, and the method comprises the steps: defining a weight matrix of an AI model according to an edge computing node address, inputting structured data into the AI model for multi-dimensional feature analysis, and extracting high-order data representation; analyzing the processing mode through an instruction decoding method to obtain a sparseness level, and performing parameter importance evaluation through a dynamic importance scoring method based on distribution characteristics represented by high-order data and an aging weight value to generate a deformable sparse mask matrix; performing dynamic pruning fusion on a weight matrix of the AI model through sparse matrix multiplication based on the deformable sparse mask matrix to generate a sparse data processing graph; according to the method, through node resource sensing and feature mapping, the AI model weight matrix is accurately matched with the edge computing node capability, and the operation efficiency and stability of the AI model on the edge computing node are improved.
Owner:深圳市双银科技有限公司

Large language model reasoning-oriented sparse reasoning method, system, equipment and product

The invention discloses a sparse reasoning method, system, equipment and product for large language model reasoning, and relates to the technical field of artificial intelligence. The method comprises the following steps: firstly, acquiring a weight matrix of a large language model, and then carrying out pruning sparsification processing on the weight matrix to obtain a sparse matrix; the storage format of the sparse matrix is converted into a bitmap coding storage format which is suitable for tensor calculation core perception and adopts a multi-level block structure to respectively correspond to different calculation granularities in a GPU / NPU hardware architecture so as to obtain a sparse model, and then the sparse model is deployed to respond to a reasoning request to perform reasoning service; sparse matrix multiplication is completed through a sparse matrix multiplication kernel based on on-chip storage bitmap coding so as to perform reasoning, so that through an innovative sparse matrix storage format and calculation optimization, the storage efficiency and the calculation performance in the reasoning process are remarkably improved, and especially in a low-sparseness scene, the reasoning efficiency is greatly improved. And the performance blank of the existing sparse reasoning framework in the field is filled.
Owner:HEBEI TSINGHUA DEV RES INST

Neural network training method involving sparse matrix multiplication using different mask blocks and computing device

The present disclosure provides a neural network training method and a computing device, which relates to the technical field of artificial intelligence. The scheme includes: splitting a weight matrix into В weight blocks each having the same size and satisfying N:M sparsity constraint; and then by using В sparse mask sets defined for the В weight blocks respectively and В random number indices generated for the В weight blocks respectively, generating a mask matrix for implementing sparse matrix multiplication operation. There is no need to pre-train a dense weight matrix of huge-scale in advance nor to fine-tune a trained model again for the N:M sparsity constraint, which saves a lot of training time and improves training efficiency.
Owner:HUAWEI TECH CO LTD +1

Method, apparatus, device, medium and product for accelerating multiplication of double sparse matrices

The present invention discloses a method, apparatus, device, medium and product for accelerating the multiplication of double sparse matrices. The method includes: locating a set of compressed matrices matching a first sparse matrix and a second sparse matrix in the global memory of a computing chip according to the sparse matrix multiplication requirement; sequentially transferring the set of compressed matrices and the second sparse matrix from the global memory to the hardware registers of the computing chip in the form of data blocks; and gradually calculating the multiplication result of the first sparse matrix and the second sparse matrix by a sparse computing unit of the computing chip according to the data loaded in batches in the hardware registers. The technical solution of the embodiments of the present invention can give full play to the hardware acceleration performance of the sparse computing unit in the computing chip, and thus greatly optimize the overhead of computing, bandwidth and storage resources in the process of multiplying double sparse matrices, and is particularly applicable to the model calculation scenario of a mixture-of-experts model.
Owner:SHANGHAI SUIYUAN TECH CO LTD

Matrix multiplication apparatus and method based on systolic array, and electronic device

The present application is applicable to the technical field of integrated circuits and neural networks. Provided are a matrix multiplication apparatus and method based on a systolic array, and an electronic device. The matrix multiplication apparatus based on a systolic array comprises: a systolic array formed by means of arranging processing elements, wherein two adjacent processing elements in each column are connected by a shift, each processing element has transmission channels in a row direction, a column direction and a diagonal direction, the row direction is used for inputting a row vector of a first matrix, and the column direction and the diagonal direction are used for inputting a column vector of a second matrix; a plurality of first control signal channels, which are used for inputting a first control signal for controlling the operating state of each shift; and a plurality of second control signal channels, which are used for inputting a second control signal for controlling each column vector of the second matrix to be input in the column direction or the diagonal direction. Therefore, a hardware structure taking into consideration a sparse matrix multiplication operation and a dense matrix multiplication operation is achieved, and the calculation complexity of matrix multiplication is reduced.
Owner:SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD

A Sparse Sparse Matrix Multiplication Array with High DSP Resource Utilization for Graph Neural Networks Based on FPGA

The present invention discloses a sparse matrix multiplication array with high DSP resource utilization rate for graph neural networks based on FPGA, including a preprocessing module for generating required functional configuration parameters, eigenvector matrix, and adjacency matrix according to the used GNN model and its dataset, and controlling the slice size to slice the eigenvector matrix and adjacency matrix; an eigenmatrix cache module for caching the valid values of the eigenvector matrix after slicing; an adjacency matrix cache module for caching the valid values of the adjacency matrix after slicing; a congestion mitigation array for transmitting the valid values of the adjacency matrix to the pairing array according to the congestion mitigation strategy; the pairing array obtains the valid values of the adjacency matrix transmitted by the congestion mitigation array and the valid values of the eigenvector matrix transmitted by the eigenmatrix cache module, and inputs the valid values of the adjacency matrix and the valid values of the eigenvector matrix into the multiplication array in a certain order; the multiplication array performs multiplication processing according to the received valid values of the adjacency matrix and the valid values of the eigenvector matrix, and outputs a partial aggregation result.
Owner:SUN YAT SEN UNIV

Method and device for generating calculation model for sparse matrix multiplication

The invention discloses a method and a device for generating a calculation model for sparse matrix multiplication. The method comprises the following steps: acquiring a matrix data set for training; determining a first position of an all-zero data block in a first matrix and a second position of an all-zero data block in a second matrix in the matrix data set; updating network parameters of the initial deep neural network based on the first matrix, the second matrix, the first position and the second position; and repeating the above steps until a training stopping condition is reached, obtaining the target deep neural network, and determining the target deep neural network as a calculation model for sparse matrix multiplication. Therefore, through sparse information of the first position of the all-zero data block in the first matrix and the second position of the all-zero data block in the second matrix in the matrix data set, calculation saving brought by sparsity is considered in the training process of the calculation model, so that the multiplication times required by matrix multiplication are reduced, and the calculation complexity is reduced.
Owner:TSINGHUA UNIVERSITY

Hybrid expert model reasoning optimization method and device and electronic equipment

The invention provides a hybrid expert model reasoning optimization method and device and electronic equipment. The method comprises the following steps: receiving a plurality of expert models of a hybrid expert model distributed by a terminal, and receiving operation weights of the plurality of expert models, gating weights of the hybrid expert model and feature vectors of text data sent by the terminal; based on the feature vector and the gating weight of the text data, calculating selection weights of a plurality of expert models, determining a plurality of target expert models from the plurality of expert models, and generating masks of the plurality of target expert models; based on the feature vector of the text data, the masks of the multiple target expert models and the operation weights of the multiple target expert models, sparse matrix multiplication is carried out, and multiple calculation results are obtained; and summarizing the plurality of calculation results to obtain a final calculation result. According to the hybrid expert model reasoning optimization method provided by the invention, the target expert model is determined by adopting the equipment side, and sparse matrix multiplication is performed, so that the data transmission time delay is reduced, and the calculation efficiency is improved.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Sparse matrix multiplication acceleration hardware, recommendation system acceleration method, and ai chip

The application discloses sparse matrix multiplication acceleration hardware, a recommendation system acceleration method and an AI chip. A data loading unit in the hardware establishes a data carrying task for dense data in a sparse matrix multiplication task, carries the dense data to a register, a normalization acceleration unit group sends each left and right matrix block to a multiplication calculation unit group after performing a pre-multiplication activation operation according to the left and right matrix blocks obtained from the register, the multiplication calculation unit group is used for selecting a multiplication calculation unit to perform multiplication calculation according to the data size of the left and right matrix blocks, and providing a calculation result to an addition calculation unit group to perform addition calculation, and the normalization acceleration unit group is also used for performing normalization calculation and / or activation operation according to the addition calculation result when the post-multiplication normalization and / or post-multiplication activation operation is needed, and obtaining a final result. The technical scheme of the embodiment can realize fast sparse matrix multiplication calculation under limited calculation resources.
Owner:SHANGHAI YUNSUI TECHNOLOGY CO LTD

A method and device for semi-precision sparse matrix multiplication multi-core parallel of a vector processor

The application discloses a kind of semi-precision sparse matrix multiplication multicore parallel method and device for vector processor.There are three kinds of multicore parallel modes according to the dimension of matrix and the number of computing core, suitable for a variety of computing scenarios, make full use of the multicore architecture of vector processor.At the same time, it reduces the calculation redundancy under part of matrix dimension specification, improves the parallelism of sparse matrix multiplication calculation, helps to play the computing performance of vector processor.Each multicore parallel mode is to parallel multiple computing cores in the dimension of weight matrix and dense input matrix, and to realize sparse matrix multiplication in different dimensions.The theoretical calculation efficiency of sparse matrix multiplication calculation in each multicore parallel mode is obtained based on the dimension specification of two matrices.Then the multicore parallel mode with the maximum theoretical calculation efficiency is selected for sparse matrix multiplication calculation.This can automatically adapt the optimal mode to perform calculation, with high versatility and improved calculation efficiency.
Owner:NAT UNIV OF DEFENSE TECH

Calculation method and device for sparse matrix multiplication

The invention discloses a sparse matrix multiplication calculation method and device. The method comprises the following steps: acquiring a first matrix and a second matrix which need to be subjected to matrix multiplication; determining a target dimension of the basic block, and dividing the first matrix and the second matrix based on the target dimension to obtain a corresponding first target data block and a corresponding second target data block; performing dynamic planning solution on the first target data block and the second target data block to obtain an optimal calculation strategy set; and calculating the first target data block and the second target data block based on the optimal calculation strategy set to obtain a target calculation result. According to the method, the first target data block and the second target data block divided based on the basic blocks are dynamically planned and solved to obtain the optimal calculation strategy set, so that the optimal calculation strategy set automatically adapts to matrix multiplication of different sizes and sparseness levels, and a more efficient and more flexible sparse matrix acceleration scheme is provided for LLM.
Owner:TSINGHUA UNIVERSITY

Sparse matrix multiplication acceleration hardware, recommendation system acceleration method and AI chip

The invention discloses sparse matrix multiplication acceleration hardware, a recommendation system acceleration method and an AI chip. A data loading unit in the hardware establishes a data carrying task for dense data in a sparse matrix multiplication task, and carries the dense data to a register; the normalization acceleration unit group is used for sending the left and right matrix blocks to the multiplication calculation unit group after executing activation operation before multiplication according to the left and right matrix blocks obtained from the register; the multiplication calculation unit group is used for selecting multiplication calculation units to implement multiplication calculation according to the data scales of the left and right matrix blocks, and providing calculation results to the addition calculation unit group to implement addition calculation; and the normalization acceleration unit group is also used for carrying out normalization calculation and / or activation operation according to an addition calculation result when normalization after multiplication and / or activation operation after multiplication are / is needed to be carried out, so that a final result is obtained. According to the technical scheme provided by the embodiment of the invention, rapid sparse matrix multiplication calculation can be realized under limited calculation resources.
Owner:SHANGHAI YUNSUI TECHNOLOGY CO LTD

Sparse matrix multiplication in a neural network

Apparatuses, systems, and methods to enable matrix multiplication acceleration by modifying an input to apply sparsity through sparse activation filtering. In at least one embodiment, a neural network modifies pixels within an image through sparse activation filtering to enable use of one or more matrix multiplication acceleration units to perform a sparse patch embedding operation.
Owner:NVIDIA CORP

N: M sparse matrix multiplication operator optimization method and device, equipment and storage medium

The invention relates to an N: M sparse matrix multiplication operator optimization method and device, equipment and a storage medium. The method comprises the following steps of: performing task decomposition on N: M sparse matrix multiplication by using a multi-layer partitioning algorithm based on a multi-level scheduling abstraction and storage structure of a GPU (Graphics Processing Unit); carrying out sparsity perception preprocessing on the N: M sparse matrix multiplication after task decomposition, and carrying out optimization processing on input matrixes in different sparsity scenes by using a packaging strategy and a non-packaging strategy respectively; and optimizing the N: M sparse matrix multiplication workflow by adopting prefetching, and calculating the optimized input matrix according to the optimized N: M sparse matrix multiplication workflow to obtain a final calculation result. According to the embodiment of the invention, performance optimization of N: M sparse matrix multiplication is realized, wide input scenes can be adapted, and efficient reasoning on the premise of not influencing the model effect is realized.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Sparse matrix multiplication device and method and control method

The invention provides a sparse matrix multiplication device and method and a control method, and the method comprises the steps: a mask generation module reads a first matrix and a second matrix from a memory, sets the position corresponding to each non-zero element of the first matrix in a first register to be 1, and sets the position corresponding to each zero element to be 0; setting a position corresponding to each non-zero element of a second matrix in a second register to be 1, and setting a position corresponding to each zero element to be 0; the data compression module compresses each non-zero element value of the first matrix and the second matrix and the position index of each non-zero element value; the mask analysis module performs logic AND operation on values of corresponding positions of the first register and the second register to obtain an activation signal of each multiplication module; and the multiplication module performs multiplication and addition operation on corresponding non-zero element values of the first matrix and the second matrix when the activation signal is 1, and enters a dormant state when the activation signal is 0. According to the invention, hardware power consumption of sparse matrix multiplication can be reduced.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Sparse matrix operation method, device and computing equipment

The embodiments disclosed in the present application belong to the field of computing technology, and particularly relate to a sparse matrix operation method, processor, and computing device. The processor includes a sparse matrix processing unit and a sparse vector operation unit. The operation method executed by the processor includes: the sparse matrix processing unit obtains a first numerical matrix and a first index matrix corresponding to the first sparse matrix, and a second numerical matrix and a second index matrix corresponding to the second sparse matrix. The sparse vector operation unit determines the target elements that need to be dot-producted in each row element of the first numerical matrix and each column element of the second numerical matrix based on each row element of the first index matrix and each column element included in the second index matrix, and performs a dot product operation on the target elements to obtain a first result matrix of the sparse multiplication operation performed on the first sparse matrix and the second sparse matrix. By adopting the present application, the number of zero elements involved in sparse matrix multiplication can be reduced, and the efficiency of performing sparse matrix multiplication can be improved.
Owner:HUAWEI TECH CO LTD

Sparse matrix multiplication in hardware

The present disclosure relates to sparse matrix multiplication in hardware. Methods, systems, and apparatuses, including computer-readable storage media, are provided for sparse matrix multiplication. A system for matrix multiplication includes an array of sparse tiles. Each sparse tile can be configured to receive an input submatrix and an input subvector, where the input submatrix has a number of non-zero values that is equal to or less than a predetermined maximum non-zero threshold. The sparse tile can compute, through a plurality of multiplier circuits, one or more products of a vector value multiplied by respective non-zero values of the input submatrix. The sparse tile can generate a tile output vector that is an output of the sparse tile and uses the one or more products, the tile output vector being a product of applying the tile input vector to the tile input matrix.
Owner:GOOGLE LLC

A memory-aware sparse matrix multiplication method suitable for edge embedded platforms

This application relates to a memory-aware sparse matrix multiplication method suitable for edge embedded platforms. The method includes: first, dividing the sparse matrix into row blocks and storing them in column-major order as column segments; second, dividing the dense matrix and the result matrix into column blocks, with the layout as continuous row segments. Blocking parameters are determined based on the on-chip cache capacity to ensure that each block of the result matrix can reside in the cache. Active column segments of the sparse matrix are traversed, and matching dense matrix row segments are preloaded into registers for element reuse. Non-zero elements of the column segments are multiplied and added to elements of the dense row segments, and the results are accumulated into the corresponding blocks of the result matrix. After completing all column segment operations, the result matrix format is restored, and the final multiplication result is output. This method can reduce computational latency and memory power consumption.
Owner:NAT UNIV OF DEFENSE TECH

Apparatus, method and program product for accelerating unstructured sparse matrix multiplication computation

This application discloses an apparatus, method, and program product for accelerating unstructured sparse matrix multiplication computation. The apparatus includes a preprocessing unit, a partial sum generation unit, and a merging unit. The preprocessing unit is configured to convert the unstructured sparse weight matrix to be multiplied into multiple column groups by compressing non-zero elements within the matrix rows while retaining their original column indices, and aggregating non-zero elements from different rows along the column direction. The partial sum generation unit is configured to, for each column group, perform a scalar-vector multiplication operation on each non-zero element within the group with the corresponding row data of the dense matrix to obtain the product result, and combine all product results of each column group as a partial sum data block. The merging unit is configured to accumulate all partial sum data blocks to obtain a merged result. This application significantly improves the execution efficiency of unstructured sparse SpMM in large language models.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Linear scale electronic structure calculation method and system based on dyeing superposition state and terminal

The application provides a linear scale electronic structure calculation method and system based on a dyeing superposition state and a terminal, based on the dyeing superposition state, and is used for solving sparse operator or matrix functions with spatial locality or limited correlation length, such as a density matrix, an inverse square root of an overlap matrix, and an operator required by an orthogonal representation transformation, which are involved in a Kohn-Sham self-consistent process under a local basis representation. On the premise of maintaining linear scale complexity, the application breaks through the dependence of a traditional algorithm on sparse matrix-sparse matrix multiplication (SpMSpM), and reconstructs core calculation into regular sparse matrix-dense matrix multiplication (SpMM). Through transformation of a bottom layer calculation paradigm, the regularity of Kohn-Sham electronic structure self-consistent calculation and the parallel adaptation ability of a bottom layer hardware are significantly improved, and inherent bottlenecks, such as a large pre-factor and low hardware execution efficiency, caused by the limitation of a core operator (SpMSpM) in a traditional linear scale algorithm are fundamentally relieved.
Owner:SHANGHAI TECH UNIV

A model parameter determination method, a model inference system, a decoding device, and an electronic device

The embodiment of the present disclosure provides a model parameter determination method, a model inference system, a decoding device and electronic equipment, which relates to the technical field of model inference, and comprises the following steps: determining a to-be-decoded code stream in a first weight code stream based on a weight parameter, the first weight code stream being obtained by performing prefix encoding processing on an initial weight matrix of a current network layer; and performing prefix decoding processing on the to-be-decoded code stream to obtain target weight data, the target weight data being used for sparse matrix multiplication calculation of the current network layer. The method provided by the present disclosure compresses the data transmission amount through prefix encoding processing, reduces the bandwidth requirement, and simultaneously reduces the bandwidth overhead during model network inference through the cooperative processing of prefix decoding of the to-be-decoded code stream determined by the weight parameter part, thereby improving the model processing efficiency.
Owner:BEIJING X RING TECHNOLOGY CO LTD

Sparse matrix multiplication in hardware

The present disclosure relates to sparse matrix multiplication in hardware. Methods, systems, and apparatuses, including computer-readable storage media, are provided for sparse matrix multiplication. A system for matrix multiplication includes an array of sparse tiles. Each sparse tile can be configured to receive an input submatrix and an input subvector, where the input submatrix has a number of non-zero values that is equal to or less than a predetermined maximum non-zero threshold. The sparse tile can compute, through a plurality of multiplier circuits, one or more products of a vector value multiplied by respective non-zero values of the input submatrix. The sparse tile can generate a tile output vector that is an output of the sparse tile and uses the one or more products, the tile output vector being a product of applying the tile input vector to the tile input matrix.
Owner:GOOGLE LLC

Data processing method, system and equipment for sparse matrix multiplication and storage medium

The invention provides a data processing method for sparse matrix multiplication, and the method comprises the following steps: determining a first target accumulator based on the number of non-zero elements of a processing unit of a first sparse matrix, the number of non-zero elements of a corresponding processing unit in a second sparse matrix, and a first preset selection strategy, performing analog matrix multiplication based on the first target accumulator to obtain the number of non-zero elements of a corresponding output unit of the output matrix; determining a second target accumulator based on the number of non-zero elements of a processing unit of the first sparse matrix, the number of non-zero elements of a corresponding output unit of the output matrix and a second preset selection strategy; the target accumulator is a hybrid accumulator at least comprising a merging accumulator; and using a second target accumulator to accumulate and sum the products of the non-zero elements in the processing units of the first sparse matrix and the non-zero elements in the corresponding processing units in the second sparse matrix to obtain the corresponding output units of the output matrix. According to the invention, the operation performance of sparse matrix multiplication is improved.
Owner:NAT UNIV OF DEFENSE TECH