Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

44 results about "General matrix" patented technology

Heterogeneous computing thread block optimal scheduling method and system based on dynamic topology mapping

The invention belongs to the field of parallel computing architecture optimization, and relates to a matrix multiplication acceleration method and system based on dynamic computing resource mapping, and the method comprises the steps: constructing a dynamic topology model driven by tensor dimension features, and generating a thread block distribution mode according to matrix parameters and GPU hardware information; constructing a multi-dimensional resource scheduling strategy library, dynamically selecting an optimal thread block distribution strategy from the multi-dimensional resource scheduling strategy library, and generating a binding relationship between the thread blocks and the data blocks; calculating collaborative access logic of thread blocks and storage hierarchies based on block parameters and dynamic mapping function optimization; distributed calculation is carried out, calculation and data transmission are parallelized through pipelining and a double-buffering mechanism, and result aggregation across calculation units is completed synchronously through atomic operation and a barrier. According to the method, discontinuous memory access conflicts can be effectively reduced, the execution efficiency of the calculation instruction and the utilization rate of the cache space are improved, the parallel calculation process of accelerating and optimizing the general matrix multiplication is realized, and the data processing efficiency is improved.
Owner:SOUTH CHINA UNIV OF TECH

Information processing method and device, electronic equipment and storage medium

The invention discloses an information processing method and device, electronic equipment and a storage medium, and is applied to the field of information processing. The information processing method comprises the steps that a first utilization rate is determined in combination with a data flow between a memory and a shared memory when a processor executes a general matrix multiplication operator, and the first utilization rate is used for indicating performance evaluation of the general matrix multiplication operator at a processor level; determining a second utilization rate in combination with a data stream between a shared memory and a register file in a single computing unit when the processor executes the general matrix multiplication operator, the second utilization rate being used for indicating performance evaluation of the general matrix multiplication operator at the computing unit level; and based on the smaller one of the first utilization rate and the second utilization rate, determining the theoretical performance of the processor when executing the universal matrix multiplication operator. At present, only chip-level performance evaluation is considered in performance evaluation of a GEMM operator, and double-level-dimension theoretical performance evaluation is provided, so that the evaluation result is more accurate.
Owner:SHANGHAI BIREN TECH CO LTD

Memory architecture-oriented dual-precision general matrix multiplication optimization method and system

The invention belongs to the related technical field of high-performance computing, and provides a memory architecture-oriented dual-precision general matrix multiplication optimization method and system in order to solve the problems of limited computing power and access efficiency and the like in the prior art. Decomposing the matrix into a plurality of sub-matrix blocks according to the slave core array topology; the slave core receives the sub-matrix blocks issued by the master core, divides the sub-matrix blocks into small sub-matrix blocks based on a uniform blocking rule, loads the small sub-matrix blocks to an independent buffer area of a local data memory based on a DMA double-buffer protocol, divides the small sub-matrix blocks in the buffer area into SIMD vectors according to the SIMD unit characteristics of the slave core, and sends the SIMD vectors to the slave core; vectorization calculation and caching operation are alternately switched according to an iteration period through different independent buffer areas; and after all the slave cores finish calculation, the master core collects results written back to the master memory by the slave cores to obtain a final operation result, and double breakthrough of calculation power and memory access efficiency is realized.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Large language model weight inverse quantization reasoning device and method

The invention relates to the technical field of large language model deployment, and discloses a large language model weight inverse quantization reasoning device and method.The method comprises the steps that low-precision weight data is transmitted to a high-bandwidth storage from a host and then transmitted to an on-chip storage through the high-bandwidth storage; data conversion from a low-precision format to a high-precision format is completed in the on-chip memory, the data is multiplied by an inverse quantization factor to obtain recovered high-precision weight data, and the functional unit is responsible for executing general matrix multiplication of input data and the high-precision weight data after inverse quantization. And pipeline parallel execution of the inverse quantization operation and the general matrix multiplication operation is realized through a double-buffering technology. According to the invention, on the basis of a dual-path inverse quantization architecture of the vector processing unit and a dual-buffer mechanism in the on-chip memory, the problem of hardware adaptation of low-precision calculation is solved, and efficient execution of a low-precision conversion algorithm is realized under the condition of limited hardware resources.
Owner:JIANGNAN UNIV +1

Method and device for calculating matrix multiplied by vector, computing equipment and storage medium

The embodiment of the invention provides a method and device for calculating a matrix multiplied by a vector, computing equipment and a storage medium, the matrix is a matrix of M * K, the vector is a vector of K * 1, and M and K are positive integers. The method comprises the following steps: converting an M * K matrix into a first tensor of M * (K / N) * N, N being a positive integer, Ngt; 1 and K are multiples of N; converting the vector of K * 1 into a second tensor of N * (K / N); using a tensor calculation kernel to carry out general matrix multiplication calculation of a first tensor of M * (K / N) * N and a second tensor of N * (K / N) in batches to obtain M result matrices; elements on diagonals of each of the M result matrices are added to result in a matrix of M * K multiplied by each element on a result vector of M * 1 of the vector of K * 1. According to the scheme, high-throughput and low-delay high-speed matrix vector multiplication (MMV) operation is realized by utilizing the tensor calculation kernel.
Owner:SHANGHAI BIREN TECH CO LTD

Model reasoning acceleration method and device, electronic equipment and nonvolatile storage medium

The invention discloses a model reasoning acceleration method and device, electronic equipment and a nonvolatile storage medium. The method comprises the steps that in the process that a natural language processing model carries out general matrix calculation, a first matrix and a second matrix are divided into a plurality of blocks respectively, general matrix calculation is executed through an image processing unit, and the first matrix and the second matrix are matrixes needing general matrix calculation; loading the blocks from a global memory of the image processing unit to a shared memory of the image processing unit, and performing matrix operation by adopting threads corresponding to the blocks in the shared memory; and writing a result obtained by calculation of each thread back to a corresponding position in the global memory to obtain a matrix calculation result corresponding to matrix calculation of the first matrix and the second matrix. The technical problem of low model reasoning calculation efficiency caused by high memory access delay and low data reuse rate of a global memory of an image processing unit during reasoning of a large model in related technologies is solved.
Owner:CHINA TELECOM CORP LTD

Communication calculation parallel optimization method, multiprocessor system, medium and program product

The invention discloses a communication computing parallel optimization method, a multiprocessor system, a medium and a program product, and the method comprises the steps: configuring a first computing core and a second computing core on a single task flow for a general matrix multiplication subtask allocated to a single processor; wherein the first calculation core is used for executing a general matrix multiplication subtask, and the second calculation core is used for executing a set communication task; the set communication task comprises a full accumulation operator or a protocol dispersion operator; then, starting scheduling is conducted on the first calculation core and the second calculation core according to a preset dependency mechanism, so that the general matrix multiplication subtask and the set communication task are executed asynchronously in an overlapped mode; according to the method, the problem of poor reusability of the original Kernel caused by intrusive modification of the GEMM or re-implementation of the Kernel can be effectively avoided through parallel optimization of GEMM calculation and ensemble communication operation, and the performance overhead of the processor is reduced.
Owner:SHANGHAI BIREN TECH CO LTD

Kernel selection method and device during general matrix multiplication operation, equipment and storage medium

PendingCN122044838AResource allocationBiological modelsGeneral matrixAlgorithm
The embodiment of the invention provides a kernel selection method and device during general matrix multiplication operation, equipment and a storage medium, and belongs to the technical field of computers. The method comprises the following steps: acquiring general matrix multiplication problem size data; preprocessing the general matrix multiplication problem size data to obtain general matrix multiplication problem size features; based on a pre-trained problem size encoder, mapping the general matrix multiplication problem size feature into a problem size embedded vector of a preset dimension; taking the problem size embedded vector as a query vector, and performing nearest neighbor vector search in a pre-configured vector database to obtain a kernel configuration feature with the highest similarity; and outputting the kernel configuration feature as a selection result. The method is used for improving the kernel selection efficiency and precision during the operation of the general matrix multiplication.
Owner:DAWNING INT INFORMATION IND CO LTD +1

Local perception lookup table lookup method

The invention relates to the field of data lookup, and discloses a locality perception lookup table lookup method, which comprises the following steps of: S1, reordering between rows: clustering and reordering the rows of an input vector and a weight matrix according to the same quantized value to enable elements with the same quantized value to be continuously distributed; for a common matrix GeMV scene, cache locality is improved through rearrangement and subsequent blocking operation, and repeated loading of a lookup table LUT is reduced. And the data locality of the GeMV scene is remarkably improved through inter-line reordering and a virtual mapping strategy. On the basis of reordering of input vector element values, elements with the same value are aggregated, so that the LUT line loading times are reduced to 256 times at most from being in direct proportion to the input dimension d, and frequent replacement of the LUT lines in the WRAM is avoided; meanwhile, a weight matrix is divided into BR * BC sub-matrixes, the characteristic that vertical adjacent rows share an accumulator is utilized, multi-row sub-matrix calculation is loaded at a time, and virtual rearrangement maintains the row matching relation between an input vector and the weight matrix through index mapping.
Owner:RENMIN UNIVERSITY OF CHINA

Determination method of rock mass fracture network seepage field containing complex fracture morphology

The invention discloses a calculation method of a rock mass fracture network seepage field containing a complex fracture form, and relates to the field of geotechnical engineering. The method comprises the following steps: for each node in a rock mass fracture network, determining a source-sink item at the node based on the flow of a plurality of fracture units connected with the node; substituting the source sink item at the node and the equivalent permeability coefficient matrix into an overall matrix equation of fracture network seepage for solving to obtain a pressure head value of the node; the overall matrix equation is obtained by assembling unit seepage matrix equations of all nodes in the rock mass fracture network; in the construction process of the fracture seepage unit equation, the influence of the on-way resistance coefficient on the equivalent seepage flow velocity of the fracture in the complex wall surface form and the influence of the local resistance coefficient on the equivalent seepage flow velocity of the fracture in the complex wall surface form are considered; and determining the pressure head values of all nodes in the rock mass fracture network as a rock mass fracture network seepage field. The method can improve the calculation precision of the fracture network seepage field.
Owner:JIANGSU OCEAN UNIV

Matrix multiplication operators, operational methods, devices, graphics processors, and storage media

PendingCN122132654AComplex mathematical operationsGraphicsGeneral matrix
This invention provides a method, apparatus, graphics processor, and storage medium for matrix multiplication operators, relating to the field of artificial intelligence technology. The method includes: reading two operation matrices from high-bandwidth memory; determining a first target input matrix in the general matrix buffer (GMB) based on the two operation matrices, and identifying the other operation matrix as a second target input matrix; the first target input matrix is ​​arranged in a transposed manner; and writing the first and second target input matrices into a tensor core for matrix multiplication to obtain the target output matrix from the tensor core. This invention, by identifying the data arrangement of the two operation matrices, prioritizes inputting the first target input matrix (with a transposed arrangement) into the GMB, effectively avoiding hardware performance limitations associated with non-transposed readings, ensuring the tensor core maintains full-bandwidth data input, and significantly improving the computational efficiency of the matrix multiplication operator.
Owner:广州壁仞智能科技有限公司 +1

A data processing method, apparatus, electronic device, and storage medium

ActiveCN117785031BInput/output to record carriersComputer hardwareGeneral matrix
The present disclosure relates to a data processing method, apparatus, electronic device, and storage medium. The method comprises: in response to a data processing instruction for a target input feature map, obtaining a target input data block from an input data block arrangement corresponding to the target input feature map in a memory; the input data block arrangement is based on the data throughput of a general matrix processing engine in each clock cycle, and the target input feature map is divided into multiple input data blocks and stored in the memory according to a preset arrangement; the target input data blocks are cached; and based on the cached target input data blocks, matrix multiplication operations are performed using the general matrix processing engine to obtain data processing results. The embodiments of the present disclosure improve the data processing efficiency based on Img2Col, thereby improving the data processing efficiency of electronic devices based on deep learning networks.
Owner:北京凌川科技有限公司

An address mapping method and system for GPU convolution acceleration

The application relates to the field of computer architecture and parallel computing technology, and discloses an address mapping method and system for GPU convolution acceleration, which is applied to the process of performing convolution operation based on general matrix multiplication by a GPU, dynamically remaps the memory access address of a workspace matrix, and comprises the following steps: identifying a data repetition mode in the workspace matrix; based on the data repetition mode, remapping the original address of a current memory access request to a target address by using an address mapping strategy; performing a data access operation by using the target address; if the target data is cached in a cache level of the GPU, directly returning the same value in the original access address; otherwise, loading data from a global memory and storage. The application avoids frequent global memory and storage access caused by scattered storage of repeated data in a traditional method, significantly improves the cache hit rate, and further breaks through the bottleneck of low cache utilization of an existing strategy.
Owner:SHANDONG UNIV

Computing method and apparatus for general matrix multiplication

A computing method and apparatus for general matrix multiplication (GEMM) are provided. The method comprises: in response to determining that a host needs to obtain a first calculation result of multiplication of first weight data and an input vector, performing a multiplication operation of second weight data and the input vector to obtain a second calculation result, wherein the first weight data satisfies a first arrangement rule for the host, and the second weight data is obtained by arranging the first weight data based on a second arrangement rule for a PIM apparatus different from the first arrangement rule; and obtaining the first calculation result based on the second calculation result.
Owner:SAMSUNG (CHINA) SEMICONDUCTOR CO LTD +1

Methods, apparatus, computing devices, and storage media for computing a matrix multiplication of vectors

Embodiments of the present disclosure provide a method, apparatus, computing device and storage medium for calculating matrix-vector multiplication, where the matrix is an M*K matrix and the vector is a K*1 vector, M and K are positive integers. The method comprises: converting the M*K matrix into an M*(K / N)*N first tensor, where N is a positive integer, N>1, and K is a multiple of N; converting the K*1 vector into an N*(K / N) second tensor; using a tensor computing core to batch-process the general matrix multiplication of the M*(K / N)*N first tensor and the N*(K / N) second tensor to obtain M result matrices; and adding the elements on the diagonal of each of the M result matrices to obtain each element on the M*1 result vector of the M*K matrix multiplied by the K*1 vector. The above scheme uses a tensor computing core to implement high-throughput, low-latency high-speed matrix-vector multiplication (MMV) operation.
Owner:SHANGHAI BIREN TECH CO LTD

Systems and methods for activation sparse and kernel sparse general matrix multiplication in neural networks

ActiveUS12670370B1General matrixAlgorithm
A system and method for performing multiplication for a neural network, e.g. for data of one or more layers in a neural network, may include loading a portion of a compressed version of a sparse input matrix into a cache memory; uncompressing a subset of the data in the portion of the compressed version of the sparse input matrix; and multiplying a sparse kernel matrix by the subset of the data using a set of instructions which are themselves created based on the sparse kernel matrix.
Owner:RED HAT INC

Method for accelerating a general matrix multiplication on a computer system

PCT designated stageWO2026008165A1Biological modelsMachine learningGeneral matrixTheoretical computer science
The present invention relates to a computer-implemented method, computer program, non-transitory, computer-readable medium comprising a program code and computer system for accelerating a General Matrix Multiplication (GEMM) on a computer system, and to methods, systems and computer programs / computer-readable media using such a computer-implemented method. A computer-implemented method for accelerating a General Matrix Multiplication, GEMM, on a computer system, comprises obtaining (310) a data structure comprising a representation of a first set of matrix layout variants known to be more efficient than a second set of matrix layout variants when being used to perform, by the computer system, a GEMM, with each matrix layout variant defining whether at least one of the matrices involved in the GEMM are to be transposed, and selecting (350) a matrix layout variant of the first set of matrix layout variants based on matrices involved in a GEMM to be performed, and performing (370) the GEMM on the matrices using the selected matrix layout variant.
Owner:NEC LAB EURO GMBH

Multi-low-rank adaptive fine tuning model reasoning optimization method and device

The invention provides a multi-low-rank adaptive fine tuning model reasoning optimization method and device. The method comprises a step of constructing a fusion input matrix and obtaining basic output by using a pre-training model, a step of carrying out fusion and feature extraction on input side low-rank matrixes of a plurality of LoRA adapters, and a step of carrying out fusion on output side low-rank moments of the plurality of LoRA adapters to obtain a correction term. The device comprises a corresponding processing module, namely, according to the method and the device, a plurality of small matrix multiplications of a plurality of LoRA fine tuning models are fused into a single general matrix multiplication (GEMM) operation, so that the calculation intensity is improved, the memory addressing operation is reduced, and the calculation efficiency is improved. The utilization rate of the GPU computing unit can be remarkably improved, so that the efficiency of multi-LoRA fine-tuning model reasoning is improved, and the multi-LoRA fine-tuning model reasoning can have better real-time performance and expandability.
Owner:FUDAN UNIVERSITY +1

A method for accelerating arbitrary precision sparse matrix multiplication based on tensor cores

The application discloses a method for accelerating arbitrary precision sparse matrix multiplication and addition operation based on a tensor core, which is in combination with deep learning and the program characteristics of parallel processing of sparse matrices, and is in combination of an arbitrary precision sparse matrix multiplication and addition operation operator and a parallel computing acceleration device, decomposes a general matrix into a sparse matrix of a set format, blocks the sparse matrix and specifies a corresponding parallel computing hardware unit for the sparse matrix, uses shared memory to reduce memory conflict and access delay, calls a high-level language API function library to execute the loading, calculation and storage process of the calculation data, and finally reduces the output data of the function, so that the final result of the arbitrary precision sparse matrix multiplication and addition operation on the parallel computing acceleration device is obtained. The application can fully utilize the parallel computing resources such as the tensor core on the A100, effectively reduce the program calculation amount and communication overhead, and greatly improve the program running performance.
Owner:XI AN JIAOTONG UNIV

A code optimization method and optimization device

ActiveCN119690442BIntelligent editorsRequirement analysisGeneral matrixAlgorithm
A code optimization method and an optimization device are provided. The method is applied to the optimization device, and includes: determining M first general matrix vector multiplication (GEMV) calculations in a code to be optimized based on an identification rule; determining N second GEMV calculations in the M first GEMV calculations based on a merging rule; inserting general matrix multiplication (GEMM) calculations corresponding to the N second GEMV calculations into an insertion position of the GEMM calculations in the code, and deleting the N second GEMV calculations. This scheme can realize automatic identification and merging of GEMV calculations, and solve the problem of waste of matrix unit computing power.
Owner:HUAWEI TECH CO LTD

Data processing method, device, computer equipment and storage medium

Embodiments of the present application disclose a data processing method, apparatus, computer equipment, and storage medium. A neural network processor includes a global memory, a transfer buffer, and multiple computing units. A plurality of sub-matrices to be updated are obtained by performing block processing on a matrix to be updated, and the plurality of sub-matrices to be updated are evenly distributed to each computing unit. A first sub-matrix corresponding to each sub-matrix to be updated is determined in a first matrix, and the first sub-matrix is ​​moved from the global memory to a preset buffer of the computing unit corresponding to each sub-matrix to be updated. A second sub-matrix corresponding to each sub-matrix to be updated is determined in a second matrix, and the second sub-matrix is ​​moved from the global memory to the transfer buffer. A target first sub-matrix is ​​obtained from the preset buffer, and a target second sub-matrix is ​​obtained from the transfer buffer. A general matrix multiplication operation is performed on each sub-matrix to be updated, a target first sub-matrix, and a target second sub-matrix by the computing unit to obtain an updated sub-matrix.
Owner:PENG CHENG LAB

Dual-sparse neural processing unit with multi-dimensional routing of non-zero values

A general matrix-matrix (GEMM) accelerator core includes first and second buffers, a control logic circuit, and a first processing element (PE). The first buffer receives a elements of a first matrix A of activation values. The second buffer receives b elements of a second matrix B of weight values. The control logic circuit replaces a zero-valued a element in a first column of the first buffer with a nonzero-valued a element that is within a maximum borrowing distance of a location of the zero-valued a element in the first column of the first buffer. The PE receives a elements from the first column of the first buffer including the nonzero-valued element a selected to replace the zero-valued a element and receives b elements from locations in the second buffer that correspond to locations in the first buffer from where the a elements have been received by the PE.
Owner:SAMSUNG ELECTRONICS CO LTD

A matrix inversion method and system based on iterative calculation of Cholesky decomposition

ActiveCN119441699BComplex mathematical operationsGeneral matrixConjugate transpose
The present application discloses a matrix inversion method and system for iterative calculation based on Cholesky decomposition, which relates to the field of DSP system optimization technology. The method includes obtaining a target source matrix; performing a first iterative process on the target source matrix based on Cholesky decomposition to generate an upper triangular matrix; performing a second iterative process on the upper triangular matrix to generate an inverse matrix of the upper triangular matrix; performing a conjugate transpose process on the inverse matrix of the upper triangular matrix to generate a lower triangular matrix; wherein the lower triangular matrix is ​​stored in the form of whole column storage; the storage mode of the inverse matrix of the upper triangular matrix is ​​converted into a sequential storage form; performing matrix multiplication process on the inverse matrix of the upper triangular matrix and the lower triangular matrix to generate the inverse matrix of the target source matrix. The present application replaces cumulative summation with iteration, adopts complex multiplication and addition optimization calculation, supports multi-parallel operation, parallelizes zero-filling operation, can adapt to general matrix multiplication module, and reduces calculation time and area overhead.
Owner:NANJING UNIV

Data processing method and apparatus, computer device and storage medium

PCT designated stageWO2026076765A1Energy efficient computingPhysical realisationGeneral matrixAlgorithm
Disclosed in the embodiments of the present application are a data processing method and apparatus, a computer device and a storage medium. A neural network processor comprises a global memory, a transfer buffer and a plurality of computing units. The method comprises: partitioning a matrix to be updated, so as to obtain a plurality of sub-matrices to be updated, and equally allocating the plurality of sub-matrices to be updated to the computing units; determining, in a first matrix, a first sub-matrix corresponding to each sub-matrix to be updated, and transferring the first sub-matrix from the global memory to a preset buffer of the computing unit corresponding to the sub-matrix to be updated; determining, in a second matrix, a second sub-matrix corresponding to each sub-matrix to be updated, and transferring the second sub-matrix from the global memory to the transfer buffer; acquiring target first sub-matrices from the preset buffers, and acquiring target second sub-matrices from the transfer buffer; and performing, by means of each computing unit, general matrix multiplication on each sub-matrix to be updated, the target first sub-matrix and the target second sub-matrix, so as to obtain an updated sub-matrix.
Owner:PENG CHENG LAB

End-side oriented self-attention reasoning method, engine and multi-head fast decoding circuit

PendingCN122264127ABiological modelsInference methodsGeneral matrixAlgorithm
The application belongs to the technical field of circuits and systems, and discloses an end-side-oriented self-attention reasoning method, an engine and a multi-head fast decoding circuit. The self-attention reasoning method realizes self-attention reasoning calculation by sequentially traversing the cache key-value pairs through a token-by-token pipeline, dynamically maintaining the maximum attention score statistics, the softmax normalization factor cumulative value and the attention result cumulative vector. In the process, each key-value pair is only accessed and processed once, without the need to store the intermediate attention scores or to perform block-shaped softmax or secondary traversal. Through the above method and in combination with the multi-head fast decoding circuit designed in cooperation with the self-attention reasoning method, the storage consumption of self-attention calculation can be reduced without additional hardware parallelism, high-precision Attention calculation and low-precision general matrix operation can be uniformly executed, and multi-head fast decoding can be supported, so that the end device can also efficiently realize LLM decoding.
Owner:HUAZHONG UNIV OF SCI & TECH

Method for optimizing implementation of convolution operator for tensor computing unit

A kind of tensor computing unit convolution operator optimization implementation method faces, convolution operator is represented by the DSL of deep learning compiler, implicit general matrix multiplication is obtained by coordinate transformation to convolution calculation;Then after scheduling optimization is carried out to the convolution operator, the optimal search parameter is obtained by searching and the CUDA C code is generated through the back end of deep learning compiler, then the generated CUDA C code is integrated into neural network, the inference speed of convolutional neural network on NVIDIAGPU platform is improved.The application can improve the performance of automatic code generation of convolution operator in half-precision calculation, and provide guarantee for the performance of automatic code generation of fusion operator in neural network inference calculation.
Owner:SHANGHAI JIAOTONG UNIV

Large language model indexing method and device, computer equipment and storage medium

The embodiment of the invention discloses an indexing method and device for a large language model, computer equipment and a storage medium, and the method comprises the steps: generating an expert mask table based on a comparison result of a generalized coding sequence outputted by a router and a preset numerical value; executing a first type of prefix sum operation on the expert mask table to generate a first auxiliary index table, and integrating tail row data of the first auxiliary index table under a plurality of batch processing scenes to generate a second auxiliary index table; executing a second type of prefix sum operation on the tail row of the second auxiliary index table to generate a third auxiliary index table; according to the first auxiliary index table, the second auxiliary index table and the third auxiliary index table, rearranging input feature data to obtain rearranged feature data; performing expert-level batch general matrix multiplication on the rearranged feature data; and based on the gating weight output by the router, performing weighted fusion operation on results of the batch general matrix multiplication operation to obtain a final output feature of a big language model MOE architecture reasoning stage.
Owner:YUAN LI (BEI JING) BAN DAO TI JI SHU YOU XIAN GONG SI

Processing to accelerate distributed matrix multiplication operations

This disclosure relates to accelerating the processing of distributed matrix multiplication operations. The method proposed herein can efficiently perform operations such as General Matrix Multiplication (GEMM). The data required for such operations can be prefetched as needed, for example, immediately before a specific computation is performed on a pair of data blocks. Prefetching for subsequent computations can be performed during the current computation. After the current computation, the results can be accumulated in subsequent computations; for example, the accumulation operation can be offloaded using a data processing unit, thus freeing up one or more processing cores for other computations. Such operations can be performed in parallel until each block of the resulting matrix has been processed. In at least one embodiment, a single computational task can also be partitioned among multiple worker processes (e.g., processing units).
Owner:MELLANOX TECHNOLOGIES LTD(IL)

Weight inverse quantization matrix multiplication module and related equipment

The invention provides a weight inverse quantization matrix multiplication module and related equipment. A general matrix multiplication unit is used for sending read control information to a direct memory access unit; the direct memory access unit is used for generating a corresponding read request according to the read control information, sending the read request to the off-chip memory and receiving target data fed back by the off-chip memory, and the target data at least comprises target weight data; the target weight data are weight data matched with the target row identifier and the target column identifier at the same time in the quantized weight matrix; and the direct memory access unit is used for carrying out inverse quantization processing on the target weight data and providing the inverse quantized weight data to the general matrix multiplication unit. The inverse quantization processing is completed by the direct memory access unit, and the inversely quantized weight data is provided for the general matrix multiplication unit, so that global cache is not needed, the memory access bandwidth of the global cache is not occupied, and extra and unnecessary power consumption is avoided.
Owner:CIX TECH (SUZHOU) CO LTD

Method for Accelerating General Matrix Operations for Codebook-Based Quantization Models and Computing Device Using the Same

PendingKR1020260140094AGeneral matrixTheoretical computer science
The present invention relates to a method for accelerating general matrix operations for a codebook-based quantization model and a computing device for the same. A method for accelerating general matrix operations (GEMM: General Matrix Multiply) for a codebook-based quantization model, executed by a processor according to one embodiment of the present invention, may include: a step of reshaping an input activation vector to generate at least one input segment; a step of performing matrix operations between the codebook of the quantization model and the input segment to generate a partial sum book (Psum book: Partial sum book) containing partial sums corresponding to each code in the codebook; a step of extracting partial sums corresponding to each code included in the code matrix of the quantization model by referring to the partial sum book; and a step of generating an output vector corresponding to the input activation vector using the extracted partial sums.
Owner:NAVER CLOUD CORP