Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

51 results about "Matrix partitioning" patented technology

Matrix transposition method of DDR3 read-write controller based on AXI4 bus

The invention discloses a matrix transposition method of a DDR3 read-write controller based on an AXI4 bus. The matrix transposition method solves the problems that an existing transposition method is low in data access efficiency, serious in storage resource waste and poor in system portability. The method comprises the following steps: dividing an original echo matrix into a plurality of sub-matrix blocks, mapping and storing the sub-matrix blocks into a Bank of DDR3 to obtain a three-dimensional mapping result; reading data in each sub-matrix block in the three-dimensional mapping result by utilizing a DDR3 read control module, and writing the read data into an RAM (Random Access Memory) of an FPGA (Field Programmable Gate Array) to obtain a transposed matrix; writing the transposed matrix into the original Bank of the DDR3 by using a DDR3 write control module to obtain the DDR3 in which the data is written; according to the invention, low on-chip storage resource occupation independent of matrix scale is realized, and high-efficiency processing requirements of large-scale echo data in satellite-borne SAR real-time imaging processing are better met.
Owner:XIDIAN UNIV

Quantization method, reasoning method and equipment of industry large model and storage medium

The embodiment of the invention provides a quantification method and reasoning method of an industry large model, equipment and a storage medium. When an original weight matrix of a network layer in an industry large model is quantified, values of matrix elements on diagonals of an original Hessian matrix are obtained based on an input data matrix of the network layer and can represent activation values, and the positions of columns in the original weight matrix and the original Hessian matrix of the network layer are reordered according to an activation value descending mode. Carrying out matrix partitioning processing on the reordered first weight matrix and the first Hessian matrix respectively, and quantizing the first sub-weight matrix based on the quantization parameter of each sub-weight matrix in the first weight matrix in sequence by taking the partitioned matrix as granularity; the method comprises the steps of quantizing a first weight matrix, updating other first weight sub-matrixes which are not quantized based on the quantized first weight sub-matrix, and finally determining a quantization result of an original weight matrix of a network layer based on a quantization result of each first weight sub-matrix in the first weight matrix, so that the size of an industry large model is compressed.
Owner:HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD

Matrix multiplication implementation method and device, electronic equipment, storage medium and program product

The invention relates to the technical field of artificial intelligence, and provides a matrix multiplication implementation method and device, electronic equipment, a storage medium and a program product.The method comprises the steps that calculation cores on a chip are divided into a plurality of calculation groups based on a plurality of matrix multiplication operations needing to be executed at the same time, each calculation group at least comprises two calculation cores, each calculation group corresponds to one matrix multiplication operation; and controlling each calculation group to execute the corresponding matrix multiplication operation, and obtaining a result matrix of each matrix multiplication operation. According to the method, the calculation cores are grouped, and each calculation group executes one matrix multiplication operation in parallel, so that the function of executing a plurality of matrix multiplication operations in parallel on the chip is realized; the number of calculation cores in each calculation group is reduced relative to the whole chip, one matrix in each matrix multiplication operation is divided into a smaller number of sub-matrixes, and the number of rows of the sub-matrixes is large, so that the number of rows of the sub-matrixes can cover the minimum calculation granularity of the calculation cores, and the utilization rate of the calculation cores is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Method for performing matrix multiplication operation by processor comprising plurality of computing units

The invention relates to the technical field of computers, and provides a method for executing matrix multiplication by a processor comprising a plurality of computing units. In the method, a quantization parameter matrix is divided into a plurality of parameter blocks, and the parameter blocks are sequentially and respectively distributed to a plurality of calculation units in a non-repeated manner, so that at most one calculation unit is distributed with less than a first number of continuous parameter blocks in the plurality of parameter blocks, each computing unit is allocated a first number of consecutive parameter blocks in the plurality of parameter blocks; each calculation unit is used for carrying out dequantization on the distributed parameter blocks; each of the plurality of calculation units obtains a portion of the input matrix that should be multiplied by the allocated parameter block, and performs matrix multiplication on the dequantized allocated parameter block and the portion of the input matrix to complete matrix multiplication of the input matrix and the model parameter. Therefore, only one-time solution quantization needs to be carried out on the quantized parameter matrix globally, calculation resources are saved, and the reasoning efficiency is improved.
Owner:SHANGHAI INFINIGENCE AI INTELLIGENT TECHNOLOGY CO LTD

Matrix memory reading method applied to sublimation NPU

The invention relates to the field of heterogeneous parallel computing, in particular to a matrix memory reading method applied to mercuric chloride NPU, which comprises the following steps of: dividing a matrix into a regular matrix and an irregular matrix according to feature information of the matrix; arranging memory data of the regular left matrix and the regular right matrix into a format required by the L1 cache for matrix multiplication; according to the equivalent bandwidth quantization preprocessing income, judging whether to use a vector operation unit to preprocess memory data arrangement on the irregular matrix; when the equivalent bandwidth after preprocessing is larger than the original bandwidth, preprocessing of memory data arrangement is carried out on the irregular matrix, and otherwise, preprocessing of memory data arrangement is not carried out on the irregular matrix; and reading the preprocessed matrix to a cache by using an efficient memory reading interface, and carrying out matrix multiplication. According to the method, the problem that the memory bandwidth cannot be efficiently utilized in a memory access bottleneck scene can be optimized, high-bandwidth memory reading can be realized during matrix multiplication, and the overall calculation efficiency of a system is improved.
Owner:SOUTH CHINA UNIV OF TECH

Acceleration unit configured for multi- dimensional block-scaled matrices

To perform matrix multiplication operations for one or more applications, a processing system includes an acceleration unit (AU) having a block-scaled dot-product circuitry configured to multiply a first matrix by a second matrix. To this end, the block-scaled dot-product circuitry first partitions the first matrix into one or more multi-dimensional scaled blocks and the second matrix also into one or more multi-dimensional scaled blocks. The block-scaled dot-product circuitry next determines dot products of respective portions of the first matrix and corresponding portions of the second matrix using the multi-dimensional scaled blocks of the matrices and then combines these dot products to determine the dot product of the first matrix and the second matrix.
Owner:XILINX INC

Method and system for providing vector sparsification in neural networks

The present disclosure relates to a method and system for providing vector sparsification in a neural network. In some embodiments, an exemplary method for providing vector sparsification in a neural network includes partitioning a matrix associated with the neural network into a plurality of vectors, selecting a first subset of non-zero elements from the plurality of vectors to form a pruned matrix, and outputting the pruned matrix and performing the neural network using the pruned matrix. This method can address the irregular workload as a bottleneck for performing most of the existing neural networks.
Owner:ALIBABA GROUP HOLDING LTD

Large-scale 2D convolution operator acceleration method for DSP platform

This invention discloses a large-scale two-dimensional convolution operator acceleration method for DSP platforms, belonging to the field of digital signal processing. This method selects the im2col or col2im algorithm to rearrange data based on the convolution type; adopts a three-segment matrix partitioning strategy to adapt to L3 cache capacity; utilizes an EDMA buffer ping-pong architecture to achieve pipeline parallelism for data transmission and computation; implements SIMD instruction-level optimization, using DMPYSP and DADDSP instructions in conjunction with pipeline optimization to achieve four FP32 multiplication and addition operations per cycle; and implements multi-core parallel scheduling, using OpenMP to achieve task-level and data-level parallelism. This method has been tested on a TI TMS320C6678 platform to achieve efficient inference of SAR target detection networks, providing a feasible solution for real-time inference of CNN networks on DSP platforms.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS +1

Multi-head attention calculation task optimization method and system based on multiple processing units

The invention discloses a multi-head attention computing task optimization method and system based on multiple processing units, belongs to the technical field of artificial intelligence, and aims to solve the technical problem of how to efficiently map complex computing tasks of the MHA to a tile type array architecture and improve the hardware utilization rate and the computing performance. Comprising the following steps: logically constructing a processing cluster by a plurality of processing units; dividing the input matrix into a plurality of data blocks, and dividing the data blocks into a plurality of data pieces; distributing each data piece to a corresponding processing unit in a row multicast or column multicast mode through a hardware collective communication primitive of the network-on-chip; each processing unit in the processing cluster carries out attention calculation based on the received data fragments, and Softmax normalization operation is carried out on the local attention score matrix; and through hardware collective communication primitives of the network-on-chip, local output results of all the processing units in the row or the column where the edge processing unit is located are aggregated.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

A data processing architecture, chip, matrix multiplication and neural network calculation method

PendingCN122477461AAlgorithmParallel computing
A data processing architecture, chip, matrix multiplication and neural network calculation method, the data processing architecture (20) comprises: a node array of P rows and Q columns; P×Q nodes (10) in the node array respectively store a first submatrix obtained by dividing a first matrix, and a second submatrix obtained by dividing a second matrix; each node (10) is configured to send the second submatrix stored in the node to other P-1 nodes in the column where the node is located, respectively, to perform multiplication calculation on the first submatrix stored in the node and the transposed matrix of the plurality of second submatrices, respectively, to obtain a plurality of first type submatrix calculation results; according to a preset first correspondence relationship, the first type submatrix calculation results corresponding to other nodes in the row where the node is located are respectively sent to the corresponding nodes, and the first type submatrix calculation result corresponding to the node calculated is added to the first type submatrix calculation result received from other nodes to obtain a first result submatrix of the node.
Owner:SUNMMIO SCIENCE & TECHNOLOGY (BEIJING) CO LTD

Implementation method, device and medium of a general fixed-point matrix multiplier based on FPGA high-performance computing architecture

This invention discloses a method, apparatus, and medium for implementing a general-purpose fixed-point matrix multiplier based on a high-performance computing architecture in an FPGA. The method includes: designing matrix partitioning strategies at different levels based on the parallelism of the AI ​​engine array resources on the Versal ACAP platform; data scheduling and multiplexing based on the data packet stream and data packet exchange of the AI ​​engine and AXI stream, accelerating the kernel through matrix partitioned multiplication, and designing data scheduling and multiplexing strategies; and implementing a high-throughput vectorized matrix multiplication pipeline on the AI ​​engine vector processor. This invention achieves multi-level partitioning of matrix multiplication based on the AXI stream transport protocol and AI engine array on the Versal ACAP platform, enabling efficient utilization of hardware resources, effectively improving data reuse rate, and achieving high parallelism, while achieving high computational speed under the high-speed clock of the AI ​​engine. This invention can be widely applied in the field of high-performance computing.
Owner:SOUTH CHINA UNIV OF TECH

Data processing method and device, electronic equipment, storage medium and program product

The invention provides a data processing method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of large models.The method comprises the steps that an input matrix is divided into a plurality of sub-matrixes, and the number and the position index of each sub-matrix are determined based on a grouping staggering method; executing corresponding sub-matrix operation based on at least one calculation thread block in the thread block network, and transmitting an operation result of the corresponding sub-matrix based on at least one communication thread block corresponding to the at least one calculation thread block; based on the serial number sequence corresponding to each sub-matrix, storing the operation result of each sub-matrix in a cache, and based on the serial number and the position index corresponding to each sub-matrix and the operation results of the sub-matrixes sequentially stored in the cache, determining an output matrix; therefore, the communication thread blocks are arranged between the calculation thread blocks, so that the calculation thread blocks can synchronously perform sub-matrix operation in the communication process of the communication thread blocks, and the resource utilization rate of the graphics processor is improved.
Owner:INSPUR (SHANDONG) COMPUTER TECH CO LTD

High-order matrix operation method and device, electronic equipment and storage medium

The invention discloses a high-order matrix operation method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a plurality of target matrixes to be operated, dividing each target matrix into a plurality of square matrixes, and distributing each square matrix to a corresponding chip; a ping-pong transmission algorithm is adopted through a plurality of chips, and on-chip matrix multiplication is completed on the distributed square matrix; respectively corresponding multiplication results are transmitted to a target chip in parallel by adopting a ping-pong transmission algorithm through the plurality of chips; and performing accumulation operation on the multiple multiplication operation results through the target chip to obtain a target operation result corresponding to the target matrix. According to the technical scheme of the embodiment of the invention, the multiplication efficiency of the high-order matrix and the utilization rate of hardware resources can be improved.
Owner:SHANGHAI SMARTLOGIC TECHNOLOGY LTD

A heterogeneous edge distributed collaborative full-volume large language model parallel inference method

PendingCN122287796ALinguistic modelAlgorithm
This invention discloses a heterogeneous edge distributed collaborative parallel inference method for a large language model, relating to the field of distributed edge computing technology. The method includes: determining the calculation formulas for the parameter quantity and inference time of the self-attention layer allocated to each edge device; determining the calculation formulas for the parameter quantity and inference time of the multilayer perceptron allocated to each edge device; obtaining the total memory constraint; optimizing the minimum computation time of the self-attention layer and the multilayer perceptron under the total memory constraint to obtain the attention head allocation ratio of the self-attention layer and the parameter matrix partitioning ratio of the multilayer perceptron; distributing the parameters of the large language model to each edge device according to the allocation and partitioning ratios, and performing parallel inference of the large language model through each edge device. This method improves the segmentation accuracy of the large language model parameters on edge devices.
Owner:SHENZHEN UNIV

Data processing architecture, chip, matrix multiplication and neural network calculation method

The invention discloses a data processing architecture, a chip, a matrix multiplication method and a neural network calculation method. A data processing architecture (20) includes: an array of nodes of P rows Q columns; p * Q nodes (10) in the node array respectively store a first sub-matrix obtained by dividing a first matrix X to be subjected to matrix multiplication and a second sub-matrix obtained by dividing a second matrix W to be subjected to matrix multiplication; the node (10) is set to respectively send the first sub-matrix stored by the node to other Q-1 nodes in the same row as the node and respectively send the second sub-matrix stored by the node to other P-1 nodes in the same column as the node, a result submatrix corresponding to the node is obtained through calculation according to the submatrixes stored by the node and received by the node; wherein the sub-matrix comprises a first sub-matrix and a second sub-matrix; and the result sub-matrixes of the P * Q nodes form a result matrix Y obtained by performing matrix multiplication on the first matrix X and the second matrix W.
Owner:SUNMMIO SCIENCE & TECHNOLOGY (BEIJING) CO LTD

Shared memory allocation method, device, medium, equipment and product for matrix calculation

The application discloses a shared memory allocation method and device for matrix calculation, a medium, equipment and products, and the method comprises the following steps: when the second storage body with smaller storage capacity is sufficient to accommodate the first multiplier matrix, allocating the first multiplier matrix with larger memory requirement to the second storage body, and allocating the second multiplier matrix to the first storage body; checking whether the storage body with the largest remaining space is sufficient to accommodate the bias matrix; if yes, allocating the bias matrix to the storage body with the largest remaining space and performing matrix fusion calculation; if not, if the memory requirement of the bias matrix is not greater than the total remaining space, dividing the bias matrix into two sub-matrices, and allocating the two sub-matrices to the first storage body and the second storage body respectively to perform matrix fusion calculation; and allocating the result matrix obtained through matrix fusion calculation to the memory occupation space of the bias matrix through multiplexing operation. The application can maximize the utilization of the remaining space of the shared memory, and significantly improve the success rate of memory allocation.
Owner:SHANGHAI BIREN TECH CO LTD

Matrix multiplication implementation method and device, electronic equipment, storage medium and program product

The application relates to the technical field of artificial intelligence, and provides a matrix multiplication implementation method and device, electronic equipment, a storage medium and a program product, the method comprising the following steps: based on multiple matrix multiplication operations to be simultaneously executed, dividing computing cores on a chip into multiple computing groups, each computing group comprising at least two computing cores, and each computing group corresponding to one matrix multiplication operation; controlling each computing group to respectively execute the corresponding matrix multiplication operation, and obtaining a result matrix of each matrix multiplication operation. According to the application, the computing cores are grouped, each computing group executes one matrix multiplication operation in parallel, and the function of executing multiple matrix multiplication operations on the chip in parallel is realized; moreover, the computing cores in each computing group are reduced relative to the entire chip, for one matrix in each matrix multiplication operation, the matrix is divided into a smaller number of sub-matrices, the number of rows of the sub-matrices is relatively large, the number of rows of the sub-matrices can cover the minimum calculation granularity of the computing cores, and the utilization rate of the computing cores is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Computing system, model training method and computing device

A computing system, a model training method and a computing device. The computing system comprises a first processing unit and a plurality of second processing units, wherein the first processing unit can instruct a plurality of local second processing units to divide a matrix of computation results into a plurality of blocks; and the plurality of second processing units in the computing system perform matrix multiplication computation on submatrices obtained after division by the first processing unit, and sequentially send, on the basis of the generation order of blocks in a submatrix of computation results, the blocks in the submatrix of computation results. Each second processing unit partitions a submatrix of computation results into a plurality of blocks, such that the time for the second processing unit to send a computation result each time can be reduced, thereby preventing the communication bandwidth of a processing unit or a memory, which receives results from a plurality of second processing units, from remaining in an idle state for a long time. The data volume of each block is relatively small, so that the processing unit or the memory receiving the results from the plurality of second processing units is less prone to a communication bottleneck, thereby achieving the effect of reducing the cost.
Owner:HUAWEI TECH CO LTD

Quantum logic gate compression storage method and device and electronic equipment

This application relates to the field of quantum computing technology, and in particular to a method, apparatus, and electronic device for compressed storage of quantum logic gates, comprising: creating a decision graph and using the root node of the decision graph as the current node; dividing the matrix into four sub-matrices with the same number of rows and columns, creating a child node corresponding to each sub-matrix, and associating the pointer of the current node with each child node; when a sub-matrix satisfies a preset termination condition, using the child node corresponding to the sub-matrix as a leaf node and storing the leaf node as a scalar value of the corresponding sub-matrix; recursively decomposing the quantum gate matrix and shared sub-matrices by creating a decision tree, decomposing complex matrix operations into smaller sub-problems, reducing redundant computations; and using the decision graph to explore the sparsity and regularity of the quantum gate matrix, efficiently representing the quantum gate matrix and state vector, thereby achieving compressed storage of matrices in quantum circuits and avoiding redundant computations.
Owner:ORIGIN QUANTUM COMPUTING TECH (HEFEI) CO LTD

Sequence processing method and processing unit cluster

The invention discloses a sequence processing method and a processing unit cluster, the processing unit cluster comprises at least one processing unit group, the processing unit group comprises at least one processing unit, and the method comprises the following steps: distributing a plurality of subsequences split from a data sequence to a plurality of processing units in the processing unit group; the processing unit converts the subsequences into a query matrix and a key matrix; determining a first hash code of each first element in the query matrix and a second hash code of each second element in the key matrix; obtaining a plurality of query matrix blocks divided from the query matrix and a plurality of key matrix blocks divided from the key matrix; based on the first hash code and the first position index of the first element and the second hash code and the second position index of the second element, screening out a target key matrix block meeting a matching rule with the query matrix block from the plurality of key matrix blocks; and executing attention calculation based on each first element in the query matrix block and each second element in the target key matrix block.
Owner:LENOVO SOFTWARE

Data processing method and system and related equipment

The invention discloses a data processing method and device and related equipment, which can perform LU decomposition of R columns of sub-matrixes on a to-be-decomposed matrix needing to be decomposed in each iteration process in a process of performing LU decomposition calculation on applied to-be-processed data, namely a dense matrix, and calculate a matrix multiplication taking a product of R and sub-matrix dimensions as a dimension. The dimension of matrix multiplication calculation is large, and efficient calculation of matrix multiplication can be achieved. In this way, the dense matrix can be divided into the sub-matrixes with the small granularity, it is guaranteed that the calculation efficiency of matrix multiplication is high on the premise that process loads are balanced, and then the overall efficiency of matrix decomposition calculation is improved.
Owner:HUAWEI TECH CO LTD

A data processing acceleration method and device for sparse matrix multiplication operator

The present invention provides a data processing acceleration method and device for a sparse matrix multiplication operator, comprising: obtaining a sparse matrix multiplication operator to be processed during data processing, compressing the sparse matrix, and storing the compressed sparse matrix in a form including a non-zero element array, a row pointer array, a global column index array, and a local column index array. The global column index records the column index of the non-zero elements in a set of consecutive preset number of sparse rows divided by the sparse matrix, and the global column index is not repeated in each column index array. The local column index records the position index of each non-zero element in the global column index array. The global column index is used to access the corresponding row data in the dense matrix and read it into a temporary buffer area. The local column index is used to access the corresponding row data in the dense matrix in the temporary buffer area. The product of the non-zero element and the corresponding row data in the dense matrix is ​​calculated and accumulated, and the result matrix is ​​output. The present invention can improve data processing speed and reduce memory access overhead.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Data processing architecture, chip, matrix multiplication computation method, and neural network computation method

A data processing architecture, a chip, a matrix multiplication computation method, and a neural network computation method. The data processing architecture (20) comprises a node array having P rows and Q columns. P×Q nodes (10) in the node array respectively store first sub-matrices obtained by partitioning a first matrix X on which matrix multiplication is to be performed, and second sub-matrices obtained by partitioning a second matrix W on which matrix multiplication is to be performed. Each node (10) is configured to separately send the first sub-matrix stored in the node to the other Q-1 nodes in the same row as the node, separately send the second sub-matrix stored in the node to the other P-1 nodes in the same column as the node, and on the basis of sub-matrices stored in and received by the node, compute a result sub-matrix corresponding to the node, wherein the sub-matrices include the first sub-matrices and the second sub-matrices, and the result sub-matrices of the P×Q nodes constitute a result matrix Y obtained by performing matrix multiplication on the first matrix X and the second matrix W.
Owner:SUNMMIO SCIENCE & TECHNOLOGY (BEIJING) CO LTD

Method for automatically dividing assembly structures

The invention belongs to assembly structure division in the field of assembly sequence planning, and particularly relates to a method for automatically dividing sub-assemblies, which comprises the following specific steps of: 1, acquiring six interference matrixes of an assembly in the direction of an overall coordinate axis; 2, acquiring a static interference matrix, namely a connection matrix, of the assembly; 3, performing C + + programming on the connection matrix and the gravity direction interference matrix, and generating a support matrix based on an intersection taking mode; 4, firstly defining a basic part of the assembly body by utilizing the support matrix based on an actual assembly site; step 5, dividing sub-assemblies based on the support matrix; and 6, for the obtained sub-assemblies and basic parts, the assemblies are defined into an assembly hierarchy tree form conforming to an assembly site. The invention defines a sub-assembly division method based on an actual assembly site, the method can be used as a reference in a subsequent assembly sequence planning process, the assembly sequence of the sub-assembly layers can be planned firstly, and then the sequence of parts in each sub-assembly is planned.
Owner:SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI

A method for storing a parity check matrix, a data decoding method, an apparatus, and an electronic device.

This application discloses a parity check matrix storage method, data decoding method, apparatus, and electronic device, relating to the field of data processing technology. The parity check matrix storage method includes: obtaining the base matrix corresponding to the parity check matrix; dividing the base matrix into multiple sub-base matrices with the same data format; for each sub-base matrix, determining the shift value corresponding to each non-zero item in each row of the sub-base matrix, and the column index in the sub-base matrix, forming multiple sets of data and row end marker information corresponding to each row; storing the multiple sets of data and row end marker information corresponding to each row by column to obtain the corresponding storage matrix. Thus, since the number of columns and / or rows in each sub-base matrix is ​​reduced relative to the base matrix, the redundant space required for alignment when storing the column index and shift value of each row of the base matrix can be reduced. Compared with the scheme of storing based on the column index and shift value of the non-zero item in each row of the base matrix, the storage space can be reduced.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

Matrix multiplication hardware acceleration method and hardware acceleration circuit

The invention discloses a matrix multiplication hardware acceleration method and a hardware acceleration circuit, and belongs to the field of artificial neural network acceleration calculation and chips. Dividing the matrix into a plurality of coarse blocks in advance according to the size of a data cache module; the data reading module reads the plurality of coarse blocks from the external storage module according to a matrix multiplication sequence, allocates different fine blocks of the matrix A to each matrix multiplication unit in the matrix multiplication array, transpose the matrix B and allocates one fine block to each matrix multiplication unit in the matrix multiplication array; the matrix multiplication array can obtain a plurality of parts and data in each calculation; and each matrix multiplication unit in the matrix multiplication array obtains the part of the input matrix, which should be multiplied by the distributed fine blocks, and the part is subjected to matrix multiplication so as to complete the matrix multiplication of the input matrix and the model parameters of the large model.
Owner:58TH RES INST OF CETC

Shared memory allocation method and device for matrix calculation, medium, equipment and product

The invention discloses a shared memory allocation method and device for matrix calculation, a medium, equipment and a product, and the method comprises the steps: allocating a first multiplier matrix with a larger memory demand to a second memory bank and allocating the second multiplier matrix to the first memory bank when the second memory bank with a smaller storage capacity is enough to accommodate the first multiplier matrix; checking whether the memory bank with the maximum residual space is enough to accommodate the bias matrix; if yes, distributing the bias matrix to the memory bank with the maximum residual space, and executing matrix fusion calculation; if not, under the condition that the memory requirement of the bias matrix is not larger than the total residual space, the bias matrix is divided into two sub-matrixes, and the two sub-matrixes are distributed to the first memory bank and the second memory bank respectively so as to execute matrix fusion calculation; and through multiplexing operation, distributing a result matrix obtained by matrix fusion calculation to a memory occupation space of the bias matrix. According to the method, the residual space of the shared memory can be utilized to the maximum extent, and the memory allocation success rate is remarkably improved.
Owner:SHANGHAI BIREN TECH CO LTD