Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

34 results about "General matrix" patented technology

Memory architecture-oriented dual-precision general matrix multiplication optimization method and system

The invention belongs to the related technical field of high-performance computing, and provides a memory architecture-oriented dual-precision general matrix multiplication optimization method and system in order to solve the problems of limited computing power and access efficiency and the like in the prior art. Decomposing the matrix into a plurality of sub-matrix blocks according to the slave core array topology; the slave core receives the sub-matrix blocks issued by the master core, divides the sub-matrix blocks into small sub-matrix blocks based on a uniform blocking rule, loads the small sub-matrix blocks to an independent buffer area of a local data memory based on a DMA double-buffer protocol, divides the small sub-matrix blocks in the buffer area into SIMD vectors according to the SIMD unit characteristics of the slave core, and sends the SIMD vectors to the slave core; vectorization calculation and caching operation are alternately switched according to an iteration period through different independent buffer areas; and after all the slave cores finish calculation, the master core collects results written back to the master memory by the slave cores to obtain a final operation result, and double breakthrough of calculation power and memory access efficiency is realized.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Large language model weight inverse quantization reasoning device and method

The invention relates to the technical field of large language model deployment, and discloses a large language model weight inverse quantization reasoning device and method.The method comprises the steps that low-precision weight data is transmitted to a high-bandwidth storage from a host and then transmitted to an on-chip storage through the high-bandwidth storage; data conversion from a low-precision format to a high-precision format is completed in the on-chip memory, the data is multiplied by an inverse quantization factor to obtain recovered high-precision weight data, and the functional unit is responsible for executing general matrix multiplication of input data and the high-precision weight data after inverse quantization. And pipeline parallel execution of the inverse quantization operation and the general matrix multiplication operation is realized through a double-buffering technology. According to the invention, on the basis of a dual-path inverse quantization architecture of the vector processing unit and a dual-buffer mechanism in the on-chip memory, the problem of hardware adaptation of low-precision calculation is solved, and efficient execution of a low-precision conversion algorithm is realized under the condition of limited hardware resources.
Owner:JIANGNAN UNIV +1

Method and device for calculating matrix multiplied by vector, computing equipment and storage medium

The embodiment of the invention provides a method and device for calculating a matrix multiplied by a vector, computing equipment and a storage medium, the matrix is a matrix of M * K, the vector is a vector of K * 1, and M and K are positive integers. The method comprises the following steps: converting an M * K matrix into a first tensor of M * (K / N) * N, N being a positive integer, Ngt; 1 and K are multiples of N; converting the vector of K * 1 into a second tensor of N * (K / N); using a tensor calculation kernel to carry out general matrix multiplication calculation of a first tensor of M * (K / N) * N and a second tensor of N * (K / N) in batches to obtain M result matrices; elements on diagonals of each of the M result matrices are added to result in a matrix of M * K multiplied by each element on a result vector of M * 1 of the vector of K * 1. According to the scheme, high-throughput and low-delay high-speed matrix vector multiplication (MMV) operation is realized by utilizing the tensor calculation kernel.
Owner:SHANGHAI BIREN TECH CO LTD

Model reasoning acceleration method and device, electronic equipment and nonvolatile storage medium

The invention discloses a model reasoning acceleration method and device, electronic equipment and a nonvolatile storage medium. The method comprises the steps that in the process that a natural language processing model carries out general matrix calculation, a first matrix and a second matrix are divided into a plurality of blocks respectively, general matrix calculation is executed through an image processing unit, and the first matrix and the second matrix are matrixes needing general matrix calculation; loading the blocks from a global memory of the image processing unit to a shared memory of the image processing unit, and performing matrix operation by adopting threads corresponding to the blocks in the shared memory; and writing a result obtained by calculation of each thread back to a corresponding position in the global memory to obtain a matrix calculation result corresponding to matrix calculation of the first matrix and the second matrix. The technical problem of low model reasoning calculation efficiency caused by high memory access delay and low data reuse rate of a global memory of an image processing unit during reasoning of a large model in related technologies is solved.
Owner:CHINA TELECOM CORP LTD

Communication calculation parallel optimization method, multiprocessor system, medium and program product

The invention discloses a communication computing parallel optimization method, a multiprocessor system, a medium and a program product, and the method comprises the steps: configuring a first computing core and a second computing core on a single task flow for a general matrix multiplication subtask allocated to a single processor; wherein the first calculation core is used for executing a general matrix multiplication subtask, and the second calculation core is used for executing a set communication task; the set communication task comprises a full accumulation operator or a protocol dispersion operator; then, starting scheduling is conducted on the first calculation core and the second calculation core according to a preset dependency mechanism, so that the general matrix multiplication subtask and the set communication task are executed asynchronously in an overlapped mode; according to the method, the problem of poor reusability of the original Kernel caused by intrusive modification of the GEMM or re-implementation of the Kernel can be effectively avoided through parallel optimization of GEMM calculation and ensemble communication operation, and the performance overhead of the processor is reduced.
Owner:SHANGHAI BIREN TECH CO LTD

Kernel selection method and device during general matrix multiplication operation, equipment and storage medium

PendingCN122044838AResource allocationBiological modelsGeneral matrixAlgorithm
The embodiment of the invention provides a kernel selection method and device during general matrix multiplication operation, equipment and a storage medium, and belongs to the technical field of computers. The method comprises the following steps: acquiring general matrix multiplication problem size data; preprocessing the general matrix multiplication problem size data to obtain general matrix multiplication problem size features; based on a pre-trained problem size encoder, mapping the general matrix multiplication problem size feature into a problem size embedded vector of a preset dimension; taking the problem size embedded vector as a query vector, and performing nearest neighbor vector search in a pre-configured vector database to obtain a kernel configuration feature with the highest similarity; and outputting the kernel configuration feature as a selection result. The method is used for improving the kernel selection efficiency and precision during the operation of the general matrix multiplication.
Owner:DAWNING INT INFORMATION IND CO LTD +1

Local perception lookup table lookup method

The invention relates to the field of data lookup, and discloses a locality perception lookup table lookup method, which comprises the following steps of: S1, reordering between rows: clustering and reordering the rows of an input vector and a weight matrix according to the same quantized value to enable elements with the same quantized value to be continuously distributed; for a common matrix GeMV scene, cache locality is improved through rearrangement and subsequent blocking operation, and repeated loading of a lookup table LUT is reduced. And the data locality of the GeMV scene is remarkably improved through inter-line reordering and a virtual mapping strategy. On the basis of reordering of input vector element values, elements with the same value are aggregated, so that the LUT line loading times are reduced to 256 times at most from being in direct proportion to the input dimension d, and frequent replacement of the LUT lines in the WRAM is avoided; meanwhile, a weight matrix is divided into BR * BC sub-matrixes, the characteristic that vertical adjacent rows share an accumulator is utilized, multi-row sub-matrix calculation is loaded at a time, and virtual rearrangement maintains the row matching relation between an input vector and the weight matrix through index mapping.
Owner:RENMIN UNIVERSITY OF CHINA

Determination method of rock mass fracture network seepage field containing complex fracture morphology

The invention discloses a calculation method of a rock mass fracture network seepage field containing a complex fracture form, and relates to the field of geotechnical engineering. The method comprises the following steps: for each node in a rock mass fracture network, determining a source-sink item at the node based on the flow of a plurality of fracture units connected with the node; substituting the source sink item at the node and the equivalent permeability coefficient matrix into an overall matrix equation of fracture network seepage for solving to obtain a pressure head value of the node; the overall matrix equation is obtained by assembling unit seepage matrix equations of all nodes in the rock mass fracture network; in the construction process of the fracture seepage unit equation, the influence of the on-way resistance coefficient on the equivalent seepage flow velocity of the fracture in the complex wall surface form and the influence of the local resistance coefficient on the equivalent seepage flow velocity of the fracture in the complex wall surface form are considered; and determining the pressure head values of all nodes in the rock mass fracture network as a rock mass fracture network seepage field. The method can improve the calculation precision of the fracture network seepage field.
Owner:JIANGSU OCEAN UNIV

Matrix multiplication operators, operational methods, devices, graphics processors, and storage media

PendingCN122132654AComplex mathematical operationsGraphicsGeneral matrix
This invention provides a method, apparatus, graphics processor, and storage medium for matrix multiplication operators, relating to the field of artificial intelligence technology. The method includes: reading two operation matrices from high-bandwidth memory; determining a first target input matrix in the general matrix buffer (GMB) based on the two operation matrices, and identifying the other operation matrix as a second target input matrix; the first target input matrix is ​​arranged in a transposed manner; and writing the first and second target input matrices into a tensor core for matrix multiplication to obtain the target output matrix from the tensor core. This invention, by identifying the data arrangement of the two operation matrices, prioritizes inputting the first target input matrix (with a transposed arrangement) into the GMB, effectively avoiding hardware performance limitations associated with non-transposed readings, ensuring the tensor core maintains full-bandwidth data input, and significantly improving the computational efficiency of the matrix multiplication operator.
Owner:广州壁仞智能科技有限公司 +1

An address mapping method and system for GPU convolution acceleration

The application relates to the field of computer architecture and parallel computing technology, and discloses an address mapping method and system for GPU convolution acceleration, which is applied to the process of performing convolution operation based on general matrix multiplication by a GPU, dynamically remaps the memory access address of a workspace matrix, and comprises the following steps: identifying a data repetition mode in the workspace matrix; based on the data repetition mode, remapping the original address of a current memory access request to a target address by using an address mapping strategy; performing a data access operation by using the target address; if the target data is cached in a cache level of the GPU, directly returning the same value in the original access address; otherwise, loading data from a global memory and storage. The application avoids frequent global memory and storage access caused by scattered storage of repeated data in a traditional method, significantly improves the cache hit rate, and further breaks through the bottleneck of low cache utilization of an existing strategy.
Owner:SHANDONG UNIV

Computing method and apparatus for general matrix multiplication

A computing method and apparatus for general matrix multiplication (GEMM) are provided. The method comprises: in response to determining that a host needs to obtain a first calculation result of multiplication of first weight data and an input vector, performing a multiplication operation of second weight data and the input vector to obtain a second calculation result, wherein the first weight data satisfies a first arrangement rule for the host, and the second weight data is obtained by arranging the first weight data based on a second arrangement rule for a PIM apparatus different from the first arrangement rule; and obtaining the first calculation result based on the second calculation result.
Owner:SAMSUNG (CHINA) SEMICONDUCTOR CO LTD +1

Methods, apparatus, computing devices, and storage media for computing a matrix multiplication of vectors

Embodiments of the present disclosure provide a method, apparatus, computing device and storage medium for calculating matrix-vector multiplication, where the matrix is an M*K matrix and the vector is a K*1 vector, M and K are positive integers. The method comprises: converting the M*K matrix into an M*(K / N)*N first tensor, where N is a positive integer, N>1, and K is a multiple of N; converting the K*1 vector into an N*(K / N) second tensor; using a tensor computing core to batch-process the general matrix multiplication of the M*(K / N)*N first tensor and the N*(K / N) second tensor to obtain M result matrices; and adding the elements on the diagonal of each of the M result matrices to obtain each element on the M*1 result vector of the M*K matrix multiplied by the K*1 vector. The above scheme uses a tensor computing core to implement high-throughput, low-latency high-speed matrix-vector multiplication (MMV) operation.
Owner:SHANGHAI BIREN TECH CO LTD

Systems and methods for activation sparse and kernel sparse general matrix multiplication in neural networks

ActiveUS12670370B1General matrixAlgorithm
A system and method for performing multiplication for a neural network, e.g. for data of one or more layers in a neural network, may include loading a portion of a compressed version of a sparse input matrix into a cache memory; uncompressing a subset of the data in the portion of the compressed version of the sparse input matrix; and multiplying a sparse kernel matrix by the subset of the data using a set of instructions which are themselves created based on the sparse kernel matrix.
Owner:RED HAT INC

Method for accelerating a general matrix multiplication on a computer system

PCT designated stageWO2026008165A1Biological modelsMachine learningGeneral matrixTheoretical computer science
The present invention relates to a computer-implemented method, computer program, non-transitory, computer-readable medium comprising a program code and computer system for accelerating a General Matrix Multiplication (GEMM) on a computer system, and to methods, systems and computer programs / computer-readable media using such a computer-implemented method. A computer-implemented method for accelerating a General Matrix Multiplication, GEMM, on a computer system, comprises obtaining (310) a data structure comprising a representation of a first set of matrix layout variants known to be more efficient than a second set of matrix layout variants when being used to perform, by the computer system, a GEMM, with each matrix layout variant defining whether at least one of the matrices involved in the GEMM are to be transposed, and selecting (350) a matrix layout variant of the first set of matrix layout variants based on matrices involved in a GEMM to be performed, and performing (370) the GEMM on the matrices using the selected matrix layout variant.
Owner:NEC LAB EURO GMBH

Multi-low-rank adaptive fine tuning model reasoning optimization method and device

The invention provides a multi-low-rank adaptive fine tuning model reasoning optimization method and device. The method comprises a step of constructing a fusion input matrix and obtaining basic output by using a pre-training model, a step of carrying out fusion and feature extraction on input side low-rank matrixes of a plurality of LoRA adapters, and a step of carrying out fusion on output side low-rank moments of the plurality of LoRA adapters to obtain a correction term. The device comprises a corresponding processing module, namely, according to the method and the device, a plurality of small matrix multiplications of a plurality of LoRA fine tuning models are fused into a single general matrix multiplication (GEMM) operation, so that the calculation intensity is improved, the memory addressing operation is reduced, and the calculation efficiency is improved. The utilization rate of the GPU computing unit can be remarkably improved, so that the efficiency of multi-LoRA fine-tuning model reasoning is improved, and the multi-LoRA fine-tuning model reasoning can have better real-time performance and expandability.
Owner:FUDAN UNIVERSITY +1

A method for accelerating arbitrary precision sparse matrix multiplication based on tensor cores

The application discloses a method for accelerating arbitrary precision sparse matrix multiplication and addition operation based on a tensor core, which is in combination with deep learning and the program characteristics of parallel processing of sparse matrices, and is in combination of an arbitrary precision sparse matrix multiplication and addition operation operator and a parallel computing acceleration device, decomposes a general matrix into a sparse matrix of a set format, blocks the sparse matrix and specifies a corresponding parallel computing hardware unit for the sparse matrix, uses shared memory to reduce memory conflict and access delay, calls a high-level language API function library to execute the loading, calculation and storage process of the calculation data, and finally reduces the output data of the function, so that the final result of the arbitrary precision sparse matrix multiplication and addition operation on the parallel computing acceleration device is obtained. The application can fully utilize the parallel computing resources such as the tensor core on the A100, effectively reduce the program calculation amount and communication overhead, and greatly improve the program running performance.
Owner:XI AN JIAOTONG UNIV

A code optimization method and optimization device

ActiveCN119690442BIntelligent editorsRequirement analysisGeneral matrixAlgorithm
A code optimization method and an optimization device are provided. The method is applied to the optimization device, and includes: determining M first general matrix vector multiplication (GEMV) calculations in a code to be optimized based on an identification rule; determining N second GEMV calculations in the M first GEMV calculations based on a merging rule; inserting general matrix multiplication (GEMM) calculations corresponding to the N second GEMV calculations into an insertion position of the GEMM calculations in the code, and deleting the N second GEMV calculations. This scheme can realize automatic identification and merging of GEMV calculations, and solve the problem of waste of matrix unit computing power.
Owner:HUAWEI TECH CO LTD

Dual-sparse neural processing unit with multi-dimensional routing of non-zero values

A general matrix-matrix (GEMM) accelerator core includes first and second buffers, a control logic circuit, and a first processing element (PE). The first buffer receives a elements of a first matrix A of activation values. The second buffer receives b elements of a second matrix B of weight values. The control logic circuit replaces a zero-valued a element in a first column of the first buffer with a nonzero-valued a element that is within a maximum borrowing distance of a location of the zero-valued a element in the first column of the first buffer. The PE receives a elements from the first column of the first buffer including the nonzero-valued element a selected to replace the zero-valued a element and receives b elements from locations in the second buffer that correspond to locations in the first buffer from where the a elements have been received by the PE.
Owner:SAMSUNG ELECTRONICS CO LTD

Data processing method and apparatus, computer device and storage medium

PCT designated stageWO2026076765A1Energy efficient computingPhysical realisationGeneral matrixAlgorithm
Disclosed in the embodiments of the present application are a data processing method and apparatus, a computer device and a storage medium. A neural network processor comprises a global memory, a transfer buffer and a plurality of computing units. The method comprises: partitioning a matrix to be updated, so as to obtain a plurality of sub-matrices to be updated, and equally allocating the plurality of sub-matrices to be updated to the computing units; determining, in a first matrix, a first sub-matrix corresponding to each sub-matrix to be updated, and transferring the first sub-matrix from the global memory to a preset buffer of the computing unit corresponding to the sub-matrix to be updated; determining, in a second matrix, a second sub-matrix corresponding to each sub-matrix to be updated, and transferring the second sub-matrix from the global memory to the transfer buffer; acquiring target first sub-matrices from the preset buffers, and acquiring target second sub-matrices from the transfer buffer; and performing, by means of each computing unit, general matrix multiplication on each sub-matrix to be updated, the target first sub-matrix and the target second sub-matrix, so as to obtain an updated sub-matrix.
Owner:PENG CHENG LAB

End-side oriented self-attention reasoning method, engine and multi-head fast decoding circuit

PendingCN122264127ABiological modelsInference methodsGeneral matrixAlgorithm
The application belongs to the technical field of circuits and systems, and discloses an end-side-oriented self-attention reasoning method, an engine and a multi-head fast decoding circuit. The self-attention reasoning method realizes self-attention reasoning calculation by sequentially traversing the cache key-value pairs through a token-by-token pipeline, dynamically maintaining the maximum attention score statistics, the softmax normalization factor cumulative value and the attention result cumulative vector. In the process, each key-value pair is only accessed and processed once, without the need to store the intermediate attention scores or to perform block-shaped softmax or secondary traversal. Through the above method and in combination with the multi-head fast decoding circuit designed in cooperation with the self-attention reasoning method, the storage consumption of self-attention calculation can be reduced without additional hardware parallelism, high-precision Attention calculation and low-precision general matrix operation can be uniformly executed, and multi-head fast decoding can be supported, so that the end device can also efficiently realize LLM decoding.
Owner:HUAZHONG UNIV OF SCI & TECH

Method for optimizing implementation of convolution operator for tensor computing unit

A kind of tensor computing unit convolution operator optimization implementation method faces, convolution operator is represented by the DSL of deep learning compiler, implicit general matrix multiplication is obtained by coordinate transformation to convolution calculation;Then after scheduling optimization is carried out to the convolution operator, the optimal search parameter is obtained by searching and the CUDA C code is generated through the back end of deep learning compiler, then the generated CUDA C code is integrated into neural network, the inference speed of convolutional neural network on NVIDIAGPU platform is improved.The application can improve the performance of automatic code generation of convolution operator in half-precision calculation, and provide guarantee for the performance of automatic code generation of fusion operator in neural network inference calculation.
Owner:SHANGHAI JIAOTONG UNIV

Large language model indexing method and device, computer equipment and storage medium

The embodiment of the invention discloses an indexing method and device for a large language model, computer equipment and a storage medium, and the method comprises the steps: generating an expert mask table based on a comparison result of a generalized coding sequence outputted by a router and a preset numerical value; executing a first type of prefix sum operation on the expert mask table to generate a first auxiliary index table, and integrating tail row data of the first auxiliary index table under a plurality of batch processing scenes to generate a second auxiliary index table; executing a second type of prefix sum operation on the tail row of the second auxiliary index table to generate a third auxiliary index table; according to the first auxiliary index table, the second auxiliary index table and the third auxiliary index table, rearranging input feature data to obtain rearranged feature data; performing expert-level batch general matrix multiplication on the rearranged feature data; and based on the gating weight output by the router, performing weighted fusion operation on results of the batch general matrix multiplication operation to obtain a final output feature of a big language model MOE architecture reasoning stage.
Owner:YUAN LI (BEI JING) BAN DAO TI JI SHU YOU XIAN GONG SI

Processing to accelerate distributed matrix multiplication operations

This disclosure relates to accelerating the processing of distributed matrix multiplication operations. The method proposed herein can efficiently perform operations such as General Matrix Multiplication (GEMM). The data required for such operations can be prefetched as needed, for example, immediately before a specific computation is performed on a pair of data blocks. Prefetching for subsequent computations can be performed during the current computation. After the current computation, the results can be accumulated in subsequent computations; for example, the accumulation operation can be offloaded using a data processing unit, thus freeing up one or more processing cores for other computations. Such operations can be performed in parallel until each block of the resulting matrix has been processed. In at least one embodiment, a single computational task can also be partitioned among multiple worker processes (e.g., processing units).
Owner:MELLANOX TECHNOLOGIES LTD(IL)

Weight inverse quantization matrix multiplication module and related equipment

The invention provides a weight inverse quantization matrix multiplication module and related equipment. A general matrix multiplication unit is used for sending read control information to a direct memory access unit; the direct memory access unit is used for generating a corresponding read request according to the read control information, sending the read request to the off-chip memory and receiving target data fed back by the off-chip memory, and the target data at least comprises target weight data; the target weight data are weight data matched with the target row identifier and the target column identifier at the same time in the quantized weight matrix; and the direct memory access unit is used for carrying out inverse quantization processing on the target weight data and providing the inverse quantized weight data to the general matrix multiplication unit. The inverse quantization processing is completed by the direct memory access unit, and the inversely quantized weight data is provided for the general matrix multiplication unit, so that global cache is not needed, the memory access bandwidth of the global cache is not occupied, and extra and unnecessary power consumption is avoided.
Owner:CIX TECH (SUZHOU) CO LTD

Architectures and instruction sets for non-general matrix multiplication operations

Architectures and instruction sets for non-general matrix multiplication operations are described. An example programmable processor, which implements these architectures and instruction sets, includes multiple processing engines, multiple buffers (with at least one buffer storing a multidimensional array), and multiple iterator tables stored in one of the multiple buffers. The multiple processing engines and multiple buffers are tightly coupled within an area of the programmable processor to enable high-bandwidth data communications between components of the programmable processor. Furthermore, the multiple iterator tables include an iterator table for each of the multiple buffers, and an iterator table enables read or write access to a respective buffer based on using multiple offset-stride tuples stored in the iterator table to derive addresses to access the multidimensional array stored in the respective buffer. Memory elements used by the multiple processing engines to perform operations using the multiple iterator tables consist of the multiple buffers.
Owner:RGT UNIV OF CALIFORNIA

Image recognition method and device based on implicit general matrix multiplication

ActiveCN115293335BNeural learning methodsMultiplexingGeneral matrix
The application discloses an image recognition method and device based on implicit general matrix multiplication, and the method comprises the following steps: acquiring dimension information of an expected output matrix according to structure parameters of an input image and structure parameters of a convolution kernel; sequentially taking original data points of N*N order in the expected output matrix as aggregated data points to acquire an aggregated output matrix corresponding to the expected output matrix; wherein N is a positive even number; taking data points in the aggregated output matrix as block base points, acquiring the expected output matrix based on implicit general matrix multiplication, and recognizing the input image according to the expected output matrix. The technical scheme of the embodiment of the application realizes the multiplexing of data loading when the input matrix reads data from the physical layer, reduces the data loading time, improves the calculation efficiency of the heterogeneous hardware accelerator in executing the convolution operation, and avoids the performance decline problem caused by the different data loading logics of the boundary points and the non-boundary points.
Owner:DAWNING INFORMATION IND (BEIJING) CO LTD

Weight-sparse neural processing unit with multi-dimensional routing of non-zero values

A general matrix-matrix (GEMM) accelerator core includes first and second buffers, and a processing element (PE). The first buffer receives a elements of a matrix A of activation values. The second buffer receives b elements of a matrix B of weight values. The matrix B is preprocessed with a nonzero-valued b element replacing a zero-valued b element in a first row of the second buffer based on the zero-valued b element being in the first row of the second buffer. Metadata is generated that includes movement information of the nonzero-valued b element to replace the zero-valued b element. The PE receives b elements from a first row of the second buffer and a elements from the first buffer from locations in the first buffer that correspond to locations in the second buffer from where the b elements have been received by the PE as indicated by the metadata.
Owner:SAMSUNG ELECTRONICS CO LTD

Data processing method and device, computer equipment and storage medium

The invention discloses a data processing method and device, computer equipment and a storage medium. The method comprises the following steps: constructing a first fusion weight matrix according to a column direction splicing mode and three groups of weight matrixes Q, K and V required by an attention mechanism in a Transform model; performing uniform low-bit quantization processing on the first fusion weight matrix to generate a second fusion weight matrix; performing general matrix multiplication calculation on an activation input matrix and the second fusion weight matrix to obtain a fusion output matrix; segmenting and recovering the fusion output matrix according to a preset dimension, and extracting Q, K and V matrixes; and performing error correction processing on the extracted Q, K and V matrixes to obtain corrected Q, K and V matrixes. The invention relates to the field of artificial intelligence model reasoning optimization, and aims to realize efficient joint calculation of Q, K and V matrixes, reduce storage and calculation cost and improve reasoning performance through weight fusion representation, quantitative collaborative optimization and matrix multiplication reconstruction.
Owner:RED BRICK INTELLIGENT MODEL (SHANGHAI) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Convolution and general matrix multiplication accelerator, method and electronic equipment

The invention provides a convolution and general matrix multiplication accelerator, a convolution and general matrix multiplication method and electronic equipment, and relates to the technical field of computers. The main structure of the convolution and general matrix multiplication accelerator is that the convolution and general matrix multiplication accelerator comprises at least two input caches, at least one calculation engine, an engine control register and an output cache; each calculation engine comprises at least one unit array, each unit array comprises at least one multiplication accumulator array and a summation accumulator, and each multiplication accumulator array comprises at least two multiplication accumulators; and the engine control register is used for controlling the working mode of the calculation engine to carry out different types of convolution operations or general matrix multiplication operations. According to the method, convolution operation and general matrix multiplication operation of the same array can be realized, and hardware resources are saved, so that calculation of a convolutional neural network and a neural network of an attention mechanism architecture or other neural networks depending on convolution and general matrix multiplication operation is accelerated.
Owner:CHENGDU SINO MICROELECTRONICS TECH CO LTD

Deep learning-based ill-conditioned matrix SVD (Singular Value Decomposition) preprocessing method, equipment and medium

The invention discloses an ill-conditioned matrix SVD decomposition preprocessing method and device based on deep learning, and a medium, and belongs to the technical field of numerical algebra and deep learning. The method comprises the following steps: constructing an iterative deep neural network learning framework, extracting matrix features by using a convolutional layer, and constructing an orthogonal matrix through a House holder reflection decomposition method; training the network by using a mixed loss function, forcing the network to output an approximate diagonal matrix by punishing off-diagonal elements, and performing multiple rounds of iterative optimization by taking a learning result as the input of a new round of training; a precondition is constructed based on a matrix obtained through training, an original ill-conditioned linear equation set is preprocessed, and a preconditioned equation set is solved through an iterative algorithm. The method can be adapted to an ill-conditioned general matrix without a special structure, effectively solves the problems of insufficient approximation precision and poor generalization ability of a traditional method, and remarkably improves the numerical stability, convergence efficiency and calculation precision of high-dimensional ill-conditioned matrix solution.
Owner:10TH RES INST OF CETC +1