Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

19 results about "Large matrices" patented technology

Parallel optimization method of low-rank adapter and task perception scheduling system

The invention relates to the technical field of large-model lightweight fine tuning, and discloses a parallel optimization method of a low-rank adapter and a task awareness scheduling system.The parallel optimization method comprises the steps that an increment matrix of the low-rank adapter is fragmented to a tensor parallel group according to rows, the tensor parallel group comprises a plurality of computing devices, and the computing devices are used for computing the increment matrix of the low-rank adapter; the fragmentation granularity is dynamically determined according to the equipment hardware capability and the dimension of the increment matrix; dynamically scheduling the tasks with strong conflicts to different task parallel groups based on the calculated inter-task gradient conflict coefficient; and for a plurality of tasks scheduled to the same task parallel group, adapter calculation is merged into unified large matrix operation by adopting a micro-batch processing technology, and LayerNorm and adapter projection calculation are merged into a single calculation kernel by adopting a kernel fusion technology. According to the method, communication redundancy and synchronization overhead in distributed training are reduced, task interference during multi-task parallel is effectively eliminated, and the problem of computing resource fragmentation is solved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

An EIGD time-delay power system stability analysis method and device

The application belongs to the technical field of power systems, and discloses an EIGD time-delay power system stability analysis method and device, wherein the method comprises the following steps: a linear model of the time-delay power system is established, and the differential equation in the model is converted into an abstract Cauchy problem by using an infinitesimal generator; a group of discrete points is selected for each time-delay interval of the time-delay power system, a discrete function space is established according to the discrete points, and the time-delay variable is discretized to generate a low-order partial infinitesimal generator discretization matrix; displacement and inverse transformation are performed on the discretization matrix to obtain an inverse matrix, and the characteristic value of the system is obtained according to the inverse matrix; the characteristic value is checked by using the Newton method to obtain the accurate characteristic value and the characteristic vector, and the stability of the time-delay power system is analyzed and judged. The method solves the problem of large matrix LU decomposition calculation amount in the existing EIGD time-delay power system characteristic value sparse calculation method based on the DDAE model.
Owner:ELECTRIC POWER RESEARCH INSTITUTE OF STATE GRID SHANDONG ELECTRIC POWER COMPANY +2

Multi-thread redundancy acceleration method and system based on matrix operator splitting

PendingCN121834105AImplement selective redundancyEnsure computing reliabilityComplex mathematical operationsAlgorithmPerformance computing
The invention discloses a multi-thread redundancy acceleration method and system based on matrix operator splitting, and belongs to the technical field of high-performance computing and fault-tolerant computing, and the method comprises the steps: for a target large matrix operator, splitting the target large matrix operator into a plurality of sub-matrix blocks with uniform sizes by adopting a partitioning strategy, determining an operation dependency relationship and a data interaction rule of each sub-matrix block; constructing a weighted index method based on output sensitivity and error influence, and realizing quantitative evaluation of the criticality of the sub-matrix blocks; distributing redundant thread pairs for the high-criticality sub-matrix blocks, distributing single-thread execution for the low-criticality sub-matrix blocks, and realizing selective redundancy; after all the threads are executed, the calculation result of each sub-matrix block is recovered, and finally, the results of all the sub-matrix blocks after verification are integrated to obtain the final calculation result of the target matrix operator. According to the method, selective redundancy can be realized, the redundancy calculation overhead is reduced while the calculation reliability is guaranteed, and the overall calculation efficiency is improved.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Fast simultaneous feasibility test with electrical power grid perturbations over states and contingencies

Fast simultaneous feasibility testing (SFT) for management of an electrical power grid is achieved through various innovations. The computation problem relates to evaluation of candidate solutions for external power flows into a power grid, with respect to predetermined constraints and contingencies. A perturbation approach with precomputation is extended to encompass grid states (e.g. time periods) in addition to contingencies. Advantages derive from: fewer factorizations or inversions of large matrices; decoupling of state-dependent and contingency-dependent perturbations, leaving relatively few perturbations jointly dependent on both state and contingency; or making approximations by discarding small jointly dependent terms. Significant computation reductions allow a single workstation to perform SFT for 36 hours of a day-ahead cycle.
Owner:BATTELLE MEMORIAL INST

A high-efficiency simulation method for quantum system based on parallel reduction order

This invention relates to the field of traffic safety technology, specifically to an efficient simulation method for quantum systems based on parallel order reduction. The method includes the following steps: receiving the physical parameters of the quantum system and simulation requirements, whereby the physical parameters include electron mass, potential well size, and initial wave packet parameters; and the simulation requirements include simulation physical time and accuracy requirements. Based on the physical parameters, a time-dependent Schrödinger equation describing the dynamic behavior of the quantum system is established, and the Schrödinger equation is rearranged into matrix form using the Kronecker product. This invention utilizes Arnoldi to reduce the order of the matrix-form Schrödinger equation. By reducing the order, the large matrix in the original space is projected into a relatively small subspace, improving the efficiency of the simulation solution. The reduced-order Schrödinger equation is solved using a time-parallel algorithm. This approach overcomes the time step limitation imposed by the CFL condition, ensuring that a stable solution can be obtained with fewer time steps.
Owner:ANHUI UNIV

Method for iteratively creating a matrix from base elements

A method for iteratively generating a matrix of base elements that includes forming at least one base matrix, applying modification operators iteratively, which are able to transform the base matrix or any matrix arising therefrom into a modified matrix, applying expansion operators iteratively, which form a larger matrix from a plurality of optionally modified smaller matrices from the preceding iteration by copying, rotation or reflection by virtue of parts of the larger matrix being filled with the optionally modified smaller matrices, performing virtual experiments within the scope of which the properties of a created matrix are examined by systematic creation of values deliberately containing errors from which an error signal is derivable, and in order to create a complex matrix, forming all permutations of next larger matrices by applying expansion operators and then evaluating by means of virtual experiments, the next larger matrices selected with the smallest error signals, and then forming the next larger matrices successively therefrom by applying expansion operators and evaluating by means of virtual experiments until the error signal drops below a given limit.
Owner:POLDI MICROELECTRONICS GMBH

Hybrid analog-digital matrix processor

Techniques for computing matrix operations for arbitrarily large matrices on a limited size hybrid analog-digital matrix processor are described. Techniques for gain adjustment in a limited size hybrid analog-digital matrix processor are described that enable the system to achieve higher energy efficiency, greater physical density, and improved numerical accuracy. In some embodiments, these techniques maximize the prediction accuracy of GEMM-based convolutional neural networks using low-precision data representations.
Owner:LIGHT MATERIALS CO

Hardware adaptive multi-model scheduling

Modem deep neural network (DNN) models have many layers with a single layer potentially involving large matrix multiplications. Such heavy calculation brings challenges to deploy such DNN models on a single edge device, which has relatively limited computation resources. Therefore, multiple and even heterogeneous edge devices may be required for applications with stringent latency requirements. Disclosed in the present patent documents are embodiments of a model scheduling framework that schedules multiple models on a heterogeneous platform. Two different approaches, model first scheduling (MFS) and hardware first scheduling (HFS), are presented to allocate a group of models for a service into corresponding heterogeneous edge devices, including CPU, VPU and GPU. Experimental results prove the effectiveness of the MFS and HFS methods for improving the inference speed of single and multiple AI-based services.
Owner:BAIDU USA LLC +1

Method and apparatus for memory-aware stream processing for transformer-based foundation models

Embodiments of the present application provide systems and methods for performing matrix-based computations in computing devices, such as an edge computing device, which includes a MAC unit for matrix multiplication, a VEC unit for Softmax operations, and shared on-chip cache memory. The method involves parallel processing of matrix multiplication and Softmax operations to produce intermediate and final result tiles, which may improve memory usage and computational efficiency. Memory constraints are managed using selective overwriting strategies, allowing improved pipelined execution by dynamically reallocating on-chip memory. These techniques may enable reduced latency, improved throughput, and efficient handling of large matrices in machine learning applications, such as attention mechanisms. The method may further employ tiling and adaptive optimization routines to ensure computations fit within on-chip memory, which may reduce the need for off-chip memory access. Embodiments may be useful for AI or machine learning applications in resource-constrained environments.
Owner:HUAWEI TECH CO LTD

A MATRIX MULTIPLICATION ALGORITHM WITH TIME COMPLEXITY OF O(n 2), EXTENDABLE FOR SOLVING SYSTEMS OF LINEAR EQUATIONS AND IMPLEMENTABLE ON PARALLEL PROCESSING ARCHITECTURES

This invention introduces a parallel matrix multiplication algorithm achieving O(n²) time complexity through: 1. Binary Fusion Technique: Multi-layer bitwise data representation. Vectorized counting-logical operations (AND + popcount). Kronecker-based weighted reconstruction. 2. Key Advantages: Eliminates floating-point rounding errors. Reduces memory bandwidth requirements. Hardware-agnostic design (GPU / FPGA / ASIC compatible). 3. Applications: Machine learning acceleration. Cryptographic systems optimization. Large-scale scientific computing. The algorithm transforms numerical operations into parallel bitwise processes while maintaining full precision. Compared to conventional methods (e.g. Strassen), it demonstrates superior scalability for large matrices without recursive decomposition.
Owner:SADEGHI DANIAL

Rapid decoding method and decoding circuit of Transform decoder

The invention belongs to the technical field related to artificial intelligence, and discloses a rapid decoding method and decoding circuit of a Transform decoder. The fast decoding method comprises the following steps: selecting a plurality of spaced decoding layers as an exit judgment layer, performing reasoning calculation on a token coding vector through the decoding layers in sequence to obtain a reasoning vector, if the token coding vector reaches the exit judgment layer, performing early-quit judgment after the reasoning calculation of the exit judgment layer is completed, using an early-quit matrix for early-quit judgment, and using an early-quit matrix for early-quit judgment; the quit judgment layer judges whether the quit judgment layer can execute quit or not by judging the norm value of the difference between the reasoning vector of the quit judgment layer and the reasoning vector of the previous layer of the quit judgment layer relative to the premature exit matrix, if the premature exit condition is met, reasoning is ended, and otherwise, the quit judgment layer does not execute quit. And continuing to enter the next decoding layer. According to the method, fast and lightweight early-backward judgment in the decoding stage can be realized, and the decoding generation speed of the large model is remarkably improved.
Owner:HUAZHONG UNIV OF SCI & TECH

A MATRIX MULTIPLICATION ALGORITHM WITH TIME COMPLEXITY OF O(n 2), EXTENDABLE FOR SOLVING SYSTEMS OF LINEAR EQUATIONS AND IMPLEMENTABLE ON PARALLEL PROCESSING ARCHITECTURES

This invention introduces a parallel matrix multiplication algorithm achieving O(n²) time complexity through: 1. Binary Fusion Technique: Multi-layer bitwise data representation. Vectorized counting-logical operations (AND + popcount). Kronecker-based weighted reconstruction. 2. Key Advantages: Eliminates floating-point rounding errors. Reduces memory bandwidth requirements. Hardware-agnostic design (GPU / FPGA / ASIC compatible). 3. Applications: Machine learning acceleration. Cryptographic systems optimization. Large-scale scientific computing. The algorithm transforms numerical operations into parallel bitwise processes while maintaining full precision. Compared to conventional methods (e.g. Strassen), it demonstrates superior scalability for large matrices without recursive decomposition.
Owner:SADEGHI DANIAL

Method for iteratively generating matrix from base elements

The invention relates to a method for generating a matrix from basic elements (sensors, antennas, electronics). According to the invention, a base matrix is first formed, which contains at least all base elements once, modification operators, such as rotation and mirroring, are applied iteratively, by means of which the base matrix or a matrix derived from the base matrix can be converted into a modified matrix, expansion operators are applied iteratively, and the modified matrix is converted into a modified matrix, such as rotation and mirroring, by means of which expansion operators, such as rotation and mirroring, by means of rotation and mirroring, by means of rotation and mirroring. The expansion operator forms a larger matrix from a plurality of optionally modified matrices of previous iterations by copying, rotating or mirroring by filling portions of the larger matrix with optionally modified smaller matrices, performing a virtual experiment, and performing a virtual experiment. Wherein the properties of the generated matrix are researched by systematically generating pertinently erroneous values from which error signals can be derived, in order to generate a complex matrix, first all permutations of the next larger matrix are formed by an extension operator and evaluated by a virtual experiment, and then all permutations of the next larger matrix are evaluated by a virtual experiment. Subsequently, a next larger matrix with the minimum error signal is selected, and then the next larger matrix is gradually formed from these matrices by an expansion operator and evaluated by a virtual experiment until the error signal does not exceed a predetermined limit.
Owner:PURDY MICROELECTRONICS CO LTD

A three-dimensional ISAR imaging method based on multi-observation tensor sparse representation

The application discloses a three-dimensional ISAR imaging method based on multi-observation tensor sparse representation, and belongs to the field of inverse synthetic aperture radar imaging. In view of the fact that the existing ISAR imaging technology has low imaging resolution when data is insufficient and low imaging quality when noise is large, the application adopts a multi-observation method, utilizes the correlation between signals of multiple channels, and improves the inclusiveness of data. Compared with a traditional algorithm, the reconstruction effect is better under a small signal-to-noise ratio, the resolution of ISAR imaging is improved, a tensorization method is adopted to directly process three-dimensional data, an inverse matrix of an unfolded large matrix is not needed, the efficiency of the algorithm is improved, the three-dimensional structure information of the reconstructed signal is considered, the imaging accuracy is further improved, and the precision and quality of three-dimensional ISAR imaging are improved. The application realizes three-dimensional ISAR imaging based on multi-observation tensor sparse representation, is suitable for the field of inverse synthetic aperture radar imaging, and improves the imaging precision and imaging quality under a low signal-to-noise ratio.
Owner:BEIJING INST OF TECH

Time series data processing method

The invention relates to the technical field of artificial intelligence, in particular to a time series data processing method, an autocorrelation coefficient matrix is a matrix with a larger scale, the scales of the matrix are different along with different values of tau, for example, when tau is equal to 1, a sub-sequence Y1 is equal to x1, x2,..., xL, a sub-sequence Y2 is equal to x2, x3,..., xL + 1, and from the perspective of an original sequence X, the sub-sequence Y1 is equal to x1, x2,..., xL + 1; the difference between the positions of only initial elements of the two subsequences Y1 and Y2 is 1, and so on, when tau is taken as other values, the difference between the initial positions of the subsequences relative to the original sequence X is tau, autocorrelation coefficients among all the subsequences form an autocorrelation coefficient matrix, the autocorrelation coefficients are calculated by taking the autocorrelation coefficients as the core, and the difference between the initial positions of all the subsequences is 1; according to the method, future time sequence data can be predicted according to historical time sequence data, prediction can be performed by integrating multiple time spans, meanwhile, the relationship between the time sequence and historical small sequences of multiple time spans is considered, the minimum unit of a problem is considered to be a small sequence, and the historical small sequences of large time spans can be processed.
Owner:SHENZHEN QINGRONG ZHIHUI TECH CO LTD

A signal delay estimation method based on block processing

The application discloses a signal delay estimation method based on block processing, comprising the following steps: acquiring CSI data and preprocessing (validity test, noise filtering, selecting valid index); obtaining overlapping data blocks by block processing according to a preset block quantity and overlapping ratio; building Hankel matrix and enhancement matrix for each block, and obtaining delay estimation value through singular value decomposition, subspace separation and characteristic polynomial solving; assigning weight to each estimation value in combination with signal quality evaluation; and determining the final result based on the delay estimation value and weight information. In the application, a large matrix is disassembled through block overlapping, the singular value decomposition complexity and storage occupation are reduced, and lightweight hardware is adapted; the processing time consumption is shortened, the real-time demand in multiple scenes is met, the application can be run on ordinary equipment, and positioning deviation caused by delay is avoided.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

An OTA differential method based on size matrix cutting

The application discloses an OTA differential method and system based on size matrix cutting and a computer readable storage medium, and the method comprises the following steps: cutting source files and target files by using small matrices to obtain corresponding matrix data blocks, and the file data reading window size of the small matrix is a first window value; eigenvalues of the matrix data blocks of the source files and the target files are calculated respectively, source data feature tables and target data feature tables are constructed, the same matrix small blocks in the target files and the source files are filtered out by comparing the two feature tables; the positions of blank matrices in the target files and the matrix small blocks are marked respectively; the marked target files and the source files are rearranged, and the differential data is differentiated by using a large matrix, and the file data reading window size of the large matrix is a second window value. The application improves the differential data precision and can adapt to diversified upgrade data packets.
Owner:CHENGDU DESAY SV KAWA TECHNOLOGY CO LTD

Fast decoding method and decoding circuit for transformer decoder

This invention belongs to the field of artificial intelligence technology and discloses a fast decoding method and decoding circuit for a Transformer decoder. The fast decoding method includes: selecting multiple interval decoding layers as exit judgment layers; the token encoded vector sequentially passes through the decoding layers for inference calculation to obtain the inference vector; if an exit judgment layer is reached, an early termination judgment is performed after completing the inference calculation of that exit judgment layer; the early termination judgment uses an early termination matrix, which is equal to the product of the vocabulary matrix and its transpose; the exit judgment layer determines whether it can terminate by judging the norm of the difference between the inference vector of that layer and the inference vector of the previous layer relative to the early termination matrix; if the early termination condition is met, the inference ends; otherwise, it continues to the next decoding layer. This method can achieve fast and lightweight early termination judgment in the decoding stage, significantly improving the speed of large model decoding and generation.
Owner:HUAZHONG UNIV OF SCI & TECH

Systems and methods incorporating fast lipschitz constant estimation for neural networks

A compositional approach to estimating Lipschitz constants for deep feed-forward neural networks is disclosed herein. We first obtain an exact decomposition of the large matrix verification problem into smaller sub-problems. Then, leveraging the underlying cascade structure of the network, we develop two algorithms. The first algorithm explores the geometric features of the problem and enables us to provide Lipschitz estimates that are comparable to existing methods by solving small semidefinite programs (SDPs) that are only as large as the size of each layer. The second algorithm relaxes these sub-problems and provides a closed-form solution to each sub-problem for extremely fast estimation, altogether eliminating the need to solve SDPs. The two algorithms represent different levels of trade-offs between efficiency and accuracy.
Owner:PURDUE RES FOUND