Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

120 results about "Weighing matrix" patented technology

In mathematics, a weighing matrix W of order n and weight w is an n × n (0,1,-1)-matrix such that WWᵀ=wIₙ, where Wᵀ is the transpose of W and Iₙ is the identity matrix of order n. For convenience, a weighing matrix of order n and weight w is often denoted by W(n,w). A W(n,n) is a Hadamard matrix and a W(n,n-1) is equivalent to a conference matrix.t

Aircraft part production quality optimization method based on data analysis

The invention relates to the technical field of aircraft manufacturing, and discloses an aircraft part production quality optimization method based on data analysis. The method comprises the following steps: collecting multi-process real-time processing parameters and quality inspection data of a production line, and generating a dynamic quality weight matrix according to a process parameter coupling degree; extracting process path difference characteristics of qualified products and unqualified products in historical batches, and encoding the process path difference characteristics as a quality evolution chain matched with adjacent process parameter mutation relevance; optimizing process parameters at a quality evaluation node, driving a parameter combination to iterate to minimize quality fluctuation, generating an adjustment amount, updating an evolution chain constraint coefficient, and synchronously constructing a stability evaluation function of a correlation weight matrix and a defect propagation path; triggering process compensation according to an adjustment amount gradient, and verifying the cohesion through a tolerance rule by taking a reference parameter matched with the target evolution chain as compensation data; and generating a feedback matrix by using the quality fluctuation index and the compensation result, and correcting the mapping relation between the weight matrix and the evolution chain.
Owner:CHENGDU SEN BO PRECISION MASCH CO LTD

Calculation device, calculation method, and program

The invention relates to a calculation device, a calculation method, and a program. This computing device (10) computes a matrix product of a sparse matrix and a weight matrix in a neural network, the sparse matrix having a structure in which a predetermined number of non-zero elements are included in each block of a predetermined size. The sparse matrix is defined by a local index representing the position of the non-zero element in each block and the value of the non-zero element, and the arithmetic device (10) performs: a process for acquiring, from the weight matrix, an element corresponding to the position of the non-zero element in the sparse matrix by referring to the local index; and multiplication and accumulation operation of non-zero elements of the sparse matrix and corresponding elements of the weight matrix is carried out.
Owner:DENSO CORP

Apparatus, data structure and method for adjusting weights of neural networks of models

Apparatus, data structure and method of adjusting weights of a neural network of a model, processing inputs of the model and outputting outputs of the model, in which the method comprises: providing training data, the training data comprising the inputs of the model and reference truth values for the outputs of the model, the reference truth values corresponding to the inputs of the model in the training data; providing an adjustment method set for adjusting the weight; determining principal component decomposition of a weight matrix including weights; determining a feature value of a covariance matrix corresponding to the feature vector; rearranging the eigenvectors in the matrix in an order that produces a monotonically decreasing order of eigenvalues associated with the eigenvectors; rearranging the weights in the weight matrix according to the rearrangement sequence of the feature vectors in the matrix; dividing the weight matrix into groups of weights; associating at least one of the groups with an adjustment method selected from the set of adjustment methods; and adjusting the weights in the at least one group on the training data using an adjustment method.
Owner:ROBERT BOSCH GMBH

Feature selection method and data reconstruction method for shoreline image feature selection

The invention discloses a feature selection method and a data reconstruction method for shoreline image feature selection. The method comprises the following steps: S1, integrating original data into a sample data set matrix X; s2, initializing a reconstruction transformation matrix W, a transformation matrix Q, an eigenvalue regression coefficient matrix A, a coordinate basis matrix B, a coding matrix E, an auxiliary matrix F and an auxiliary matrix D as unit matrixes, initializing a weight vector p and a mean vector v of a data set as unit vectors, and initializing a weight matrix P = diag (p); and S3, updating the coding matrix E, the auxiliary matrix F and the like based on the data set matrix X, the mean vector v of the data set, the transformation matrix Q, the eigenvalue regression coefficient matrix A, the coordinate basis matrix B and the weight matrix P. According to the invention, the accuracy of data reconstruction and the validity of feature selection can be improved.
Owner:THIRD INSTITUTE OF OCEANOGRAPHY STATE OCEANI C ADMINISTRATION

Weight data processing method of neural network model, electronic equipment and storage medium

The invention provides a weight data processing method of a neural network model, electronic equipment and a storage medium, and relates to the technical field of machine learning, and the method comprises the steps: obtaining a weight matrix of a trained neural network model; determining a scaling factor and a zero offset of the weight matrix according to the maximum weight value and the minimum weight value in the weight matrix; based on the scaling factor and the zero offset, performing quantization processing on the weight matrix to obtain a quantized weight matrix; and splitting the quantized weight matrix into a preset number of low-rank matrixes. In the embodiment of the invention, the weight matrix is quantized based on the scaling factor and the zero offset, so that the high precision of the matrix quantization process can be ensured. Through quantization processing and matrix splitting of the weight matrix, dual compression is realized, the compression rate of the weight matrix can be improved, and the storage space of the weight matrix is reduced.
Owner:ALIBABA CLOUD COMPUTING CO LTD

Simulation processing system

A simulation system and method for implementing a model based on an iterative neural network, the system comprising: a simulation vector-matrix multiplication circuit that encodes a weight matrix of the model based on the iterative neural network; and an analog non-linear circuit that encodes a non-linear function arranged in a feedback loop configured to return an output signal from the non-linear circuit as input to the vector-matrix multiplication circuit, wherein the system is configured to output a solution vector of values of the model based on the iterative neural network upon convergence of the system.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Spatial intelligent world modeling method and system based on global NURBS parameter domain

The invention belongs to the technical field of space intelligent modeling, and relates to a space intelligent world modeling method based on a global NURBS parameter domain. Consistent mapping and management are carried out on the multi-Patch geometric objects in a global NURBS parameter domain; generating and optimizing a control point matrix and a weight matrix for geometric expression based on a global NURBS parameter domain; realizing automatic smooth transition of the multi-Patch geometry in the splicing area based on a continuity keeping mechanism of curvature and normal constraint; executing NURBS curved surface subdivision operation according to the curvature change rate and the geometric error threshold value; and performing global solution on the control point matrixes of all the Patches, and outputting a space intelligent world model which is continuous and differentiable in a global range and has consistent parameters. According to the method, the continuity and controllability of geometric modeling are remarkably improved. The invention further provides a spatial intelligent world modeling system based on the global NURBS parameter domain.
Owner:BEIJING FEIDU TECH CO LTD

A long-distance slag conveying numerical self-adaptive control method based on CFD

The application provides a long-distance slag conveying numerical self-adaptive control method based on CFD, relates to the technical field of slag conveying pipeline analysis, and comprises the following steps: constructing a pipeline real-time data set; generating a calculation grid; combining a mathematical model and a dynamic boundary condition to construct a CFD dynamic simulation system of a pneumatic conveying pipeline; performing solution calculation on the CFD dynamic simulation system to obtain a flow parameter distribution model in the slag conveying pipeline; performing feature extraction on the pipeline real-time data set and the flow parameter distribution model, mining the correlation between features and fault types through a weight matrix and a cross-term weight matrix, obtaining a multi-modal fault diagnosis judgment result, and generating a slag conveying pipeline self-adaptive control strategy. The technical problems that the simulation accuracy is limited, it is difficult to support high-reliability design, a multi-physical field data fusion mechanism is missing, accurate fault diagnosis cannot be achieved, dynamic self-adaptive analysis capability is lacked, and the timeliness of fault early warning and regulation strategies is insufficient in the prior art are solved.
Owner:LIAONING GUOYUAN ELECTRIC POWER TECH CO LTD

Storage method and device for weight matrix

The invention relates to a storage method and device for a weight matrix. The invention provides a weight matrix storage method, which comprises the following steps that: a high-bit weight matrix and a low-bit weight matrix which are subjected to generalized sparse processing are obtained from a large language model, the positions of elements in the high-bit weight matrix and the low-bit weight matrix are in one-to-one correspondence, and the positions of elements in the high-bit weight matrix and the positions of elements in the low-bit weight matrix are in one-to-one correspondence; the positions of the valid data elements in the high bit weight matrix and the positions of the valid data elements in the low bit weight matrix have complementarity; generating a mask identifier list according to the positions of the valid data elements in the high bit weight matrix; combining the high-bit weight matrix and the low-bit weight matrix into a combined weight matrix according to complementarity; and storing the merge weight matrix and the mask identifier list.
Owner:MOFFETT AI TECHNOLOGY SHENZHEN CO LTD

A model structured pruning method and device, and a computing device cluster

PendingCN122366561AAlgorithmWeighing matrix
A structured pruning method for models includes: obtaining a first model; performing structured pruning on a first pruning unit in the first model to obtain a second model, wherein the first pruning unit is obtained by coupling the rows / columns of the weight matrix of a first operator and the rows / columns of the weight matrix of a second operator in the first model; the input of the second operator includes the output activation values ​​of the first operator, and the dimension of the output activation values ​​of the second operator in the first model is the same as the dimension of the output activation values ​​of the second operator in the second model; or the input of a third operator in the first model includes the output activation values ​​of the first operator and the output activation values ​​of the second operator, and the dimension of the output activation values ​​of the third operator in the first model is the same as the dimension of the output activation values ​​of the third operator in the second model. Thus, by using a row (or column) column-corresponding pruning strategy on the weight matrices of multiple operators in a serial or parallel manner in the model, structured pruning of the model is achieved without changing the hidden dimensions of the pruned model.
Owner:HUAWEI TECH CO LTD

Method for reasoning by using large language model, control device and storage medium

The invention provides a method for reasoning by using a large language model, a control device and a storage medium, and belongs to the technical field of large language models. The method comprises the steps that in the reasoning process of a target large language model, a weight matrix and input data of the target large language model are segmented into a plurality of data blocks; sequentially loading each data block in the plurality of segmented data blocks into a memory, and carrying out parallel optimization calculation on each data block; and combining the calculation result corresponding to each data block to obtain a reasoning result. Through block parallel computing, the reasoning performance of the large language model in a CPU environment can be remarkably improved, and an efficient and flexible solution is provided for large language model reasoning in a resource-constrained environment.
Owner:HYGON INFORMATION TECH CO LTD

Lightening method and device for shared weight of similar channels of large model

The invention discloses a large model similar channel sharing weight lightening method and device, and belongs to the field of artificial intelligence. The method comprises the steps of performing K-Means clustering on a Q / K projector weight matrix of a self-attention layer of a Transform architecture large model according to columns, and generating a cluster representative vector and a cluster label list to replace an original weight matrix to complete compression; restoring the approximate weight matrix based on the representative vector and the cluster label, and constructing an error matrix of the original matrix and the approximate matrix; performing singular value decomposition on the error matrix and retaining a key singular value to obtain two low-dimensional matrixes; during reasoning, multiplying the two low-dimensional matrixes to obtain an approximate error matrix, and adding the approximate error matrix with the approximate weight matrix to obtain a final reasoning weight; the total storage size after compression is calculated through a unified formula, and accurate quantification of storage overhead is achieved. According to the method, the performance degradation under low-bit compression is effectively relieved, the resource demand of model deployment is reduced, and the method is adaptive to various mainstream large models and different-dimension weight matrixes.
Owner:NAT UNIV OF DEFENSE TECH

Local perception lookup table lookup method

The invention relates to the field of data lookup, and discloses a locality perception lookup table lookup method, which comprises the following steps of: S1, reordering between rows: clustering and reordering the rows of an input vector and a weight matrix according to the same quantized value to enable elements with the same quantized value to be continuously distributed; for a common matrix GeMV scene, cache locality is improved through rearrangement and subsequent blocking operation, and repeated loading of a lookup table LUT is reduced. And the data locality of the GeMV scene is remarkably improved through inter-line reordering and a virtual mapping strategy. On the basis of reordering of input vector element values, elements with the same value are aggregated, so that the LUT line loading times are reduced to 256 times at most from being in direct proportion to the input dimension d, and frequent replacement of the LUT lines in the WRAM is avoided; meanwhile, a weight matrix is divided into BR * BC sub-matrixes, the characteristic that vertical adjacent rows share an accumulator is utilized, multi-row sub-matrix calculation is loaded at a time, and virtual rearrangement maintains the row matching relation between an input vector and the weight matrix through index mapping.
Owner:RENMIN UNIVERSITY OF CHINA

Large model reasoning acceleration method and device for grouping perception quantification and residual error correction

The invention relates to the technical field of artificial intelligence model optimization, in particular to a large model reasoning acceleration method and device for packet sensing quantization and residual correction, and the method comprises the steps: carrying out the statistical analysis of the weight and activation of each layer of a large model, and generating a channel feature matrix; constructing a learnable grouping mapping matrix, dividing channels into different groups, and dynamically distributing quantized bit width to obtain a grouping weight matrix; on the basis of the channel weight-activation joint sensitivity, calculating each grouping error contribution on line by using a small prediction model, and adjusting a grouping weight matrix to generate an optimized weight matrix; constructing an error control matrix to dynamically adjust the quantization error along the propagation path, and generating a correction matrix; dynamically adjusting the sparse rate and the quantization precision according to the channel feature matrix and the hardware constraint by combining a structured sparse strategy, and generating a sparse quantization matrix; in the reasoning process, model reasoning is carried out according to the correction matrix and the sparse quantization matrix, and the weight matrix and the correction matrix are updated and optimized in a closed-loop mode.
Owner:HENAN TECHN COLLEGE OF CONSTR

Model training method and device, computer equipment, readable storage medium and program product

The invention relates to a model training method and device, computer equipment, a computer readable storage medium and a computer program product. The method relates to an artificial intelligence technology. The method comprises the following steps: adding a compressor to a weight matrix in a teacher model to obtain a student model; obtaining a first output layer vector related to the training sample from the teacher model, and obtaining a second output layer vector and a prediction result related to the training sample from the student model; constructing a knowledge distillation loss function according to the first output layer vector, the second output layer vector, the prediction result and the annotation information, and obtaining a penalty term gradient according to each element of the compressor; selecting a target dimension from the compressor as a penalty term, updating a value corresponding to the target dimension by using a penalty term gradient, and updating a value corresponding to a non-target dimension by using a knowledge distillation loss function; after training is completed, the trained student model is obtained according to the compressor and the weight matrix, and the convergence effect and the convergence speed of the student model can be guaranteed.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Big language model binarization quantification method and system based on instructive alternate optimization

PendingCN122065892AHigh precisionReduce quantization errorComputer simulationsLinguistic modelAlgorithm
The invention provides a large language model binarization quantification method and system based on instructive alternate optimization, and relates to the technical field of large language model deploying.The method comprises the steps that a weight matrix is divided into an important area and an unimportant area by evaluating the influence degree of all weight parameters on the performance of a large language model; carrying out binaryzation on all areas of the weight matrix, and alternately optimizing a row vector scaling factor and a column vector scaling factor by adopting a first-order row-column alternate optimization iteration mode to obtain a first-order reconstruction weight matrix; and carrying out binaryzation again on the important region of the first-order reconstruction weight matrix, carrying out optimization by adopting a first-order and second-order row-column alternating optimization iteration mode to obtain a second-order reconstruction weight matrix, and correspondingly taking the second-order reconstruction weight matrix as a weight parameter after the large language model is quantized. The quantization error of the key weight parameter can be effectively reduced, the quantization precision is high, the quantization process completely depends on the internal structure information of the large language model, and the edge deployment compatibility is high.
Owner:GUANGDONG UNIV OF TECH

Method suitable for multi-energy complementary linkage economy evaluation

The invention relates to the technical field of multi-energy complementary energy systems, in particular to a method suitable for multi-energy complementary linkage economy evaluation, which comprises the following steps: S1, constructing a three-dimensional dynamic index system, and designing a time sequence weight matrix based on the peak-flat-valley period operation difference of an energy system; s2, according to the time sequence weight matrix, weight pre-judgment adjustment is conducted on the three-dimensional dynamic index system through LSTM, and a real-time dynamic index system is output. Through the combined design of the three-dimensional dynamic index and the time sequence weight, the total cost and total income of multi-energy linkage are comprehensively covered; the problem that a traditional static index cannot adapt to a dynamic linkage scene is effectively solved, and the scene adaptability is improved; a real-time feedback correction mechanism is introduced through a double-stage evaluation algorithm, the evaluation error rate is reduced conveniently, multi-energy complementary systems of different scales and different energy combinations can be flexibly adapted, only basic cost parameters and a weight matrix need to be adjusted, and the application range is widened.
Owner:南方电网能源发展研究院有限责任公司

Carbon footprint accounting method, system, device and medium based on optimal allocation mode

PendingCN122366846ACarbon footprintData mining
This invention provides a carbon footprint accounting method, system, device, and medium based on optimal allocation, including: acquiring accounting data and a historical score matrix, and determining a factor score set based on the accounting data; acquiring a weight matrix corresponding to all allocation methods, adaptively optimizing the weight matrix based on the accounting data to obtain an optimized weight matrix, and then calculating a comprehensive practicality score corresponding to each allocation method based on the optimized weight matrix and the factor score set; calculating a selection threshold based on the historical score matrix, allocating the accounting data using allocation methods with a comprehensive practicality score greater than the selection threshold, and performing vehicle carbon footprint accounting based on the allocation results. This invention solves the problem in the prior art where it is difficult to determine the most suitable data allocation method to be used in the carbon footprint accounting process, resulting in an inaccurate final carbon footprint accounting structure.
Owner:CHINA AUTOMOTIVE ENG RES INST

Device and method of next token prediction

A computer-implemented method of predicting a next token given a sequence of tokens by a transformer-based machine learning system. The method includes: processing an embedding of the sequence of tokens through multiple layers; determining a confidence score for individual tokens at a layer's output by multiplying the layer's output embedding with a weight matrix; at a predetermined pth layer, determining the K most likely next tokens; creating a pruned weight matrix by removing rows corresponding to other tokens; for predefined subsequent layers, determining a confidence score for only the K tokens using the pruned weight matrix; and returning the token with the highest confidence score exceeding a layer-specific threshold.
Owner:ROBERT BOSCH GMBH

A deep fake detection method and system based on orthogonal subspace decomposition and hyperspherical metric

This application belongs to the interdisciplinary field of artificial intelligence, computer vision, and network information security. It discloses a deepfake detection method and system based on orthogonal subspace decomposition and hyperspherical metric. By applying singular value decomposition to the weight matrix of a pre-trained visual model, it explicitly constructs a frozen principal subspace that preserves general semantic knowledge and a trainable orthogonal residual subspace that captures specific forgery traces, achieving orthogonal isolation of the parameter space. Simultaneously, hyperspherical metric learning is introduced into the feature space, performing L2 normalization on the features and applying alignment and uniformity losses. Combined with spherical linear interpolation, latent space data augmentation is performed while preserving the Riemannian geometric structure. Through the synergistic constraints of the parameter and feature spaces, this application can reduce the interference of fine-tuning on pre-trained general visual knowledge and improve the feature discrimination stability and cross-forgery generalization ability in deepfake detection tasks.
Owner:NANJING UNIV OF POSTS & TELECOMM

A general information extraction method based on task-specific hybrid low-rank adaptation

ActiveCN121168579BBiological modelsAlgorithmMatrix group
The application provides a general information extraction method based on task-specific mixed low-rank adaptation, comprising the following steps: selecting a pre-trained model as a base model and initializing the base model; initializing a first matrix group and a second matrix group for a weight matrix of the base model; performing forward propagation calculation on all low-rank decomposition matrices; constructing a gating matrix, fusing features and a fourth output vector to generate a fifth output vector through vector splicing; constructing a joint loss function for optimization; combining the matrices to generate a target model; inputting a natural language text information task into the target model to output a structured information extraction result. The application has the beneficial effects of effectively alleviating the task interference and model over-coupling problems of the existing multi-task learning method, thereby improving the performance of the model in the general information extraction task; integrating multiple low-rank matrices and gating matrices to improve the adaptability of the model to different task characteristics and ensure effective cross-task knowledge transfer.
Owner:TIANJIN JIZHI TECH CO LTD +1

Data processing method and related device

The invention discloses a data processing method and a related device, and relates to the field of big data services, and the method comprises the steps: selecting a variation strategy matched with the number of iterations from a preset variation strategy set according to the number of iterations of a first population of a prediction model, and performing variation processing on individuals in the first population by using an optimization factor in a variation strategy matched with the number of iterations to obtain a second population, and adjusting a weight matrix of the prediction model by using the second population under the condition that the second population meets a model storage condition, thereby reducing parameter adjustment frequency. Wherein the value of the optimization factor in each variation strategy is changed along with the change of the number of iterations, so that variation processing on the individuals by different number of iterations is different, and the diversity of the second population under different number of iterations is increased, so that the second population can be adaptive to different training stages of the prediction model; the robustness and precision of the prediction model are improved, and the stability of the prediction model is further improved.
Owner:AGRICULTURAL BANK OF CHINA

Video decoding method, video encoding method and apparatus

The application discloses a video decoding method, a video encoding method and a device. The video decoding method comprises the following steps: generating a weight matrix of a current block; constructing a motion information candidate list of the current block; adjusting the motion information candidate list based on a template region of the current block, and determining first motion information and second motion information of the current block; performing weighted prediction on the current block by using the weight matrix, the first motion information and the second motion information, so as to obtain a prediction value of the current block; and obtaining a decoding result of the current block based on the prediction value of the current block. The application can improve the coding and decoding efficiency.
Owner:ZHEJIANG DAHUA TECH CO LTD

Model weight quantification method, electronic device and program product

ActiveCN121351913ABiological modelsKnowledge based modelsLinguistic modelGreedy optimization
The invention provides a model weight quantification method, electronic equipment and a program product, and relates to the technical field of computers. The quantification method of the model weight comprises the following steps: carrying out partitioning processing on a full-precision weight matrix of a large language model, and calculating an importance score of each weight block; according to the importance score of each weight block, distributing a quantization bit width for each weight block through a greedy optimization algorithm; performing amplitude sharing processing on the full-precision weight matrix, and optimizing quantization parameters according to Hessian attributes of weight elements in the full-precision weight matrix after the amplitude sharing processing to obtain the full-precision weight matrix after elimination of discretely distributed abnormal values and the optimized quantization parameters; and calculating the difference between the full-precision activation output and the quantization activation output according to the quantization bit width, the full-precision weight matrix after eliminating the discretely distributed abnormal values and the optimized quantization parameters, and performing hierarchical feedback denoising compensation to obtain a target quantization weight matrix.
Owner:XIAMEN UNIV

Beidou multi-frequency positioning method and device, electronic equipment and storage medium

PendingCN122043508ASatellite radio beaconingDesign matrixAlgorithm
The invention provides a Beidou multi-frequency positioning method and device, electronic equipment and a storage medium. The method comprises the following steps: S1, collecting observation data and satellite ephemeris data of two target frequency points of a Beidou satellite; s2, based on the satellite ephemeris data, modeling the Beidou satellite clock correction by adopting a quadratic polynomial to obtain a clock correction modeling error, and taking the clock correction modeling error as a constraint term of an original design matrix to generate a design matrix with clock correction constraint; s3, constructing an observation vector based on the observation data, and substituting the design matrix with the clock error constraint and the observation vector into a TLS resolving model for resolving to obtain a positioning parameter initial value and a corresponding resolving residual error; s4, fitting a mapping relation between the residual error and the noise intensity based on the resolving residual error, generating a weight matrix, and performing weighted updating on the TLS resolving model by using the weight matrix to obtain an updated TLS resolving model; according to the embodiment of the invention, effective constraint on the design matrix error can be realized, and the clock error modeling precision is improved.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD

Weight matrix processing method and device, equipment, storage medium and product

PendingCN121832884AUnpack smoothlySmooth dequantizationDigital data processing detailsComplex mathematical operationsProcess engineeringTarget weight
The invention discloses a weight matrix processing method and device, equipment, a storage medium and a product. The method comprises the steps that a target weight matrix is loaded; wherein the target weight matrix is a third-precision weight matrix which is arranged in a preset format and is formed by packaging a first-precision weight matrix; carrying out unpacking and inverse quantization processing on the target weight matrix to obtain a weight matrix with second precision; wherein the third precision is higher than the second precision, and the second precision is higher than the first precision; and storing the weight matrix with the second precision to an on-chip storage unit so as to be loaded to a tensor core to execute matrix multiplication operation. According to the embodiment of the invention, the weight matrix can be compatible with the operation logic of tcore, so that the matrix multiplication operation can be smoothly completed in the tcore.
Owner:SHANGHAI BIREN TECH CO LTD

Inference acceleration method and device for neural network model, medium and product

The invention provides a reasoning acceleration method and device for a neural network model, a medium and a product. A reasoning acceleration method for a neural network model is characterized by comprising the steps that before the neural network model is deployed, an original weight matrix of a to-be-accelerated target network layer in the neural network model is recognized, and the size of the original weight matrix is n * m; the original weight matrix is decomposed into a first low-rank matrix and a second low-rank matrix through a low-rank decomposition algorithm, the size of the first low-rank matrix is n * k, the size of the second low-rank matrix is k * m, and k is smaller than the minimum value of n and m; and in a reasoning stage of the neural network model, for the input data, executing the following operations: executing matrix multiplication of the input data and the first low-rank matrix to obtain an intermediate result; and performing matrix multiplication of the intermediate result and the second low-rank matrix to obtain a final output result.
Owner:MOFFETT AI TECHNOLOGY SHENZHEN CO LTD

Intelligent suspension robust control and state observation method based on LMI

The invention discloses an LMI-based intelligent suspension robust control and state observation method, and relates to the technical field of suspension control, and the method comprises the following steps: S1, splitting a control system into a non-matching disturbance subsystem and a matching disturbance subsystem; s2, determining state feedback control of an LMID algorithm according to the weight matrix of state control; determining disturbance compensation of a disturbance observer part in an LMID algorithm based on a weight matrix of a matching disturbance subsystem and disturbance estimation, and realizing the LMID algorithm by using an LMIob observer; and S3, taking the sprung acceleration based on the IMU in the suspension system as the input quantity of the LMIob observer, and obtaining state estimation of the control system. The method is designed based on an LMID algorithm of LMI, the algorithm has disturbance observation properties, the robustness of the algorithm can be improved, a weight matrix is introduced, and flexible adjustment of state control and disturbance estimation is achieved.
Owner:BEIJING INST OF TECH

Method and system for matrix reduction in large language models

The disclosure is directed toward a method to efficiently perform a large language model operation by reducing the memory and computation resources for operating on a matrix such as a weight matrix. The weight matrix may be included in operations such as the head attention function of the large language model. The weight matrix is decomposed into multiple submatrices that each include a set of weights from the weight matrix. An input of the large language model is provided. A submatrix solution for the submatrix is determined from the input and the set of weights of the submatrix. The determining of a solution is repeated for each of the submatrices to provide multiple submatrix solutions. The resulting submatrix solutions are combined to obtain a matrix solution for the weight matrix.
Owner:CORNAMI INC