Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

79 results about "Quantification methods" patented technology

Quantization method and reasoning method of large language model and electronic equipment

The invention discloses a quantification method and reasoning method of a large language model and electronic equipment, and belongs to the technical field of large language models.The quantification method of the large language model comprises the steps that each linear layer to be quantized in the large language model is quantized; dividing a channel of the linear layer in the hidden layer dimension into a normal channel and an outlier channel; performing INT8 quantization on the first activation matrix corresponding to the normal channel in a word segmentation token dimension to obtain a second activation matrix, and performing INT4 quantization on the first weight matrix corresponding to the normal channel according to an output channel to obtain a second weight matrix; and determining an output result of the linear layer according to the second activation matrix, the second weight matrix, a third activation matrix corresponding to the outlier channel and a third weight matrix corresponding to the outlier channel.
Owner:ZTE CORP

Quantization method of large language model, related equipment and computer program product

The invention provides a large language model quantification method, related equipment and a computer program product, and the method comprises the steps: carrying out the reasoning of a to-be-quantized large language model through calibration data, obtaining the activation of the large language model, and carrying out the statistics of the activation distribution of a target layer according to channels; calculating a smoothing factor of each channel according to the activation distribution; compensating the weight of the target layer channel by channel according to the smoothing factor to obtain a compensated weight; performing 4-bit quantization on the compensated weight; in the model reasoning process, activation of a target layer is smoothed, and 16-bit quantization is carried out on the smoothed activation. According to the method, a W4A16 quantification scheme is adopted, compared with W8A8, the weight storage amount is compressed by half, meanwhile, the precision loss is controlled to be smaller than 3%, and deployment of a large language model on edge equipment is facilitated.
Owner:SHANGHAI ZHICHEN MICRO TECHNOLOGY CO LTD

Self-adaptive mixing precision quantification method, device, equipment and medium

The invention relates to a self-adaptive mixing precision quantification method and device, equipment and a medium, and the method comprises the steps: carrying out the comprehensive sorting through the product of the cosine similarity difference value of adjacent layers and the sensitivity weight of each layer, so as to guarantee that bottleneck layers which are liable to be influenced by quantification and are crucial to the final precision can be accurately recognized; therefore, precise positioning of a protection target is realized, an iterative optimization loop is introduced, an optimal solution meeting a preset performance target can be spontaneously found finally by continuously evaluating a time-precision balance point of an overall model and automatically adjusting configuration, and a suboptimal result caused by improper primary configuration is avoided; and the input of the user is simplified into a visual final performance target, and the complicated layer sorting and selection process is automatically processed in the system, so that the use threshold of the technology is reduced.
Owner:HUNAN GREAT WALL GALAXY TECH CO LTD

Quantization methods for gnb-driven multi-vendor sequential training

Method and apparatus for quantization of base station driven multi-vendor sequential training. The apparatus generates an encoder output by inputting an input CSI to a reference encoder. The apparatus quantizes the encoder output by inputting the encoder output to a quantizer to generate a quantizer output. The apparatus trains a decoder of the network entity based at least on the quantizer output to generate a training dataset. The apparatus outputs a training dataset indication comprising the training dataset to a UE, the training dataset indication comprising at least the input CSI. The apparatus communicates with the UE using the trained decoder.
Owner:QUALCOMM INC

Enterprise credit risk quantification method

The invention relates to the technical field of financial science and technology, in particular to an enterprise credit risk quantification method, which comprises the following steps of: writing finance, Internet of Things, bill chains and remote sensing data characteristics into a data lake in a homomorphic encryption manner to generate a federated event grid; a causal topology hypergraph is constructed, topology and multi-scale frequency domain features are extracted to form node representation, and node states and weights are updated through fractional order graph differential; the mapping hypergraph is a silicon light Mach-Zehnder phase, and node state superposition impact pulse is injected into a photon interference network to obtain optical readout; the spiking neural network-reinforcement learning agent reads out the generation management action to form a reserve tensor; and solving a discrete Schrodinger equation in combination with the tensor and graph Laplacian, and outputting a risk entropy potential and a prediction period default probability confidence interval through fractional order path integral correction. According to the invention, privacy protection, high-order structure identification and long-tail sensitivity are realized, and real-time and auditable credit risk assessment is realized.
Owner:SHANGHAI BEITONG ENTERPRISE CREDIT INVESTIGATION CO LTD

Quantization method and quantization device

The invention provides a quantization method and a quantization device, which can perform singular value decomposition on a first matrix formed by word embedding vectors to obtain the maximum characteristic value of each block matrix in M block matrixes, and can determine the quantity of quantization bits of each block matrix according to the maximum characteristic value of each block matrix. The method comprises the following steps of: dividing each block matrix into a plurality of block matrixes, quantifying each block matrix to obtain each quantized block matrix, combining each quantized block matrix to obtain a quantized first matrix, and representing the amount of semantic information carried in each block matrix by a maximum characteristic value of each block matrix to a certain extent, according to the embodiment of the invention, the quantized bit number can be allocated to each block matrix according to the amount of semantic information carried by each block matrix, so that different bit numbers can be allocated to the block matrixes carrying different semantic information, differential bit number allocation can be realized, semantic information loss can be reduced, and user experience can be improved.
Owner:HUAWEI TECH CO LTD

Quantification device, quantification method and quantification program

To provide a quantification device, a quantification method and a quantification program capable of quantitatively evaluating uncertainty of a response in a system using retrieval extension generation and a generation language model.SOLUTION: A quantification device 1 includes a first index calculation unit 11 for calculating a first index that shows uncertainty of an output generated to an input to a generation language model and does not consider retrieval quality in retrieval extension creation, a second index calculation unit 12 for calculating a second index showing the retrieval quality in the retrieval extension generation, an index adjustment unit 13 for normalizing the second index and adjusting it as an influence degree to an output of the retrieval quality, and an output unit 14 for multiplying the first index by an adjusted value and outputting it as a value obtained by quantifying the uncertainty of the output in the generation language model using the retrieval extension generation.SELECTED DRAWING: Figure 1
Owner:KDDI CORP

High-efficiency INT6 quantification method, device and equipment for large language model

The invention discloses an efficient INT6 quantification method, device and equipment for a large language model, and the method comprises the steps: carrying out the mixing precision quantification of the large language model, and obtaining a quantized large language model; performing bit-level data packaging on the weight and the activation value in the quantized large language model to obtain bit-level data; loading the bit-level data to a register of the GPU, and carrying out matrix product accumulation operation and weighted summation by utilizing BTC to obtain output data; and storing the output data back to a global memory of the GPU so as to complete the quantitative reasoning process of the large language model. According to the method, mixed precision quantification is carried out on a large language model by utilizing different precisions, the reasoning speed is improved through a scheduling strategy of the GPU while relatively high quantification precision is achieved, and the reasoning potential of the GPU is fully mined, so that all potential of 6-bit quantification is released.
Owner:XIDIAN UNIV

Quantization and inverse quantization method in large language model and neural network processor

The invention relates to a quantization method, an inverse quantization method and a neural network processor in a large language model. A core (namely a quantization unit) in a neural network processor is arranged to execute quantization processing in a large-scale language model reasoning process, and the quantization unit has a data partitioning function, so that the quantization processing efficiency is greatly improved compared with an existing quantization processing core which is limited by the data size (such as block wise) when receiving data to be quantized. And frequent interaction with a cache is not needed, so that online high-efficiency large-model quantification processing is realized. And performing an inverse quantization process in a large language model inference process by setting a core (i.e., an inverse quantization unit) in a neural network processor, and the inverse quantization unit having a data partitioning function, a matrix multiplication function, and a high precision accumulation function (e.g., multiplying by a corresponding quantization parameter), the inverse quantization processing does not need to frequently carry data in a plurality of cores, and the inverse quantization processing efficiency of a large model is improved.
Owner:北京凌川科技有限公司

Processing apparatus and quantization method for quantization and inverse quantization of numeric data

The invention relates to a processing device and a quantization method for quantization and inverse quantization of numeric data. In one or more aspects, a processing apparatus for numeric data quantization includes processing circuitry to determine a maximum exponent from a set of exponents of a set of digit representations of a set of digits, obtain a set of scaled exponents based on the maximum exponent, and quantize the scaled exponents based on the set of scaled exponents. And one of (i) obtaining a set of quantized significant numbers based on the set of digit representations and a set of mantissas of the set of scaled exponents, or (ii) obtaining a set of quantized mantissas based on the set of mantissas. The processing circuitry is configured to output a set of quantized digit representations of the set of digits based on the set of quantized significant numbers or based on the set of quantized mantissas and the set of scaled exponents; and outputting an offset exponential scaling factor based on the maximum exponent.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Method and system for quantizing large-scale voice model after ultra-low training

The invention discloses a quantizing method and system after ultra-low training of a large voice model, and relates to the field of voice models. Comprising the steps of weight matrix preprocessing, K-means clustering quantification, abnormal value detection and mixing precision distribution, selective abnormal value retention, model reasoning and performance evaluation. According to the method, dynamic precision distribution is realized, quantization bits are dynamically adjusted according to abnormal value density, and the resource utilization rate is optimized; key information is reserved, abnormal values are selectively reserved, quantization errors are avoided, and the model performance is improved; the generalization ability of the quantitative model is ensured through the cross-domain robustness; and an efficient and flexible solution is provided for ultra-low bit quantization.
Owner:SHANGHAI JIAOTONG UNIV

System using transformer architecture with quantization-aware non-linear approximation and near-memory computing

This invention proposes a GQA-LUT method, utilizing a genetic algorithm and LUT-based circuit to efficiently approximate non-linear operators in Transformers. It adaptively finds optimal solutions for various non-linear functions, outperforming conventional neural network methods. A novel rounding mutation (RM) algorithm enhances approximation accuracy during quantization, improving low-bit integer precision. The invention also introduces a LayerNorm folding strategy as a near-memory computing principle, reducing IO and energy overheads with a two-stage memory hierarchy. Additionally, an additive partial sum quantization method is proposed to reduce energy consumption by quantizing accumulated PSUMs in matrix multiplication, alongside a PSQ-APSQ grouping strategy and floating-point regularization.
Owner:THE HONG KONG UNIV OF SCI & TECH +1

Model quantification method, apparatus, computer device, and storage medium

Embodiments of the present application disclose a model quantization method and device, computer equipment and a storage medium, belonging to the technical field of computers. The method comprises: obtaining a first model, the first model comprising a plurality of first operators, a splicing layer and a second operator; determining an input range of each operator in the plurality of first operators and the second operator; updating a quantization parameter of each operator based on the input range of each operator and a target output range, so that the quantization parameter of each operator converges to a quantization parameter corresponding to the target output range; and performing quantization processing on network parameters in the splicing layer based on the quantization parameter corresponding to the target output range, so as to ensure that the plurality of input ranges of the splicing layer and the range of the network parameters are the same, thereby completing quantization of the splicing layer.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Quantization method and device for realizing elastic KV cache by computing power through intelligent computing cloud platform

The application provides a method and device for quantifying elastic KV cache through computing power of an intelligent computing cloud platform, and relates to the technical fields of intelligent computing centers, intelligent computing centers, computing power infrastructure and intelligent computing cloud technology.The method comprises the following steps: S1, dividing historical tokens into multiple cache blocks, quantifying KV data and writing the data into corresponding cache blocks, and selecting multiple candidate anchor points; S2, calculating block-level summary data; S3, in response to a new target token, scoring the cache blocks to generate cache block scores and screening out candidate cache blocks; S4, calculating uncertainty index data and determining a target cache block with to-be-restored precision according to the uncertainty index data; and S5, locating an upstream target anchor point and locally playing back the historical tokens based on the upstream target anchor point to generate target high-precision KV data of the target cache block.The application can greatly improve the quantification effect of KV cache of the intelligent computing cloud platform.
Owner:DATACANVAS LTD

A large language model quantization method based on orthogonal characteristics and accelerator architecture

The application belongs to the technical field of large language model quantization, and particularly relates to a large language model quantization method based on orthogonal characteristics and an accelerator architecture. The quantization method divides the activation tensor of the large language model into multiple column blocks, and allocates an FP4 quantization format to the entire activation tensor with the column block as the granularity. The concept of the column block is defined as follows: the matrix of the activation tensor is divided into multiple segments with the same number of elements, wherein each element in the segment is arranged continuously in the same row in the first dimension of the matrix, and arranged in multiple continuous columns in the second dimension; the column block includes multiple columns in the second dimension, and the number of columns in each column block is consistent with the number of elements in the segment. The application overcomes the defects existing in the existing large language model grouping quantization technology, and solves the contradiction between the precision of the large language model and the hardware efficiency.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Software technical debt identification and quantification method and system based on version control log

PendingCN122633233ARenameSoftware engineering
The present application relates to the technical field of computer software development and version control, and particularly relates to a software technical debt identification and quantification method and system based on version control logs. The method obtains version control logs containing commit records, commit parent node relationships, branch reference records, merge commit records, merge conflict records, code difference records, file renaming records and version tag records, constructs a merge propagation graph anchored by version tags and a code entity identity chain; generates a merge influence unit for a merge commit node within a version tag interval, identifies source branch introduction segments, conflict resolution segments, target branch coverage segments and pre-tag re-introduction segments on the same code entity identity chain; when the above segments are connected in the order of merge propagation edges and commit topologies, a version propagation closed loop technical debt event is generated and a technical debt quantification record is formed.
Owner:SMIC WANYE TECHNOLOGY CO LTD

Network quantization method and apparatus, and related device

This application provides a network quantization method and apparatus, and a related device. A first device receives a first model and a first KL divergence in tth aggregation from a second device. The first device determines a second model in (t + 1)th update and quantization based on the first model and the first KL divergence in the tth aggregation, and sends the second model in the (t + 1)th update and quantization to the second device. It can be learned that the first device may update a quantized local model (namely, the second model) based on global aggregation information / a model (namely, the first model and the first KL divergence). Because global knowledge is integrated, the first device can implement faster convergence when updating the quantized local model, to improve training efficiency. In addition, even if quantization causes a specific accuracy loss, this method still implements good learning performance, can meet an accuracy requirement for a local model, and helps reduce consumption of transmission bandwidth.
Owner:HUAWEI TECH CO LTD

An Ontology-Based Composite Semantic Relevance Quantification Method and System

The present invention discloses an ontology-based method and system for quantifying composite semantic relevance, including: S1, preprocessing the first composite semantics and the second composite semantics and respectively mapping them to the first node set and the second node set of the ontology in the semantic network; S2, setting a first virtual node and a second virtual node for the first composite semantics and the second composite semantics respectively; S3, quantifying the first relevance of the first virtual node to all nodes in the first node set; S4, quantifying the second relevance between the nodes in the first node set and the nodes in the second node set; S5, quantifying the third relevance of the nodes in the second node set to the second virtual node; S6, quantifying the fourth relevance of the first virtual node to the second virtual node, that is, the relevance between the first composite semantics and the second composite semantics. The present invention solves the problem that it is difficult to map real semantics to semantic nodes and the problem of mutual conversion between the relevance of real composite semantics and the relevance of node sets.
Owner:BEIJING ANDY TECH CO LTD

Model quantization method, apparatus, device, and storage medium

A model quantization method, apparatus, device, and storage medium are disclosed. The method involves inputting acquired historical lexical units into a large language model to be quantized, obtaining a set of candidate lexical units and their probability distributions; determining a first contribution value for each candidate lexical unit based on the feature space and depth-sensing weights of each linear layer in the large language model; determining a target lexical unit from the set of candidate lexical units based on the first contribution value; iteratively executing the operation of inputting the acquired historical lexical units into the large language model to be quantized until the iteration stop condition is met, generating calibration samples; using the calibration samples, quantizing the large language model to be quantized to obtain the quantized target model, which is then deployed to a hardware device. The quantized target model can be called by the hardware device to perform corresponding tasks, realizing the generation of calibration data samples from the geometric perspective of the latent manifold of the large language model to be quantized.
Owner:NANJING HOUMO TECH CO LTD

Method and apparatus of quantization configuration for artificial intelligence (AI) / machine learning (ML) models

In an aspect of the disclosure, a method, a computer-readable medium, and an apparatus are provided. The method may be performed by a UE. In certain configurations, the UE collects data samples for an artificial intelligence (AI) / machine learning (ML) model at the UE and a base station. The AI / ML model is trained at the UE or at the base station in a training stage. The UE performs, according to a quantization method, quantization of the data samples to obtain quantized data samples. The UE transmits the quantized data samples to the base station. The UE executes an encoder or a decoder of the trained AI / ML model in an inference stage. The quantization method may be a latent quantization method for a latent space, and the data samples are latent vectors of Channel State Information (CSI) samples measured by the UE.
Owner:MEDIATEK INC +2

Quantification method, system and device of end side large model based on distillation and medium

The invention belongs to the technical field of artificial intelligence, and relates to a distillation-based end side large model quantification method, system and device and a medium, and the quantification method comprises the steps: 1) carrying out simulation quantification on a target large model M composed of N transformer structure layers through weight quantification and activation quantification, 2) based on the target large model M and the simulated and quantified large model # imgabs 1 #, carrying out layer-by-layer distillation on each transformer structure layer of the simulated and quantified large model # imgabs 2 # to obtain a simulated and quantified large model # imgabs 0 # 2; and end-to-end quantization parameter optimization training is carried out on the preliminarily quantized large model # imgabs5 # based on the target large model M and the preliminarily quantized large model # imgabs4 #, so that a finally quantized large model # imgabs6 # can be obtained, and through layer-by-layer distillation and end-to-end quantization parameter optimization training based on self-distillation, the final quantized large model # imgabs6 # can be obtained. The problem that large model quantification needs a large amount of calculation power is avoided, and the method has reliability, expansibility and usability.
Owner:BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD

Quantization method for improving quantization feature

The invention provides a quantification method for improving quantification feature, the convolution speed is improved, and the model reasoning time is shortened. According to the method, deep learning Feature is marked in a mask form, a 4bit value replaces a 5bit value when Feature inference is carried out, and the quantitative expression capability of Feature is improved, so that the quantitative precision of a model is improved. Comprising the following steps: S1, quantifying weights and features; and S2, inference is carried out on the model. According to the method, the convolution speed can be increased, the bandwidth in the network reasoning process is reduced, and the model precision and generalization ability are improved. And the condition of the feature quantification range is kept, and meanwhile, the edge side does not impact the bandwidth, so that the quantification precision of the model is further improved.
Owner:HEFEI JUNZHENG TECH CO LTD

Quantization method and quantization apparatus

The present application provides a quantization method and a quantization apparatus. The method comprises: performing singular value decomposition on a first matrix composed of word embedding vectors, so as to obtain a maximum eigenvalue of each block matrix among M block matrices; on the basis of the maximum eigenvalue of each block matrix, determining the number of quantization bits for each block matrix; performing quantization on each block matrix to obtain quantized block matrices; and, merging the quantized block matrices to obtain a quantized first matrix. The maximum eigenvalue of each block matrix can to some extent characterize the amount of semantic information carried in each block matrix, and the number of quantization bits for each block matrix can be allocated on the basis of the amount of semantic information carried in each block matrix, thereby allowing different bit quantities to be allocated to block matrices carrying different semantic information, achieving differentiated bit quantity allocation and reducing loss of semantic information, and accordingly improving user experience.
Owner:HUAWEI TECH CO LTD

Optimizer quantization method, device and controller based on text generation model

The present application relates to the field of artificial intelligence technology, and in particular to an optimizer quantization method, device, and controller based on a text generation model. The optimizer quantization method includes reading a text input tensor of an optimizer, where the text input tensor is floating-point data of the first bit width; determining the gradient information of the text input tensor and processing the gradient information in blocks to obtain multiple independent blocks; quantizing the independent blocks according to a normalization constant to obtain quantization results of the independent blocks, where the quantization results are integer data of the second bit width; optimizing the quantization results to obtain optimized quantization results, and using the optimized quantization results as the first optimizer state; performing inverse quantization on the first optimizer state to obtain a second optimizer state, and updating the optimizer; quantizing the second optimizer state to return to the first optimizer state, and storing the optimized quantization results of the independent blocks, which is beneficial to reducing the video memory usage of the optimizer in the text generation model and improving the utilization rate of the graphics card.
Owner:PENG CHENG LAB

Quantization method for quantization based on reconstruction quantization operator

Quantification essentially only readjusts a numerical range, and can be roughly understood as linear mapping. The problem of adding'rough 'two words is that some papers are subjected to nonlinear quantization, but the papers are still subjected to linear quantization in the industrial circle at present. However, it can be obviously seen that the inverse quantization generally has no information loss, and the quantization generally has precision loss. It is also very well understood that the numerical value range which can be stored by float32 is larger than that of uint8, so that a large number of numerical values cannot be represented by uint8 and can only be rounded into uint8 type numerical values. The error of the quantitative model and the full-precision model also comes from a clip operation of rounding off. The invention aims to provide a quantization method for quantization based on a reconstruction quantization operator, which can perform asymmetric quantization on a trained model parameter file and perform reasoning acceleration.
Owner:SMIC FUTURE (BEIJING) TECH CO LTD

Quantization method, system and equipment based on neural network hardware and storage medium

The invention relates to a quantification method, system and device based on neural network hardware and a storage medium, and relates to the field of neural networks. The method comprises the following steps: receiving initial data of which the data type is BF16; performing layer normalization on the initial data to obtain active data of which the data type is BF16; designing a matrix multiplier according to neural network hardware; processing and calculating the activation data according to a matrix multiplier to obtain first data; performing inverse quantization on the first data to obtain inverse quantization activity data; and outputting the inverse quantization activity data. The method has the technical effects that the operation speed of the large model is improved while the performance of the large model is kept.
Owner:STORAGEX TECH INC

Quantization method, computing device, and computer-readable storage medium

The present disclosure discloses a quantization method of a neural network, a computing device and a computer readable storage medium. The computing device can be included in a combined processing device, which can further include an interface device and other processing devices. The computing device interacts with the other processing devices to jointly complete a user-specified computing operation. The combined processing device can further include a storage device connected to the computing device and the other processing devices respectively for storing data of the computing device and the other processing devices. The scheme of the present disclosure can greatly reduce the operation time required for computing quantization parameters while maintaining the required network accuracy.
Owner:ANHUI CAMBRICON INFORMATION TECH CO LTD

Quantization method for training and quantizing feature map introduced noise

The invention provides a quantization method for training quantized feature map introduced noise. The method comprises the following steps: S1, inserting a quantization node; s2, counting a quantization range; s3, quantization of the feature map is carried out; and S4, introducing a noise mechanism to obtain a final result: according to the method, the query feature is supplemented, and the noise mechanism is introduced. Compared with an existing general method, the method is higher in precision. And the training quantification model precision can be further improved. Specifically, the model identification is more accurate, and the error rate is lower.
Owner:HEFEI JUNZHENG TECH CO LTD

Quantization methods and related devices for deep learning models

The present disclosure provides a method and related apparatus for quantizing a deep learning model. The method comprises: receiving a deep learning model to be quantized; dividing the deep learning model to be quantized into sub-models; for each sub-model, selecting a quantization algorithm and a quantization strategy corresponding to the sub-model from a combination of pre-set candidate quantization algorithms and candidate quantization strategies, wherein the quantization strategy is a criterion to be followed in addition to the quantization algorithm during the quantization process; and outputting a quantized deep learning model obtained by quantizing the sub-model according to the corresponding quantization algorithm and quantization strategy. The embodiments of the present disclosure overcome the problems of low precision or high complexity in the prior art model quantization, thereby ensuring quantization precision while reducing complexity.
Owner:C SKY MICROSYST CO LTD