Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

922 results about "Floating point" patented technology

In computing, floating-point arithmetic (FP) is arithmetic using formulaic representation of real numbers as an approximation to support a trade-off between range and precision. For this reason, floating-point computation is often found in systems which include very small and very large real numbers, which require fast processing times. A number is, in general, represented approximately to a fixed number of significant digits (the significand) and scaled using an exponent in some fixed base; the base for the scaling is normally two, ten, or sixteen. A number that can be represented exactly is of the following form...

Processor and electronic equipment

The invention discloses a processor and electronic equipment. The processor includes a computing unit including a tensor core configured to perform a matrix multiplication operation using a scaling factor, and a memory, the computing unit further including a scaling factor processing module configured to determine and cache a scaling factor for each tensor associated with the matrix multiplication operation, the computing unit further comprises at least one storage module arranged on a data path between the tensor core and the memory, and the at least one storage module is exclusively occupied by the tensor core when the tensor core executes tensor related operation. And the scaling factor processing module is arranged on the at least one storage module. At present, scaling factors and floating-point number quantization are completed by a vector calculation core, so that performance is reduced, and delay becomes high, and a scaling factor processing module arranged on a storage module in a calculation unit can improve the overall execution efficiency of low-precision matrix multiplication using the scaling factors.
Owner:SHANGHAI BIREN TECH CO LTD

Frequency difference compressor based on precision perception, gradient compression method, equipment and medium

The invention provides a frequency difference compressor based on precision perception, a gradient compression method, equipment and a medium, which are used for carrying out data compression and transmission between a server and a client so as to reduce communication overhead in personalized federated learning, and relates to the technical field of data compression. The frequency difference compressor comprises an information bottleneck rarefaction unit, a frequency domain compression unit, a dynamic quantization unit and a differential coding unit. Non-key gradient redundant components of the original gradient data are removed through an information bottleneck rarefaction unit to generate sparse gradient data; generating a metadata packet containing the first N high-energy frequency domain coefficients and the positions and the number of the first N high-energy frequency domain coefficients through a frequency domain compression unit; performing adaptive bit width mapping on the high-precision floating point gradient data into low-order integer representation through a dynamic quantization unit to generate dynamic quantization gradient data; and the differential gradient is transmitted to the client through the differential coding unit, so that data compression and transmission are completed. The method solves the problems that the communication overhead is large and gradient information cannot be reserved as far as possible in gradient transmission.
Owner:XIAMEN UNIV OF TECH

Mixed data precision matrix multiplication and addition unit and calculation method

The invention provides a mixed data precision matrix multiplication and addition unit and a calculation method, the matrix multiplication and addition unit comprises a calculation unit, and the calculation unit comprises a format division module, a multiplication array module, an addition tree module, an accumulator module, a normalization module and a shift register module. The calculation unit converts the first input matrix and the second input matrix into input data in a middle floating point format; executing parallel multiplication operation on the input data to generate an intermediate product result; performing index alignment and accumulation on the intermediate product result to generate an intermediate accumulated value; accumulating the intermediate product result and the value of the third input matrix in a form of accumulating an intermediate accumulated value, and outputting an accumulated result; and converting an accumulation result into a normalized result and outputting the normalized result. The format division module supports various precisions and converts data with different widths into an intermediate floating point format, so that other hardware units can be reused, and the problems that hardware resources are complex and different model reasoning scenes are difficult to meet are solved.
Owner:NANJING UNIV

Tensor core, processor, data processing method, electronic device and storage medium

The invention discloses a tensor core, a processor, a data processing method, electronic equipment and a storage medium, and is applied to the field of tensor processing. The tensor kernel comprises a first dot multiplication unit and a scaling factor matrix multiplication processing module, the scaling factor matrix multiplication processing module comprises a second dot multiplication unit, and the first dot multiplication unit and the second dot multiplication unit support dot multiplication operations of different floating-point number precisions. The tensor core is configured to receive a first tensor, a second tensor, a scaling factor of the first tensor, a scaling factor of the second tensor, an offset term of the first tensor, and an offset term of the second tensor, perform a matrix multiplication operation using the scaling factor and the offset term using a first dot multiplication unit and a scaling factor matrix multiplication processing module, and obtaining a matrix multiplication operation result of the first tensor and the second tensor. Matrix multiplication operation using scaling factors is executed by multiplexing dot multiplication units with different precisions in a tensor kernel, extra hardware area cost is reduced, and existing hardware resources are fully multiplexed.
Owner:SHANGHAI BIREN TECH CO LTD

Picture occlusion relation processing method and device and storage medium

The invention discloses a picture occlusion relation processing method and device and a storage medium, and relates to the technical field of rendering, and the method comprises the steps: determining a vertical distance between a target element in a to-be-rendered frame and a preset base point; performing normalization processing on the vertical distance to generate a floating point value corresponding to each target element; encoding the floating point value to generate at least one integer channel value corresponding to the element; storing the integer channel value corresponding to the target element to vertex color information of the grid data corresponding to the target element to generate a rendering grid; and rendering the target element based on the rendering grid. According to the method, the normalized and coded depth value is borne through the vertex color information, so that the technical problem that the rendering performance is low due to the fact that the node tree needs to be traversed in each frame and massive elements need to be sorted in the prior art is solved, seamless fusion of the depth data of the rendered elements and the vertex colors is realized, and the rendering efficiency of the elements is improved.
Owner:SHENZHEN ZIXIAO INTERACTIVE TECH CO LTD

Multi-dimensional collaborative AI training processor benchmark evaluation device and method

PendingCN120315982AHardware monitoringProcessor modelFloating point
The invention discloses a multi-dimensional collaborative AI training processor benchmark evaluation device and method. According to end-to-end performance evaluation, TTA and throughput indexes are obtained through a standardized training process; the operator support degree evaluation adopts a dynamic parameter generation technology to evaluate the speed-up ratio and the floating point performance utilization rate of calculation-intensive and memory-access-intensive operators; the energy consumption evaluation calculates a power consumption delay product based on steady-state power consumption monitoring. And finally, integrating multi-dimensional indexes through an analytic hierarchy process model and a fuzzy theory, realizing level mapping by adopting a four-section membership function, and finishing multi-level aggregation evaluation in combination with matrix operation. According to the method, three dimensions of performance, operator support degree and energy efficiency are creatively integrated, the problems of single evaluation index and high environment dependence of the heterogeneous AI processor are solved through automatic container deployment, dynamic parameter generation and steady-state power consumption monitoring technologies, and a comprehensive and quantifiable evaluation basis is provided for processor model selection.
Owner:XIDIAN UNIV

Quantization Error Compensation for Vector Computing

A method for performing a computing task includes: extracting one or more features from a user content; converting the features to a floating point query vector; quantizing the floating point query vector; obtaining a database vector including one or more floating point feature vectors; determining a compensation vector based on a data distribution of the floating point query vector; quantizing the floating point feature vectors; determining an error function based on a difference between data distributions of i) the quantized query vector compensated with the compensation vector, and ii) the floating point query vector; determining, based on the error function, values of the compensation vector corresponding to the quantized feature vectors; combining the quantized query vectors and the values of the compensation vector to obtain one or more compensated query vectors; and performing the computing task using the compensated query vectors and the quantized feature vectors to obtain an output.
Owner:MACRONIX INTERNATIONAL CO LTD

Model Compression Method and Apparatus, and Related Device

A model compression method includes obtaining a first weight of each layer of a neural network model, where the first weight of each layer is a value of a floating-point type; and quantizing the first weight of each layer based on a quantization parameter to obtain a second weight of each layer. The second weight of each layer is a multi-bit integer. Quantities of bits of second weights of at least a part of layers are different. A quantity of quantization bits of a second weight of a layer with high sensitivity to a quantization error is greater than a quantity of quantization bits of a second weight of a layer with low sensitivity to the quantization error.
Owner:HUAWEI TECH CO LTD

V-DMC signalling improvements in displacement sub-bitstream for wavelet coefficient inter prediction with fixed-point inverse quantization

A device for decoding encoded dynamic mesh data determines a set of quantized integer coefficient values for displacement vectors of the encoded dynamic mesh data; determines a quantization parameter based on one or more syntax elements included in a displacement sub-bitstream of the encoded dynamic mesh data; inverse quantizes, based on the quantization parameter, the set of quantized integer coefficient values to determine a set of fixed-point dequantized coefficient values; determines a set of fixed-point transformed coefficient values based on the set of fixed-point dequantized coefficient values; converts the set of fixed-point transformed coefficient values to a set of floating-point transformed coefficient values; inverse transforms the set of floating-point transformed coefficient values to determine a set of reconstructed displacement vectors; and determines a reconstructed deformed mesh based on the set of reconstructed displacement vectors.
Owner:QUALCOMM INC

Data processing method based on large language model, and large language model and electronic device

Disclosed in the present application are a data processing method based on a large language model, and a large language model, an electronic device, a computer-readable storage medium and a computer program product. The method is applied to a user terminal, wherein a large language model is deployed on the user terminal, and weight parameters of linear calculation layers of the large language model are pre-quantized into format data of an integer data type. The method comprises: acquiring input data; performing vector conversion on the input data by means of an embedding layer of a large language model, so as to obtain a floating-point query vector of a floating-point data type corresponding to the input data; converting the floating-point query vector into an integer query vector of an integer data type; and performing an operation by means of weight parameters of linear calculation layers and the integer query vector, so as to obtain a query result corresponding to the input data. By means of the solution provided in the present application, a large language model can be smoothly run on a user terminal, such that the user terminal can provide services for users without needing network connectivity, and can better ensure the privacy of the users.
Owner:TAOBAO CHINA SOFTWARE

BLAS3 structured operator accelerated computing system based on Hopper architecture GPU

The invention provides a BLAS3 structured operator accelerated computing system based on a Hopper architecture GPU, and relates to the technical field of computers. The system comprises: a calculation unit discrimination module for determining a calculation unit used by a current operator during operation, and estimating the maximum row dimension upper bound of the current operator in a tensor core execution path; an instruction sensing block parameter determination module dynamically determines the optimal block size and number of the input matrix in real time; the block matrix loading and aligning module divides an input matrix and a matrix to be updated into sub-matrixes by taking the block size as a basic block and completes loading of the corresponding sub-matrixes; the operator kernel function execution module completes shared memory structured parallel loading and storage of a double-precision floating-point number array of a sub-matrix corresponding to the input matrix, and calls a tensor core to carry out multiply-add accumulation calculation; and the assembly line and concurrent scheduling module adds the block calculation tasks into corresponding task sets and performs multi-stream concurrent scheduling on the task sets.
Owner:NORTHEASTERN UNIV CHINA

Methods and systems for quantization of large language models

A method and an apparatus for storing data points are provided. The method comprises: receiving the plurality of data points, each data point of the plurality of data points being represented in a floating-point representation; quantizing each one of the first plurality of data points, by: executing, during a first quantization phase: converting each data point of the plurality of data points into a corresponding first data point of a plurality of first data points; executing, during a second quantization phase: applying, to each first data point of the plurality of first data points a clamping function, thereby converting each first data point of the plurality of first data points into a corresponding second data point of a plurality of second data points; and storing the plurality of second points for further calculations instead of the first plurality of data points.
Owner:HUAWEI TECH CO LTD

Low-bit-width high-energy-efficiency floating point storage and calculation integrated circuit based on partial pre-alignment architecture

The invention belongs to the technical field of storage and calculation integration, and particularly relates to a low-bit-width and high-energy-efficiency floating point storage and calculation integrated circuit based on a partial pre-alignment framework. The circuit comprises a memory array, a pre-calculation unit, an adder tree, a configurable arithmetic unit and a normalization unit, and supports mixed precision operation of FP8MACFP4 and FP8MACFP8. The method is characterized in that a partial pre-alignment strategy dominated by an activation value is adopted, the maximum index of the activation value is dynamically counted, the mantissa of the maximum index is aligned, and multiple partial pre-alignment intermediate results are pre-calculated and latched for reuse; in combination with a customized lookup table and a multiplexer, a pre-calculation result is directly selected to replace real-time multiplication and displacement; and through the reconfigurable hardware, the FP8MACFP8 high-precision operation is realized by utilizing the FP8MACFP4 unit combination. According to the method, complete online floating point multiplication and addition operation is realized, and excellent energy efficiency ratio and operation speed are obtained while high precision is kept.
Owner:FUDAN UNIVERSITY

Accelerator for operations between floating point matrix and integer matrix and operation method thereof

An operation accelerator that performs an operation between a floating point matrix and an integer matrix includes a first buffer storing integer matrix data; a second buffer storing floating point matrix data; a data converter to convert the floating point matrix data into an integer; and an operator to perform multiplication on the integer matrix data and integer operation target matrix data output from the data converter, wherein the data converter includes a pre-aligner to find a maximum exponent value among multiple floating point values included in the floating point matrix data, perform pre-alignment for moving a mantissa of each of floating points by a difference between the maximum exponent value and an exponent value of each of the multiple floating point values, and generate the integer operation target matrix data based on mantissas of a preset number of high-order bits extracted from among mantissas of pre-aligned floating point values.
Owner:SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION

Floating point multiply-accumulate unit facilitating variable data precision

A fused dot-product multiply-accumulate (MAC) circuit may support variable precision of floating-point data elements to perform computations in deep learning operations (e.g., MAC operations). The operating mode of the circuit may be selected based on the accuracy of the input element. The mode of operation may be an FP16 mode or an FP8 mode. In the FP8 mode, a product index may be calculated based on an index of a floating point input element. A maximum index may be selected from the one or more product indexes. A global maximum index may be selected from a plurality of maximum indexes. A product mantissa may be calculated based on a difference between the global maximum exponent and a corresponding maximum exponent and aligned with another product mantissa. The adder tree may accumulate the aligned product mantissas and compute the partial and mantissas. The portions and mantissas may be normalized using a global maximum index.
Owner:INTEL CORP

Self-adaptive repairing method for high-density NAND storage medium

The invention discloses a high-density NAND storage medium self-adaptive repairing method which comprises the following steps: a main control chip separates a target feature vector representing the real aging trend of a storage unit from original read data containing random physical noise; the main control chip deduces the target feature vector by using a full-integer recursive prediction model to obtain a health state prediction result of the storage unit in a future preset time period; wherein the health state prediction result comprises an optimal read reference voltage offset and an estimated bit error rate growth curve; and the main control chip adjusts the charge distribution pattern and programming voltage parameters of the data on the storage unit in the data writing stage according to the health state prediction result. By means of the mode, the problems of firmware assembly line blocking and performance jitter caused by floating point operation and huge model parameter loading can be solved under the limited hardware environment that the main control chip only supports integer operation and on-chip cache is extremely small.
Owner:深圳华芯星半导体有限公司

Finger-storage-multiply-accumulate three-stacked floating-point number in-memory computing system

The invention discloses a finger-storage-multiply-accumulate three-stacked floating-point number in-memory computing system, and belongs to the field of application-specific integrated circuit design. According to the system, a three-stack type floating-point number calculation framework is designed, a pointing shift and mantissa multiply-accumulate module is designed, and floating-point number multiply-accumulate calculation of global pointing is supported. Compared with a traditional parallel floating-point multiply-accumulate circuit, the parallel floating-point multiply-accumulate circuit has the advantages that a calculation process of global alignment shifting and mantissa multiply-accumulate in sequence and a three-stack type calculation framework are adopted, similar calculation precision is achieved with smaller hardware overhead, low hardware overhead and high calculation precision are both considered, and the comprehensive performance of the calculation framework and the circuit is improved.
Owner:SOUTHEAST UNIV

Floating point calculation device and method for processor, electronic equipment and storage medium

The embodiment of the invention discloses a floating point calculation device and method for a processor, electronic equipment and a storage medium, and the device comprises a mantissa alignment module which is used for carrying out the interception processing of a first mantissa of a first floating point number based on an interception digit to obtain a first addend, and carrying out the first addend based on a shift digit to obtain a second addend; the second mantissa of the second floating-point number is shifted to obtain a second addend, and the first addend and the second addend are two mantissas with aligned indexes; the interception bit number and the shift bit number are determined based on an index difference value between a first index of the first floating-point number and a second index of the second floating-point number; the adder is used for adding the first addend and the second addend to obtain a first mantissa sum; and the post-processing module is used for performing normalization processing and rounding processing on the first mantissa sum to obtain a target mantissa. Therefore, the selector processing process before the mantissa is shifted in the floating point addition calculation process can be reduced, the calculation path is shortened, and the calculation speed is improved.
Owner:MOORE THREADS TECH CO LTD

Arithmetic logic unit of processor and floating-point number calculation method

The invention provides an arithmetic logic unit of a processor and a floating-point number calculation method, and belongs to the technical field of floating-point numbers. The arithmetic logic unit comprises a format conversion module and a mixed precision multiplier-adder, the format conversion module performs format conversion on a high-precision floating-point number to obtain a low-precision floating-point number, and retains an index lost in the conversion process as index auxiliary information, and when the mixed precision multiplier-adder performs multiply-add operation on the low-precision floating-point number, the mixed precision multiplier-adder performs multiply-add operation on the low-precision floating-point number. And the index auxiliary information is used for compensating the lost index, so that the precision loss of the floating-point number can be reduced when the mixed precision multiplier-adder is used.
Owner:HUAWEI TECH CO LTD

Method for applying linear programming to CDN (Content Delivery Network) scheduling

The invention discloses a method for applying linear programming to CDN (Content Delivery Network) scheduling, which relates to the technical field of content delivery networks and comprises the steps of data preparation, strategy layer version smooth configuration, macroscopic layer and microscopic layer linear solution and online execution. Basic data are collected, cleaned and repaired, and a version change rule is set; the macroscopic layer constructs a linear programming model, and the cross-provincial bearing quota is solved with the aim of minimizing the cross-provincial cost; the micro layer takes the quota as a boundary and generates domain name class-node weight vectors in parallel; and adapting a routing request online through weighted rendezvous hashing and request features. According to the method, a dynamic cost matrix and a weight granularity control technology are integrated, the engineering problem of linear programming is solved, second-level response, approximate global optimal scheduling and accurate execution of floating-point-level weight are realized, memory overhead is reduced, smooth updating of a strategy and system stability are guaranteed, and CDN service quality and operation efficiency are improved.
Owner:YUNZHOU TIMES TECHNOLOGY CO LTD

Weight quantization method and device based on hybrid segment coding and server

The invention provides a weight quantification method and device based on mixed segment coding and a server, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a to-be-processed floating point weight and an activation value; based on the Laplacian distribution of the floating point weights, carrying out hybrid segmented coding processing on the floating point weights to obtain a coding weight set in a quantization weight coding format; and determining a multiplication coefficient and a shift value of the coding weight based on the identification bit, and carrying out multiplication and addition calculation processing on the activation value and the weight information in the data bit based on the multiplication coefficient and the shift value to obtain an output activation value corresponding to the coding weight. According to the method, the quantization precision of the neural network can be remarkably improved, and the quantization error is reduced while the compression ratio is not changed.
Owner:ZHEJIANG XINMAI SILICON CO LTD

Inverse quantization method of model data, matrix operation method and related equipment

The embodiment of the invention discloses an inverse quantization method of model data, a matrix operation method and related equipment, and the method comprises the steps: determining first floating point data to be inversely quantized in a model; regarding the first floating point data as first integer data with the same bit width as the first floating point data, and performing first type conversion processing of expanding the first bit width to a second bit width to generate intermediate floating point data; performing second type conversion processing for maintaining a second bit width on the intermediate floating point data to generate second integer data; performing format correction processing on the second integer data to generate inversely quantized second floating point data; according to the method, the inverse quantization process can be simplified, and the calculation efficiency can be improved, so that chips which are partially limited by instruction sets and lack special calculation instruction support for low-precision floating point data are adapted, and the overall performance of artificial intelligence model training is improved.
Owner:ZHONG KE JIA HE (BEI JING) KE JI YOU XIAN GONG SI

Model quantification implementation method, model and computer equipment

The invention provides a model quantization implementation method, a model and computer equipment, and the method comprises the steps: obtaining a target floating point model and a training data set, and determining the quantization configuration corresponding to each neural network layer contained in the target floating point model; the quantization configuration corresponding to different neural network layers is related to the change degree of the performance of the corresponding neural network layer compared with the target floating point model after the corresponding neural network layer is quantized, and according to the quantization configuration corresponding to each neural network layer, performing pseudo quantization on parameters of the corresponding neural network layer in the target floating point model to obtain a first pseudo quantization model; and based on the target floating point model and the training data set, performing quantitative perception training on the first pseudo-quantitative model to obtain a target quantitative model, and deploying the target quantitative model to a terminal device to execute a calculation task.
Owner:SMARTER SILICON (SHANGHAI) TECH CO LTD

Complex operation lightweight architecture and method based on FPGA (Field Programmable Gate Array)

The invention provides a complex operation lightweight architecture and method based on an FPGA, the architecture comprises a main scheduling module, a calculation module and a floating point operator module, the calculation module comprises a storage unit and a calculation unit, the storage unit obtains and caches input data of a target algorithm, and the main scheduling module carries out the calculation of the target algorithm according to the operation sequence of the target algorithm. The target algorithm is divided into a plurality of calculation steps, and corresponding operation enable signals are generated and transmitted to the calculation unit; the calculation unit obtains target parameters required by calculation steps from the storage unit according to the operation enable signal, obtains floating point operator units required by the calculation steps from the floating point operator module, and executes the calculation steps to determine an operation intermediate result; and storing the operation intermediate result in a storage unit, generating an operation completion signal and transmitting the operation completion signal to a main scheduling module to control the next calculation step until a final operation result is obtained. Therefore, the FPGA architecture with low power consumption and miniaturization is realized.
Owner:EHIWAY MICROELECTRONIC SCI & TECH (SUZHOU) CO LTD

Deep learning reasoning service performance analysis method based on kernel function trajectory

The invention provides a kernel function trajectory-based deep learning inference service performance analysis method, which comprises the following steps of: based on service indexes and hardware theoretical computing power acquired from a production cluster, defining floating point operation times per request (FPR) index to quantify service resource efficiency, and identifying high FPR hotspot services; positioning a reasoning iteration candidate boundary based on a GPU kernel function trajectory, verifying iteration integrity through fingerprint matching and chi-square test, and calculating a second reasoning iteration number IIPS and a model reasoning efficiency MIE; aiming at calculation-intensive operators on the key path, combining a dynamic Roofline model to estimate an operator theoretical performance upper limit, and based on actual execution time, calculating efficiency and a BottleScore index to identify a key bottleneck operator; and outputting targeted optimization suggestions according to analysis results of service efficiency analysis, model efficiency analysis and operator efficiency analysis. According to the method, the inference behavior pattern can be automatically identified from massive kernel trajectories, and the efficiency loss of each level is quantified.
Owner:UNIV OF SHANGHAI FOR SCI & TECH +1

Floating point accumulation

There is provided an apparatus, a system, a chip containing product, a method and a medium, the apparatus comprises decoder circuitry responsive to a floating point accumulate instruction identifying pairs of floating point operands and an accumulation source, and processing circuitry comprising a plurality of arithmetic combination units to perform an arithmetic operation to combine the pairs of operands, and summation circuitry to perform an arithmetically precise summation operation to calculate an intermediate result by summing results generated by the arithmetic combination units. The intermediate result is independent of an order in which the arithmetic results are summed. The processing circuitry comprises accumulation circuitry to accumulate the intermediate result into the accumulation source and rounding circuitry to perform a rounding operation after accumulating, and is configured to propagate additional precision information relating to the arithmetic results, the intermediate result, and / or the prior accumulation value to the rounding circuitry.
Owner:ARM LTD

Resource reuse type transcendental function calculation device and calculation method

The invention relates to transcendental function calculation, in particular to a resource reuse type transcendental function calculation device and method, and the device comprises a selector which receives input data, a data type and a control signal, and transmits the input data and the data type to a corresponding processing unit PE according to the control signal; the processing unit PE can process all types of transcendental functions and perform transcendental function calculation on input data according to data types; the result output unit is used for splicing the calculation results of the plurality of processing units PE in sequence to obtain output data; when all the processing units PE perform transcendental function calculation, the lookup table, the fixed-point multiplier, the floating-point multiplier and the floating-point adder are multiplexed; according to the technical scheme provided by the invention, the defect that the transcendental function is difficult to accurately calculate by using fewer hardware resources in the prior art can be effectively overcome.
Owner:安徽芯纪元科技有限公司

KV cache data quantification device and method

The invention provides KV cache data quantization equipment, which comprises a precision decision interface, a compressor and a storage interface, and is characterized in that the precision decision interface is used for determining a quantization precision mark of each token in combination with token quantization difficulty, and generating a mixed quantization precision instruction according to the quantization precision marks corresponding to all the tokens; the compressor is used for acquiring a floating point Key value and / or a floating point Value value corresponding to each token, quantizing the floating point Key value and / or the floating point Value value based on a mixed quantization precision instruction issued by the precision decision interface, and generating metadata with a scaling factor and a zero point of each token and corresponding quantization data; and the storage interface is used for combining the quantized data corresponding to the quantized precision marks with the low-precision bit width in pairs, associating the quantized data with the metadata through an address mapping table, and separately storing the quantized data and the metadata. According to the invention, the storage space occupied by the KV cache can be reduced, and the reasoning speed of the model is improved.
Owner:HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD

Storage and calculation integrated processor and processing method based on POSIT format

According to the storage and calculation integrated processor based on the POSIT format and the processing method, the calculation complexity during network training can be greatly reduced by utilizing the improved POSIT data format, meanwhile, due to the fact that the characteristic of dynamic bit width is reserved, the numerical value representation range is widened, the neural network training precision is guaranteed, and the processing efficiency is improved. The problems that in the prior art, low-precision training supports multiple floating points or integer data formats, and POSIT data types are not supported are solved. And a large number of optimization means are needed to ensure the precision, so that the technical scheme is complicated, the data format lacks flexibility, and network parameters cannot be effectively represented.
Owner:XIDIAN UNIV HANGZHOU RES INST +1

System for post-training quantization of large language models

Post-training quantization of weight values and activation values substantially reduces the memory and processing requirements of floating-point (FP) large language models (LLMs). A quantization parameter training process is performed on the FP LLM to determine quantization parameters. Weight-activation scaling may be applied to linear modules of the LLM, including down projection layers, enabling subsequent per-tensor quantization for activation values. The weight and activations values of the FP LLM are quantized from FP to integer values. Different layers may have different integer sizes. For example, weight values may be reduced to 4 bit integers and activation values to 8 bit integers. Layers within the model are modified to operate on the integer values. For example, an integer SiLU module may provide an integer approximation of a sigmoid-weighted linear unit activation function.
Owner:AMAZON TECH INC