Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

26 results about "Variable precision" patented technology

Variable precision logic is concerned with problems of reasoning with incomplete information and resource constraints. It offers mechanisms for handling trade-offs between the precision of inferences and the computational efficiency of deriving them.

Floating point multiply-accumulate unit facilitating variable data precision

A fused dot-product multiply-accumulate (MAC) circuit may support variable precision of floating-point data elements to perform computations in deep learning operations (e.g., MAC operations). The operating mode of the circuit may be selected based on the accuracy of the input element. The mode of operation may be an FP16 mode or an FP8 mode. In the FP8 mode, a product index may be calculated based on an index of a floating point input element. A maximum index may be selected from the one or more product indexes. A global maximum index may be selected from a plurality of maximum indexes. A product mantissa may be calculated based on a difference between the global maximum exponent and a corresponding maximum exponent and aligned with another product mantissa. The adder tree may accumulate the aligned product mantissas and compute the partial and mantissas. The portions and mantissas may be normalized using a global maximum index.
Owner:INTEL CORP

Precision target optimization method and system adaptive to variable precision arithmetic logic unit, medium, terminal and program product

The invention provides a precision target optimization method and system adaptive to a variable precision arithmetic logic unit, a medium, a terminal and a program product. The method comprises the following steps: acquiring an output feature set of each group of an upper layer; the precision generation network layer generates a corresponding precision target according to the output feature set, and the ALU calculation layer generates a prediction result according to the generated precision target; the teacher model generates a reference target and a real label according to the output feature set; constructing a total loss function according to the calculated task loss, precision generation loss and adversarial loss; performing back propagation optimization on the student model based on the constructed total loss function; repeatedly and iteratively training the student model until convergence to obtain a final student model; and deploying the final student model to generate an optimal precision target corresponding to each group. According to the method provided by the invention, the fine precision adjustment of the bit granularity can be realized, the adaptive ability of the model is enhanced, and the matching degree of the precision and the task demand is improved.
Owner:SHANGHAI GUANGYU XINCHEN TECHNOLOGY CO LTD

Dual-group quantization of key-value tensors in transformer-based models

KV tensors in a transformer model may be quantized in a dual-group manner. A key tensor or value tensor may be segmented into groups along both the token dimension and the channel dimension. A group size may be determined based on hardware constraint and model accuracy. The group size may indicate the total number of tokens or the total number of channels in each group. A part of the tensor may be segmented into groups having the group size, while the rest of the attention tensor may constitute an additional group having a larger size. The larger group may include keys or values corresponding to one or more recent tokens. Different groups within the tensor are quantized to variable precision levels with group-specific scale and zero-point values. The quantized tensor may be cached and used in a matrix multiplication operation in an attention module of the transformer model.
Owner:INTEL CORP

METHOD AND DEVICE FOR ROUNDING IN CALCULATIONS WITH VARIABLE PRECISION

ActiveDE602023020690T2Testing MethodsVariable precision
Owner:COMMISSARIAT A LENERGIE ATOMIQUE ET AUX ENERGIES ALTERNATIVES

Floating-point sum-of-accumulate unit for easy variable data precision

PendingJP2026528679AEngineeringFloating point
A fused dot product (MAC) circuit may support variable precision of floating-point data elements to perform calculations (e.g., MAC operations) in deep learning operations. The operating mode of the circuit may be selected based on the precision of the input elements. The operating mode may be FP16 mode or FP8 mode. In FP8 mode, the product exponent may be calculated based on the exponent of the floating-point input elements. The maximum exponent may be selected from one or more product exponents. The global maximum exponent may be selected from multiple maximum exponents. The product mantissa is calculated based on the difference between the global maximum exponent and the corresponding maximum exponent and may be aligned with another product mantissa. The adder tree may accumulate the aligned product mantissas and calculate a partial sum mantissa. The partial sum mantissa may be normalized using the global maximum exponent.
Owner:INTEL CORP

Variable precision and mixed-type representation of multiple layers in a network

In one example, an apparatus comprises a plurality of execution units including at least a first type of execution unit and a second type of execution unit and logic, at least partially including hardware logic, to: expose an embedded projection operation in at least one of a load instruction or a store instruction; determine a target precision level for the projection operation; and load the projection operation at the target precision level. Other embodiments are also disclosed and claimed.
Owner:INTEL CORP

Method and system for identifying sentiment metaphor of low-quality data based on fuzzy granular ball modeling

ActiveCN119884867BSemantic analysisNeural learning methodsMixed noiseVariable precision
The application discloses a low-quality data sentiment metaphor recognition method and system based on fuzzy granular ball modeling, and relates to the field of natural language processing.The method comprises the following steps: S1, obtaining low-quality text data and performing labeling and semantic preprocessing to obtain text data containing linguistic information; S2, inputting the text data containing linguistic information into a word embedding model to obtain sentiment metaphor word vectors and form a word vector matrix; S3, generating a granularity list satisfying a containment threshold by using fuzzy granular ball calculation, performing feature reduction on the word vector matrix by using a variable precision dependency function, and obtaining a reduced matrix; and S4, dividing the reduced matrix into a training set and a test set, training and predicting a convolutional neural network, and obtaining a sentiment metaphor recognition result.The application selects features by using fuzzy granular ball calculation, deletes redundant features, improves the feature extraction efficiency of the convolutional neural network model, and solves the problem of mixed noise information in a large amount of text information acquisition.
Owner:HUAQIAO UNIVERSITY

A method and system for automatically mixed precision optimization of programs

The application discloses a kind of compilation methods for automatically mixed precision optimization procedure, first the preprocessing of the source code file of the program to be optimized is carried out based on the static error analysis technique of chain automatic differentiation, then the precision sensitivity of floating point variable in program is analyzed using the static error analysis technique based on the chain automatic differentiation, precision insensitive variable is determined, and variable information is stored in JSON file;Again, the source code file of the program to be optimized is used as input, the file is traversed by using variable information search tool, the information of all variables in the program is obtained to form a variable configuration file, and a configuration file of variable precision search space is formed according to the precision configuration scheme of the current program;Using the result of error analysis, the variable precision search space is reduced, and the optimized variable precision search space file is formed;The application can solve the technical problem that the execution efficiency of the optimized program cannot be improved by the automatic mixed precision optimization technology based on error analysis.
Owner:HUNAN UNIV

Design method and use method of variable precision calculation unit applied to quantized neural network convolution layer

The application provides a design method and use method of a variable precision calculation unit applied to a quantized neural network convolution layer, shift operation is performed on a vector with a precision of 1, and the shift bit number, dimension and number of an operation basic block are determined. After the operation basic blocks are arranged, an addition tree is connected to form an array, and a variable precision fusion calculation unit is obtained. According to the actual precision of the current convolution layer activation value and the weight value, the variable precision fusion calculation unit array is configured, and each column of the fusion calculation unit is calculated in parallel. After the current convolution layer operation is completed, the two-dimensional array is reconfigured according to the data precision of the next convolution layer, and the operation of all convolution layers is completed in this way. The problems that the operation unit in the prior art is prone to causing waste of calculation resources in convolution operation less than 4 bits, and cannot dynamically adopt different quantization bit widths to process data according to the actual situation of the convolution layer are solved, and the resource utilization rate and the calculation efficiency are improved.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

Variable precision tokenization for adaptive machine-learned models

An example computer-implemented method includes receiving runtime content for rendering using a user interface, wherein the runtime content is selectable at a first precision corresponding to individual content elements that compose the runtime content. The example computer-implemented method includes receiving data describing a runtime input that instructs selection of a portion of the runtime content. The example computer-implemented method includes tokenizing, using a variable precision tokenizer, the runtime content to obtain a tokenized representation of the runtime content. The example computer-implemented method includes processing, using a machine-learned model, the tokenized representation of the runtime content to predict a predicted selection boundary, wherein the predicted selection boundary defines a selection at a second precision corresponding to individual tokens that compose the tokenized representation of the runtime content. The example computer-implemented method includes outputting the predicted selection boundary.
Owner:GOOGLE LLC

Variable-precision Tanh activation function fitting method

The invention discloses a precision-variable Tanh activation function fitting method, which comprises the following steps of: dividing the positive axis input of a Tanh function into a linear region (0, 1.156) and a saturation region (1.156, + infinity) based on the characteristics of the Tanh function, and setting target absolute precision epsilon 0 to determine a demarcation point x0 of a linear function y = x and a trapezoidal function, y = x fitting is adopted in a linear region, it is ensured that the error does not exceed epsilon 0 through segmented iteration in a trapezoidal region, optimization constant tempset fitting is adopted in a saturation region, and negative semi-axis fitting or negative number rejection processing is achieved through the property of a Tanh odd function. By means of the mode, it can be guaranteed that the accuracy of the AI model is reduced within a certain range, hardware resources are greatly reduced, and the requirements for low-power-consumption, small-area and high-speed deployment of the neural network terminal are met.
Owner:NANJING UNIV

Design method and use method of variable precision calculation unit applied to quantization neural network convolutional layer

The invention provides a design method and a use method of a variable precision calculation unit applied to a quantization neural network convolutional layer, and the method comprises the steps: carrying out the shift operation of a vector with the precision of 1, and determining the shift digit, dimension and number of an operation basic block; and arranging operation basic blocks and then connecting the operation basic blocks by using an adder tree to form an array to obtain a fusion calculation unit with variable precision. According to the actual precision of the activation value and the weight value of the current convolutional layer, a variable-precision fusion calculation unit array is configured, and all columns of fusion calculation units carry out parallel calculation; and after the operation of the current convolutional layer is completed, according to the data precision of the next convolutional layer, the two-dimensional array is reconfigured, and the operation of all the convolutional layers is completed by parity of reasoning. The problems that in the prior art, an operation unit is prone to causing waste of calculation resources in convolution operation smaller than 4 bits, different quantization bit widths cannot be flexibly and dynamically adopted to process data according to the actual situation of a convolution layer, and the resource utilization rate and the calculation efficiency cannot be improved are solved.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

A multi-granularity clustering-based scheme for key-value cache compression

The key-value (LV) cache in this application accelerates inference in large language models (LLMs) by allowing attention operations to scale linearly, rather than quadratically, with the total sequence length. Since context lengths are long in modern LLMs, the KV cache size may exceed the model size, potentially negatively impacting throughput. To address this issue, a multi-granularity clustering-based scheme for KV cache compression is implemented. Clusters created at different clustering levels with variable precision approximate the key and value tensors corresponding to less important terms. Precision loss is reduced by using proxies generated at finer-grained clustering levels of subsets of more salient attention heads. More salient attention heads have a greater impact on model accuracy than less salient attention heads. When the impact on accuracy is low, latency is improved by retrieving proxies from a subset of less salient attention heads from faster memory.
Owner:INTEL CORP

Method and device for variable precision computing

The present disclosure relates to a floating-point computation circuit comprising: an internal memory (104, 114) storing one or more floating-point values in a first format; status registers (124) defining a plurality of floating-point number format types associated with corresponding identifiers, each format type indicating at least a maximum size (BIS, MBB); and a load and store unit (108, 118) for loading floating-point values from and storing floating-point values to an external memory (120, 122), the load and store unit (108, 118) being configured: to receive, in relation with a first store operation, a first floating-point value from the internal memory (104, 114) and a first of said identifiers; and to convert the first floating-point value from the first format to a first external memory format having a maximum size (BIS, MBB) defined by the floating-point number format type designated by the first identifier.
Owner:COMMISSARIAT A LENERGIE ATOMIQUE ET AUX ENERGIES ALTERNATIVES

Methods for coding, transmitting, decoding and processing elements in a vector space or semantic data

Disclosed are methods for coding, transmitting, decoding and processing elements in a vector space or semantic data with variable precision. The invention relates to a method for coding an element (SD) belonging to a vector space comprising a plurality of subsets respectively identified by sequences of bits, in which the element (SD) is coded in the form of a sequence of bits (S) identifying a subset containing the element (SD), and to a transmission method carried out from a transmitting entity (100), the transmission method comprising: - transmitting, to a receiving entity (200), a sequence of bits (S) obtained by coding an element (SD) in a vector space in accordance with the coding method according to the invention; and to associated decoding and processing methods.
Owner:ORANGE SA

Tensor arithmetic unit based on RISC-V instruction set and intelligent processor

The invention provides a tensor arithmetic unit based on an RISC-V instruction set, and the tensor arithmetic unit comprises a microinstruction splitting and scheduling unit which receives a decoded macroscopic tensor instruction and an operation size parameter thereof, splits the macroscopic tensor instruction into microinstruction sequences according to a fixed physical scale of a reconfigurable calculation array, and transmits the microinstruction sequences to the reconfigurable calculation array; a zigzag traversal sequence and microinstruction scheduling are achieved through quintuple hardware circulation, and the loading, using and replacing sequence of the data blocks in the block register array is planned; the reconfigurable computing array is used for executing matrix multiply-accumulate operation of various data formats by designing a reconfigurable data path and fusing and multiplexing a floating point multiplier and a multi-precision accumulation tree under data paths with different bit widths; and the hardware multi-buffer unit is a multi-buffer architecture automatically managed by hardware and is matched with a zigzag traversal sequence to realize pipeline overlapping of data prefetching and calculation execution. The invention further provides an intelligent processor. Therefore, the invention can efficiently execute tensor operation with variable precision and variable scale.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

A reconfigurable activation function hardware device adapted for deep learning hardware accelerators

The application provides a reconfigurable activation function hardware device suitable for deep learning hardware accelerator, comprising a function type judging unit, a ReLU calculation unit, a simplified function calculation unit, a variable precision unit and an optimized function calculation unit. The application makes full use of the correlation between different nonlinear activation function calculation expressions, can realize approximate calculation of nine commonly used activation functions of neural networks, i.e. ReLU function, ReLU6 function, PReLU function, Leaky ReLU function, Sigmoid function, Tanh function, Swish function, H-Sigmoid function and H-Swish function, thereby adapting to multifunctional deep learning hardware accelerator, achieving a good balance between calculation resources and approximate accuracy, and having the characteristics of high calculation efficiency, flexibility, reconfigurability and the like.
Owner:NANJING UNIV

Seat Hall offset prevention control method and system based on variable precision and application

The invention belongs to the technical field of automobile seat control, and provides a seat anti-Hall offset control method and system based on variable precision and application, and the method comprises the steps: obtaining the track length of a seat motor operation track; comparing the total Hall number with a reference value, dynamically selecting a percentage amplification coefficient calculation strategy, and setting the minimum step precision of the percentage position in the software, so that the minimum step precision is equal to or corresponds to one Hall number; multiplying the target position percentage by a percentage amplification coefficient to obtain a target percentage value in the software, and calculating to obtain a target Hall position; and after the seat motor runs to the target Hall position and stops, dividing the current percentage value in the software by the percentage amplification coefficient to obtain a real percentage position of the seat motor and feeding back the real percentage position so as to perform closed-loop control on the position of the seat motor. According to the method, no hardware change is involved, and the problem of Hall offset caused by repeated movement of the seat is solved through a software adjusting method with variable precision.
Owner:ANHUI JIANGHUAI AUTOMOBILE GRP CORP LTD

Tumor gene classification method based on variable precision fuzzy rough set

The invention discloses a tumor gene classification method based on a variable precision fuzzy rough set, and relates to the technical field of mining and bioinformatics crossing, and the method comprises the following steps: constructing a variable precision fuzzy rough set model based on a pseudo-overlap function, and defining a fuzzy positive domain and fuzzy dependency degree used for measuring the correlation degree between attributes and decisions; providing a feature selection algorithm based on the model, the rough fuzzy set and the fuzzy dependency degree; a feature selection algorithm is combined with an intelligent classifier to be applied to tumor gene classification; according to the method, the variable-precision fuzzy rough set model based on the pseudo-overlap function is provided, so that the processing capability of fuzzy information is enhanced, the uncertainty of a fuzzy relationship in a discourse domain can be flexibly dealt with, a solid theoretical basis is provided for accurately depicting the association between attributes and decisions, and the adaptability and expression capability of complex fuzzy data are improved.
Owner:SHAANXI UNIV OF SCI & TECH

Neural inference processing unit with flexible precision

ActiveCN114787823BDigital computer detailsInference methodsAlgorithmMatrix multiplier
Neural inference chips are provided. A neural core of a neural inference chip includes a vector-matrix multiplier; a vector processor; and an activation unit, which is operatively coupled to the vector processor. The vector-matrix multiplier, vector processor, and / or activation unit are adapted to operate at variable precision.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Variable-precision SNN in-memory computing macro circuit and forward reasoning method

The invention relates to the field of artificial intelligence and brain-like chips, in particular to a variable-precision in-memory computing macro circuit and a forward reasoning method, and the method comprises the steps: mapping a pulse neural network (full connection or convolutional neural network) formed based on LIF neurons into the variable-precision in-memory computing macro circuit, and selecting signals according to different precisions, and forward reasoning processes with different precisions are realized in the storage and calculation array, so that the energy consumption and delay of network calculation are greatly reduced, and the energy efficiency of the processor is greatly improved.
Owner:CHONGQING UNIV

FPGA-based parallelization method for floating-point multiplication

ActiveCN116028012BSign bitParallel computing
The application discloses a kind of parallelization multiplication operation methods of floating point number based on FPGA, comprising the following steps: based on the representation under IEEE754 different precision, design a kind of low-precision storage mode that can be blocked, including the design of exponent block and floating block;Then design arbitrary variable bit floating block fixed-point addition, variable bit floating block fixed-point multiplication is realized by the way of multiplication pool;Then using FPGA, the multiplication calculation result between the sign bit and the exponent bit of the multiplicand and the multiplier is obtained by exclusive or operation and fixed-point addition, the multiplication calculation result of the significand bit of the multiplicand and the multiplier is obtained by fixed-point multiplication;Finally, the obtained multiplication calculation result is normalized.The application can be effectively applied in memory computing, with the increase of the number of blocks, the data calculation delay can be significantly reduced, and the variable precision can improve the flexibility of calculation.
Owner:SOUTHEAST UNIV