Low-Bit-Width Data Quantization With Vector Precision Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing quantization techniques for machine learning models fail to achieve satisfactory compression rates and precision, especially at very low bit widths, leading to inefficient resource utilization.

Innovation Solution

A method involving the extraction of first vectors from a matrix, creation of objective functions, and determination of mapping parameters to transform these vectors into second vectors with a lower data width while maintaining a predetermined precision condition, using a zero-point and scaling parameter for quantization and inverse quantization processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If data is compressed from higher bit width to lower bit width using existing quantization techniques, then resource consumption is reduced, but data precision becomes unsatisfactory

Engineering Contradiction:
Improveresource consumptionVSAvoiddata precision
Core Design Contradiction:
Loss of energyVSMeasurement precision

Solution Approach 1:

The patent segments the quantization process into multiple stages: extracting first vectors from the matrix, creating objective functions for each vector, determining second vectors with lower bit width, and establishing mapping parameters. This segmented approach allows precise control at each stage to maintain overall data precision while achieving compression.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes key parameters including bit width (from first bit width to second bit width), vector representations (from first vectors to second vectors), and introduces mapping parameters with zero-point and scaling factors. These parameter changes enable flexible control over the trade-off between compression ratio and precision.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If data is compressed to very low bit widths, then compression rate improves, but quantization precision deteriorates

Engineering Contradiction:
Improvecompression rateVSAvoidquantization precision
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces second vectors as an intermediary representation between the original high-precision first vectors and the final compressed data. These second vectors serve as a bridge that enables low-bit-width compression while maintaining precision through the mapping parameter that connects back to the original vector space.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent performs preliminary actions by extracting first vectors and creating objective functions before the actual quantization process. This preliminary preparation allows the system to determine optimal mapping parameters and zero-points in advance, ensuring precision is maintained even at very low bit widths.

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If mapping parameters are used to transform vectors to lower bit width, then data compression is achieved, but maintaining predetermined precision condition becomes challenging

Engineering Contradiction:
Improvedata sizeVSAvoidprecision condition
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent implements feedback through objective functions that evaluate the precision condition between original first vectors and quantized third vectors. The mapping parameters and zero-points are determined based on these objective functions, creating a feedback loop that ensures the predetermined precision condition is met while achieving compression.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250299029A1Method, apparatus, device and medium for quantizing data
Publication Date: 2025.09.25 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250299029A1 patent drawing
  • US20250299029A1 patent drawing
  • US20250299029A1 patent drawing

AI summary

Methods, apparatuses, devices, and media for quantizing data are provided. In a method, a plurality of first vectors is extracted from a matrix to be quantized. A plurality of objective functions respectively associated with the plurality of first vectors is created. The plurality of objective functions respectively comprises the plurality of first vectors, a plurality of second vectors respectively corresponding to the plurality of first vectors, and a mapping parameter for respectively mapping the plurality of second vectors to a plurality of third vectors. The plurality of second vectors and the mapping parameter are determined based on the plurality of objective functions. For a first vector in the plurality of first vectors, the mapping parameter enables a difference between a third vector corresponding to the first vector in the plurality of third vectors and the first vector to meet a predetermined condition.