High-precision quantification method for non-uniformly distributed data

Through the dual-layer optimization mechanism of dynamic programming and particle swarm optimization, the reconstruction error and resource waste problems of traditional uniform quantization methods in non-uniform data are solved, and high-precision quantization and efficient compression are achieved, which are suitable for edge computing and AI chips.

CN120354065AActive Publication Date: 2025-07-22GANYUE MEDICAL TECH (CHENGDU) CO LTD

Patent Information

Application Number
CN202510845442.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-22
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

Traditional uniform quantization methods cannot adapt to non-uniform distributed data, resulting in significant increase in reconstruction errors and waste of codebook resources, making it difficult to meet the real-time processing requirements of edge devices and AI chips.

Method used

The dynamic programming algorithm is used for local homogenization, and the quantization interval endpoints are iteratively tuned with the particle swarm optimization algorithm to minimize mean square error to build an efficient codebook.

Benefits of technology

Significantly reduce quantization errors, improve compression efficiency, meet the real-time requirements of edge computing nodes and AI chips, and reduce hardware power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354065A_ABST
    Figure CN120354065A_ABST
Patent Text Reader

Abstract

The invention discloses a high-precision quantization method for non-uniformly distributed data, and aims to solve the problems that reconstruction errors are remarkably increased and codebook resources are wasted when a traditional uniform quantization algorithm processes complex data distribution. According to the scheme, firstly, a dynamic programming algorithm is used for carrying out local homogenization processing on data, and optimal segmentation is achieved by minimizing the sum of squares of a weighting range; and then introducing a particle swarm optimization algorithm to carry out iterative tuning on the end points of the quantization interval, and constructing an efficient codebook by taking the minimization of a mean square error (MSE) as a target. Experimental results show that the quantization errors of the method on a 128-dimensional data set and a 420-dimensional data set are respectively reduced by 93.52% and 98.05% compared with those of a traditional method, meanwhile, hardware compatibility is kept, and the method is suitable for scenes such as edge computing nodes, AI chips and the like which have strict requirements on compression efficiency and real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence and machine learning, and specifically relates to the quantization compression theory and technology for non-uniformly distributed data. Background Art

[0002] With the rapid development of edge computing, the Internet of Things, and deep learning, the demand for efficient storage and transmission of high-dimensional non-uniform data (such as gradient features of images, word vectors in natural language processing, spatio-temporal sequences collected by sensors, etc.) has become increasingly urgent. Traditional uniform quantization methods cannot adapt to the non-uniform distribution characteristics of data, and it is difficult to meet the real-time processing requirements of edge devices (such as intelligent cameras, wearable sensors) and AI chips (such as NPUs, TPUs) in terms of compression efficiency and accuracy.

[0003] Traditional uniform quantization assumes that data follows a uniform distribution and uses equally spaced quantization intervals. However, actual data often exhibits complex distribution characteristics. When the data has a multi-modal distribution, the histogram of image features usually has multiple peaks, and uniform quantization will result in too large quantization intervals in the peak regions, significantly increasing the reconstruction error; when the data has a skewed distribution, the distribution of high-frequency and low-frequency words in text word vectors is extremely unbalanced, and uniform quantization will allocate too many codebooks in the low-frequency word regions, causing resource waste. Experiments show that in non-uniform distribution scenarios, the reconstruction error of uniform quantization is relatively high, and about 65% of the codebooks only correspond to less than 5% of the sample volume. Summary of the Invention

[0004] To solve the above problems, the present invention proposes a high-precision quantization method for non-uniformly distributed data, which realizes the balance between quantization error and codebook efficiency through a two-layer optimization mechanism of "dynamic programming segmentation + swarm intelligence optimization".

[0005] A high-precision quantization method for non-uniformly distributed data includes the following steps: S1. Data preprocessing and segmentation: For high-dimensional non-uniform data, first sort and remove duplicates the feature values of each dimension to generate an ordered unique value sequence , and at the same time count the occurrence frequencies of each value ; Based on the dynamic programming algorithm, divide the unique value sequence into mutually exclusive sub-intervals , and the division objective is to minimize the weighted range sum of squares, that is: , where represents the number of samples in the sub-interval ; S2. Initialization of quantization parameters: For each sub-interval , use the Min-Max normalization method to initialize the quantization interval endpoints , where , ; The quantization step size is defined as: , In the formula, is the initial quantization step size, is the quantization bit width, which determines the quantization level number of each sub-interval; S3. Particle swarm intelligence optimization parameters: Taking the minimization of the mean square error MSE as the optimization goal, the particle swarm optimization algorithm is used to globally search for the quantization interval endpoints ; The position parameter of the particle swarm is defined as the interval center point and the width , to decouple the optimization of the interval position and range, where is the lower bound of the quantization interval, is the upper bound of the quantization interval; The search space expands with the initial solution as the center, specifically: , where is the initial interval width, is the expansion factor, used to balance the search range and accuracy; The initial particles are generated by grid sampling, and the grid density is set to to ensure the uniform distribution of the initial solution, is the grid particle width; The fitness function is defined as the reciprocal of MSE, guiding the particles to iterate towards the region with the minimum error; The particle velocity update formula is: , where is the velocity of the particle at time t, is the mass at time t, and the inertia weight decays linearly with the number of iterations; is the learning factor, is a random number in the interval [0,1], and are the individual optimal position and the global optimal position of the particle respectively; S4. Data quantization: According to the optimized parameters , where is the quantization step size, map the data in the sub-interval to the discrete codebook space, and the specific quantization formula is: , where is the data corresponding quantization code, is the rounding operation.

[0006] Furthermore, the state transition equation of the dynamic programming algorithm is as follows: , where represents the minimum weighted range sum of squares for partitioning the first unique values into intervals; is the th unique value after sorting; is the cumulative frequency of the first values.

[0007] Furthermore, the parameter search space of particle swarm optimization is dynamically adjusted by an expansion factor ; when takes a small value, the search focuses near the initial solution for local fine-tuning; when is large, the search range expands for global exploration; a grid sampling strategy for initial particles is used to ensure uniform coverage of the solution space and avoid being trapped in local optima.

[0008] Furthermore, an inertia weight decay strategy is adopted in the particle swarm velocity update formula. A larger inertia weight is used in the early stage to encourage global search by particles, and a smaller weight is adopted in the later stage to guide particles to focus on local optimal solutions.

[0009] The beneficial effects of the present invention include: 1) Precision improvement: On the Scale-Invariant Feature Transform dataset (SIFT dataset), the mean square error of the present invention is significantly reduced, showing a significant error reduction of 93.52% compared with uniform quantization; on the audio processing dataset (MSong dataset), the mean square error is significantly reduced, and the quantization error is reduced by 98.05%; 2) Efficiency optimization: Particle swarm optimization supports GPU parallel acceleration to meet real-time quantization requirements; 3) Hardware compatibility: The quantization process only depends on addition, subtraction, multiplication, and division and integer operations, and can run on a microcontroller unit without a floating-point operation unit, with the power consumption reduced by more than 70% compared with floating-point operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a comparison chart of the quantization error results between the present invention and existing methods. DETAILED DESCRIPTION OF THE INVENTION

[0011] The technical solution of the present invention will be described in detail below in conjunction with specific embodiments.

[0012] A high-precision quantization method for non-uniformly distributed data includes the following steps: S1. Data preprocessing and segmentation: For high-dimensional non-uniform data, first sort and de-duplicate the eigenvalue of each dimension in ascending order to generate an ordered unique value sequence , and at the same time count the occurrence frequency of each value ; Based on the dynamic programming algorithm, divide the unique value sequence into mutually exclusive sub-intervals , and the division goal is to minimize the weighted range sum of squares, that is: , where represents the number of samples in the sub-interval ; S2. Initialization of quantization parameters: For each sub-interval , use the Min-Max normalization method to initialize the quantization interval endpoints , where , ; The quantization step size is defined as: , In the formula, is the initial quantization step size, is the quantization bit width, which determines the quantization level number of each sub-interval; S3. Particle swarm intelligent optimization parameters: With the minimization of the mean square error MSE as the optimization goal, use the particle swarm optimization algorithm to perform a global search on the quantization interval endpoints ; The position parameter of the particle swarm is defined as the center point and width of the interval to decouple the optimization of the interval position and range, where is the lower bound of the quantization interval, is the upper bound of the quantization interval; The search space expands with the initial solution as the center, specifically: , where is the initial interval width, is the expansion factor, which is used to balance the search range and accuracy; The initial particles are generated by grid sampling, and the grid density is set to to ensure the uniform distribution of the initial solution, is the grid particle width; The fitness function is defined as the reciprocal of MSE to guide the particles to iterate towards the region with the minimum error; The particle velocity update formula is: , where is the velocity of the particle at time t, is the mass at time t, and the inertia weight decays linearly with the number of iterations; is the learning factor, is a random number in the interval [0, 1], and are the individual optimal position and the global optimal position of the particle respectively; S4. Data Quantification: According to the optimized parameter , where is the quantization step size, map the data within the sub-interval to the discrete codebook space, and the specific quantization formula is: , where is the quantization code corresponding to the data , is the rounding operation.

[0013] Furthermore, the state transition equation of the dynamic programming algorithm is: , where represents the minimum weighted range sum of squares for dividing the first unique values into intervals; is the rd sorted unique value; is the cumulative frequency of the first values.

[0014] Furthermore, the parameter search space of particle swarm optimization is dynamically adjusted by the expansion factor ; when takes a small value, the search focuses near the initial solution for local fine-tuning; when is large, the search range expands for global exploration; use the grid sampling strategy of the initial particles to ensure uniform coverage of the solution space and avoid falling into local optima.

[0015] Furthermore, an inertia weight decay strategy is adopted in the particle swarm velocity update formula. A larger inertia weight is used in the early stage to encourage the particles to search globally, and a smaller weight is adopted in the later stage to guide the particles to focus on the local optimal solution.

[0016] Example 1

[0017] Taking the global vector dataset (GloVe dataset) as an example, the image feature quantization process is as follows:

[0018] Step 1. Input data: 1 million 300-dimensional image feature vectors, and the eigenvalue range is [0, 1000].

[0019] Step 2. Dynamic programming segmentation: 1) Set the number of segments M = 4, and divide the data into 4 sub - intervals with similar frequencies; 2) Calculate the optimal segmentation points through the state - transfer equation, and obtain the optimal segmentation points [32, 128, 200]. Divide the data into [0, 32], (32, 128], (128, 200], (200, 255]; 3) After segmentation, the dynamic ranges of each sub - interval are 32, 96, 72, and 55 respectively. The average dynamic range is compressed to 31.6% compared with traditional uniform segmentation, reducing the local quantization pressure.

[0020] Step 3. Particle swarm optimization: Initialize the parameters: (The expansion range is ±20% of the initial interval width), G = 16, generate 256 initial particles, T = 128 iterations; The inertia weight linearly decays from 0.9 to 0.4, and the learning factor ; Optimization objective: Optimize for each sub - interval's For example, the initial interval [0, 32] becomes [2, 30] after optimization, the width is compressed to 28, and the step size is adjusted from 4 (when k = 8) to 3.5, making the quantization intervals in the high - frequency central region denser.

[0021] Step 4. Quantization output: Generate 4 independent codebooks, each codebook contains 256 quantization levels; The data compression ratio reaches 8:1, compressed from 2400 bytes to 300 bytes, and the cosine similarity of the features after decompression and the original data reaches 0.987, meeting the requirements of the image retrieval task.

[0022] Example 2

[0023] Complexity analysis:

[0024] Step 1. Dynamic programming segmentation: After optimizing with the Knuth - Yao algorithm, the decision - point search range is reduced from to , where and are the optimal decision points of the previous step, and the time complexity is reduced to O(NM). For the scenario of N = 10 4 , M = 10, the calculation time is about 0.1 second. The space complexity is O(N), only need to store the cumulative frequency array and the current layer and the previous layer of the dynamic programming table , which is suitable for memory - limited devices.

[0025] Step 2. Particle swarm optimization: The initialization complexity is O(G²). When G = 16, 256 particles need to be generated for each sub - interval; The iteration complexity is O(nTG²), where n is the number of samples in the sub - interval. For n = 10 5In the scenario where T = 100, the parallel computing time of a single GPU card is approximately 50 milliseconds.

[0026] Through the double-layer optimization mechanism of "dynamic programming segmentation + swarm intelligence optimization", the present invention achieves the balance between quantization error and codebook efficiency. As Figure 1 shown, the optimized uniform quantization is the method after discarding the segmentation of the segmented optimization uniform quantization method. The experimental results show that the quantization error of the method of the present invention on the scale-invariant feature transform dataset (128 dimensions) and the audio processing dataset (420 dimensions) is reduced by 93.52% and 98.05% respectively compared with the traditional method.

[0027] Although the present invention has been described herein with reference to the illustrative embodiments of the present invention, the above embodiments are only the preferred embodiments of the present invention, and the embodiments of the present invention are not limited by the above embodiments.

[0028] The protection scope of the present invention includes but is not limited to the algorithm flow of the above high-precision quantization method, the combination mode of dynamic programming and particle swarm optimization, and the hardware implementation based on this framework, such as the dedicated integrated circuit IP core, FPGA module, and application scenarios such as data compression of edge devices and model quantization of AI chips. Those skilled in the art can design many other modifications and implementation manners, and these modifications and implementation manners will fall within the scope and spirit of the principles disclosed in this application.

Claims

1. A high-precision quantization method for non-uniformly distributed data, characterized in that, It includes the following steps: S1. Data preprocessing and segmentation: For high-dimensional non-uniform data, first sort the feature values of each dimension in ascending order and remove duplicates to generate an ordered unique value sequence , and at the same time count the occurrence frequency of each value ; Based on the dynamic programming algorithm, the unique value sequence is partitioned into mutually exclusive subintervals , and the partitioning objective is to minimize the weighted sum of squared ranges, i.e.: , Among them, represents the number of samples in the sub-interval ; S2. Quantization parameter initialization: For each sub-interval , initialize the quantization interval endpoints using the Min-Max normalization method , where , ; The quantization step size is defined as: , wherein, is the initial quantization step size, is the quantization bit width, which determines the quantization level number of each sub-interval; S3. Particle Swarm Intelligence Optimization Parameters: With the minimization of the mean square error MSE as the optimization objective, the particle swarm optimization algorithm is used to perform a global search on the quantization interval endpoints ; the position parameter of the particle swarm is defined as the interval center point and the width , to decouple the optimization of the interval position and range, where is the lower bound of the quantization interval, is the upper bound of the quantization interval; the search space is expanded centered on the initial solution, specifically: , Among them, is the initial interval width, is the expansion factor, which is used to balance the search range and accuracy; the initial particles are generated by grid sampling, and the grid density is set to to ensure the uniform distribution of the initial solution, is the grid particle width; the fitness function is defined as the reciprocal of MSE, guiding the particles to iterate towards the region with the minimum error; the particle velocity update formula is: , Among them, is the velocity of the particle at time t, is the mass at time t, and the inertia weight decays linearly with the number of iterations; is the learning factor, is a random number in the interval [0, 1], and are the individual best position and the global best position of the particle, respectively; S4. Data Quantification: According to the optimized parameters , where is the quantization step size, map the data within the sub-interval to the discrete codebook space. The specific quantization formula is as follows: , wherein, is data corresponding quantization code, is a rounding operation.

2. The method according to claim 1, wherein The state transition equation of the dynamic programming algorithm is as follows: , Among them, represents the minimum weighted range sum of squares for dividing the first unique values into intervals; is the th unique value after sorting; is the cumulative frequency of the first values.

3. The method according to claim 1, characterized in that, The parameter search space of particle swarm optimization is dynamically adjusted by an expansion factor ; when takes a small value, the search focuses on the vicinity of the initial solution for local fine-tuning; when is large, the search range expands for global exploration; a grid sampling strategy for the initial particles is used to ensure uniform coverage of the solution space and avoid falling into local optima.

4. The method according to claim 1, wherein An inertia weight decay strategy is adopted in the particle swarm velocity update formula. A larger inertia weight is used in the early stage to encourage the global search of particles, and a smaller weight is adopted in the later stage to guide the particles to focus on the local optimal solution.

Citation Information

Patent Citations

  • Graph neural network compression method and device, electronic equipment and storage medium

    CN115357554A

  • Pipeline magnetic flux leakage internal detection signal quantification method based on neural network

    CN116956711A

  • Quantitative data processing method based on cloud computing

    CN117648552A

  • Lightweight data processing method for resource-constrained industrial platform

    CN117874564A

  • Data selection method and device, storage medium and electronic equipment

    CN118036036A

Cited By

  • Deep learning model segmentation quantification method and system oriented to multi-peak feature distribution

    CN121743813A

  • Deep learning model segmentation quantization method and system for multimodal feature distribution

    CN121743813B