Sliding fence product quantification method and device

Through the sliding fence product quantization method, small-value matrix and large-value matrix are generated through initial segmentation and recursive segmentation, and K-means clustering is performed in each subspace through sliding windows. Finally, quantization results are obtained through Count-min operation, which solves the problems of long encoding and decoding time and large compression error of product quantization algorithm, and realizes efficient matrix quantization.

CN120011695APending Publication Date: 2025-05-16PEKING UNIV
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202411945551.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing product quantization algorithms have problems such as long encoding and decoding time and large compression errors.

Method used

The sliding fence product quantization method is used to generate small-value matrix and large-value matrix through initial segmentation and recursive segmentation, and K-means clustering is performed in each subspace through sliding windows, and the quantization results are finally obtained through Count-min operation.

Benefits of technology

While keeping the encoding and decoding time for not long, it effectively reduces compression errors, improves quantization accuracy, and reduces memory overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011695A_ABST
    Figure CN120011695A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of information, and particularly relates to a sliding fence product quantification method and device. The method comprises the following steps: obtaining N input D-dimensional vectors which are regarded as an N * D matrix, and carrying out initial segmentation to obtain a small-value matrix, a large-value matrix and an indication matrix; then recursive segmentation is carried out, a small-value matrix and a large-value matrix which are obtained after recursive segmentation are combined into a new matrix, and subspaces are divided in a sliding window mode; k-means clustering is executed in each subspace, and a cluster center is used as a codebook to encode all vectors in the subspace; and obtaining a final code of each element through a Count-min operation, and combining the final codes to obtain a quantization result. According to the method, improvement and optimization of a product quantization algorithm are realized, the memory overhead can be effectively reduced on the premise of ensuring the precision, and the method can be widely applied to the fields of large language model weight quantization, vector database management, KV cache optimization, image compression, image compression and the like which need efficient matrix quantization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of information technology, and in particular relates to a sliding barrier product quantization method and device. Background Art

[0002] In the development and application of large language models (LLMs) such as ChatGPT, the storage and computational efficiency of matrices is a key issue. These models have a huge number of parameters, and the storage and operation of large matrices bring a high computational burden. Matrix quantization technology records the numerical information of the matrix in a more compact form to reduce storage requirements, and the original matrix can be restored by dequantization when used. This process plays an important role in applications such as LLM weight quantization, KV (Key-Value) cache quantization, vector database management, graph compression, and image compression.

[0003] The RTN quantization algorithm is a traditional matrix element quantization technique. Generally speaking, given a matrix containing N elements, the elements are represented by A1, A2, A3...A N The RTN algorithm performs a simple rounding operation to convert each element A k Mapped to its nearest discrete value set. The core operation in this process is to round the matrix elements to a predefined quantization interval, usually using formula A k ≈Round(A k ), where the Round(·) function is used to map the element value to the closest quantization interval. RTN is simple and efficient to calculate, but the size of the quantization interval often depends on the range of the values ​​in the matrix. When there are outliers in the values, the quantization error may be large.

[0004] Vector quantization is a traditional vector compression technique. Generally speaking, given N D-dimensional real vectors V1, V2, ..., V N , the vector quantization algorithm generates a quantizer, which converts V1, V2, ..., V N A vector V in k Mapping to codebook A codeword c(i(V k )). The function i(·) is called the encoder, and the function c(·) is called the decoder. Q(V k ) as c(i(V k )) is the abbreviation of V k Quantization distortion (or compression error) is defined as MSE(Q) = ∑ 1≤k≤N |V k -Q(V k )|2 The lower the quantization distortion, the higher the quantization accuracy of the quantization algorithm. A good quantization algorithm often has lower quantization distortion.

[0005] K-means algorithm is a mainstream clustering algorithm. K-means divides N vectors into K clusters. Each vector is assigned to the nearest cluster center. The sum of the distances from each vector to the cluster center can be gradually reduced by repeatedly executing the process of "taking the mean of the vectors in the cluster as the new cluster center - re-dividing the cluster". K-means algorithm can be regarded as a vector quantization algorithm. The codebook is K vectors representing cluster centers, V1, V2, ..., V N Each vector in is encoded into an integer between 1 and K. Figure 1 This is a schematic diagram of the effect of the K-means algorithm.

[0006] The product quantization algorithm is also a vector quantization algorithm. Its process is as follows: Figure 2 As shown. For N D-dimensional real vectors V1, V2, ..., V N , the product quantization algorithm first divides each vector into M sub-vectors, each sub-vector has a dimension of D / M. Therefore, for N D-dimensional real vectors V1, V2, ..., V N , the N subvectors corresponding to the dimension form a subspace with a dimension of D / M. After completing the subspace division, the product quantization algorithm uses the K-means algorithm to cluster in each subspace to obtain the codebook and encoding. The codebook and encoding of each subspace are combined to obtain the quantization results of these N D-dimensional real vectors. Assume that V k =(V k1 , V k2 , ..., V kM ) is divided into M sub-vectors V k1 , V k2 , ..., V kM (each is D / M dimensional), V k1 , V k2 , ..., V kM The approximate estimates generated by the K-means algorithm are Q1(V k1 ), Q2(V k2 ), ..., Q M (V kM ), then V k The approximate estimate under the product quantization algorithm is Q(V k )=(Q1(V k1 ), Q2(V k2 ), ..., Q M (V kM )).

[0007] The main problems with the existing product quantization algorithm are that the encoding and decoding time is too long and the compression error is large. Summary of the invention

[0008] The present invention mainly solves the problem of improving and optimizing the product quantization algorithm, specifically, minimizing the compression error as much as possible while keeping the encoding and decoding time not too long.

[0009] The technical solution adopted by the present invention is as follows:

[0010] A sliding fence product quantization method comprises the following steps:

[0011] Get the input N D-dimensional vectors;

[0012] Considering N D-dimensional vectors as N×D matrices, and performing initial segmentation on the N×D matrices; the initial segmentation includes: considering two adjacent elements in the matrix as a pair of elements, placing the element with a smaller value in each pair of elements into a small value matrix, and placing the element with a larger value in a large value matrix, and using indicator variables to record the relative size relationship of each pair of elements, and storing the indicator variables in a first indicator matrix;

[0013] Recursively partition the small value matrix and the large value matrix after the initial partition, and generate a new small value matrix, a large value matrix and a first indicator matrix by continuing to compare and partition element pairs;

[0014] The small value matrix and the large value matrix obtained after recursive segmentation are merged into a new N×D matrix, and the merged matrix is ​​divided into subspaces by sliding windows;

[0015] K-means clustering is performed in each divided subspace, and the cluster centers obtained by clustering are used as codebooks to encode all vectors in the subspace;

[0016] The final code of each element in the merged matrix is ​​obtained through the Count-min operation, and the final codes of each element are merged to obtain the quantization result.

[0017] Furthermore, the initial segmentation also includes: assuming that the position of any element in the original matrix is ​​(a, b), then its corresponding position in the small value matrix or the large value matrix is ​​(a, [b / 2]), and the indicator variable is stored in (a, [b / 2]) of the first indicator matrix, where [·] represents the integer function; traversing the element pairs of the entire matrix in order to generate a small value matrix, a large value matrix and a first indicator matrix of size N×D / 2.

[0018] Furthermore, by specifying the number of recursive partitioning, a tree-like partitioning structure is formed. Assuming the number of recursive partitioning is l, the total size of the first indicator matrix is ​​l×N×D / 2.

[0019] Furthermore, the subspace division of the merged matrix by means of sliding windows includes: setting the window size and the distance of each sliding, taking the vector in each sliding window as a subspace, mapping each element of the merged matrix to a different subspace to form a copy of the element, and each element of the merged matrix has multiple copies.

[0020] Furthermore, the Count-min operation includes: for each element of the merged matrix, calculating the distance between the element and the cluster center of the subspace where each of its copies is located, recording the subspace number of the copy closest to the element at the corresponding position of the second exponential matrix, and taking the code of the copy closest to the element as the final code of the element.

[0021] An image compression method based on sliding barrier product quantization comprises the following steps:

[0022] Performing initial segmentation on the original image to be compressed; the initial segmentation includes: regarding the image as a pixel matrix of size N×D, regarding each pair of adjacent pixels as an element pair, storing the smaller value in the element pair into a small value matrix, storing the larger value into a large value matrix, and generating a first indicator matrix for recording the relative size relationship of each pair of pixels;

[0023] Recursively split the small value matrix and the large value matrix to form a new small value matrix, a large value matrix and a corresponding first indicator matrix;

[0024] The small value matrix and the large value matrix obtained after recursive segmentation are merged into a new N×D matrix, and the merged matrix is ​​divided into subspaces by sliding windows. The pixel vector in each sliding window is regarded as a subspace, and multiple copies are generated for each pixel.

[0025] K-means clustering is performed in each divided subspace, and the cluster centers obtained by clustering are used as codebooks to encode all vectors in the subspace;

[0026] The final code of each element in the merged matrix is ​​obtained through the Count-min operation, and the final codes of each element are merged to obtain the quantization result. The quantization result is dequantized to restore the approximate original image, thereby achieving image compression.

[0027] A method for storing weight quantization of a large language model based on sliding barrier product quantization comprises the following steps:

[0028] Performing an initial segmentation on the weight matrix of the large language model; the initial segmentation includes: regarding the weight matrix of the large language model as an N×D matrix, regarding each pair of adjacent weights as an element pair, placing the smaller value in the element pair into a small value matrix, placing the larger value into a large value matrix, and generating a first indicator matrix for recording the size relationship of each pair of weights;

[0029] Recursively split the small value matrix and the large value matrix to generate new small value matrix and large value matrix, and generate a corresponding first indicator matrix;

[0030] The small value matrix and the large value matrix obtained after recursive segmentation are merged into a new N×D matrix. The merged weight matrix is ​​divided into subspaces by sliding windows. The weight vector in each sliding window is used as a subspace, and multiple copies are generated for each weight.

[0031] K-means clustering is performed in each divided subspace, and the cluster centers obtained by clustering are used as codebooks to encode all vectors in the subspace;

[0032] The final encoding of each element in the merged matrix is ​​obtained through the Count-min operation, and the final encodings of each element are merged to obtain the quantization result of the entire weight matrix, thereby realizing the weight quantization storage of the large language model.

[0033] A sliding fence product quantization device, comprising:

[0034] An input module is used to obtain N D-dimensional vectors as input;

[0035] An initial segmentation module is used to regard N D-dimensional vectors as N×D matrices and perform initial segmentation on the N×D matrices; the initial segmentation includes: regarding two adjacent elements in the matrix as a pair of elements, placing the element with a smaller value in each pair of elements into a small value matrix, placing the element with a larger value into a large value matrix, and using indicator variables to record the relative size relationship of each pair of elements, and storing the indicator variables in a first indicator matrix;

[0036] A recursive segmentation module is used to recursively segment the small value matrix and the large value matrix after the initial segmentation, and generate a new small value matrix, a large value matrix and a first indicator matrix by continuing to compare and segment element pairs;

[0037] The subspace partitioning module is used to merge the small value matrix and the large value matrix obtained after recursive partitioning into a new N×D matrix, and divide the merged matrix into subspaces by sliding windows;

[0038] The clustering module is used to perform K-means clustering in each divided subspace, and use the cluster centers obtained by clustering as the codebook to encode all vectors in the subspace;

[0039] The count-min operation and merging module is used to obtain the final code of each element in the merged matrix through the count-min operation, and merge the final codes of each element to obtain a quantization result.

[0040] The beneficial effects of the present invention are as follows:

[0041] The present invention realizes the improvement and optimization of the product quantization algorithm, can minimize the compression error while keeping the encoding and decoding time not too long, and can effectively reduce the memory overhead under the premise of ensuring accuracy. The present invention can be widely used in fields that require efficient matrix quantization, such as large language model (LLM) weight quantization, vector database management, KV cache optimization, graph compression, image compression, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a schematic diagram of the K-means algorithm effect.

[0043] Figure 2 It is a flowchart of the product quantization algorithm.

[0044] Figure 3 It is a schematic diagram of the flow of the sliding barrier product quantization algorithm of the present invention. DETAILED DESCRIPTION

[0045] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below through specific embodiments and drawings.

[0046] The present invention proposes an improved product quantization algorithm, namely, a sliding fence product quantization algorithm, which first adjusts the local order of matrix elements, then divides the subspace in the form of a sliding window, completes K-means clustering in each subspace, and finally takes the highest precision of multiple quantized copies of each subvector as the final quantization result. In the "sliding fence" of the present invention, "fence" refers to isolating and grouping the vector elements in the original matrix according to their relative sizes, i.e., the initial segmentation and recursive segmentation steps in the algorithm flow; "sliding" refers to dividing the subspace in the form of a sliding window pane.

[0047] The process of the sliding fence product quantization algorithm of the present invention is as follows: Figure 3 Assume that the input is N D-dimensional vectors, the process of the core method of the present invention is as follows:

[0048] (1) Initial segmentation: Treat N D-dimensional vectors as an N×D matrix, and treat two adjacent elements of the matrix as a pair. Put the element with relatively smaller value into the small value matrix (small part), and the larger element into the large value matrix (big part). Use a 1-bit indicator variable to record the relative size relationship of the pair of elements. Suppose the position of any element in the original matrix is ​​(a, b), then its corresponding position in the small value matrix or the large value matrix is ​​(a, [b / 2]), and the indicator variable is stored in the indicator matrix 1 (a, [b / 2]), where [·] represents the rounding function. Traverse the element pairs of the entire matrix in order to generate the small value matrix or the large value matrix and the indicator matrix 1 (Indicator Map) of size N×D / 2.

[0049] Among them, two adjacent elements of a matrix refer to two adjacent elements in the matrix. Specifically, the matrix element is denoted as a ij (i≤N, j≤D, i and j are positive integers), i,2k-1 and Considered as a pair, compare a i,2k-1 and a i,2k The relative size of .

[0050] (2) Recursive segmentation: For the small value matrix or large value matrix after the initial segmentation, the element pair comparison and segmentation in the above step (1) can be continued to generate a new small value matrix, large value matrix, and indicator matrix 1. This process can specify the number of recursive segmentations to form Figure 3 The tree-like partitioning structure shown in . The indicator matrix 1 of each layer of partitioning can be synthesized into an indicator matrix 1 of size N×D / 2. Assuming that the number of recursive partitioning is l, the total size of the indicator matrix 1 is l×N×D / 2.

[0051] (3) Divide the subspace by sliding window: merge the matrices after the split operation into one matrix, and the size of the matrix is ​​restored to N×D. Set the window size and the distance of each slide, and divide the merged matrix into subspaces by sliding window, that is, take the vector in each sliding window as a subspace. Thus, each element of the merged matrix will be mapped to a different subspace, and the elements in different subspaces are called copies of the element, and each element has multiple copies. For example Figure 3 The window size shown in the figure is 4 and the sliding distance each time is 2.

[0052] (4) K-means clustering: K-means clustering is performed in each subspace divided in the previous step. The cluster center is used as the codebook, and all vectors in the subspace are encoded to obtain the code.

[0053] (5) Count-min operation and merging: For each element of the matrix merged in step (3), calculate the distance between the element and the cluster center of the subspace where each of its copies is located, record the subspace number of the copy closest to the element in the corresponding position in the index matrix 2, and use the code of the copy closest to the element as the final code of the element. This step is called the "Count-min" operation. The codes of each element are merged to obtain the final quantization result.

[0054] For the quantization result obtained by the above-mentioned method of the present invention, when it is necessary to restore the original matrix by inverse quantization in the actual application process, for each position of the matrix, first find the corresponding subspace according to the indicator matrix 2, then query the codebook vector of the subspace according to its encoding, restore the value of the position, and obtain the preliminary restored matrix. For the preliminary restored matrix, the position relationship of the matrix elements is recursively restored according to the value of the indicator matrix 1.

[0055] The present invention can be widely used in fields requiring efficient matrix quantization, such as large language model (LLM) weight quantization, vector database management, KV cache optimization, graph compression, and image compression.

[0056] For example, one embodiment of the present invention provides an image compression method based on sliding barrier product quantization, comprising the following steps:

[0057] 1) Initial segmentation: The original image to be compressed is processed and regarded as a pixel matrix of size N×D. Each pair of adjacent pixels is compared as a pair, and the smaller value is stored in the small value matrix and the larger value is stored in the large value matrix. At the same time, an indicator matrix 1 is generated to record the relative size relationship of each pair of pixels.

[0058] 2) Recursive segmentation: Continue the above initial segmentation operation on the small value matrix and the large value matrix to form a new small value matrix, a large value matrix and the corresponding indicator matrix 1. This process can be recursively performed multiple times to form a tree-like segmentation structure, gradually refining the information of different parts of the image. Finally, the segmentation results of each layer are merged into a total indicator matrix 1 of size l×N×D / 2.

[0059] 3) Divide the subspace by sliding window: restore the recursively partitioned matrix to an N×D matrix, set the size of the sliding window and the step size of each sliding. Divide the matrix into multiple subspaces by sliding window operation, and use the pixel vector in each sliding window as a subspace for the next quantization process. In this way, multiple copies will be generated at each pixel position for subsequent selection of the optimal encoding.

[0060] 4) K-Means clustering: K-means clustering is performed in each subspace divided in the previous step. The cluster center is used as the codebook, and all vectors in the subspace are encoded to obtain the code.

[0061] 5) Count-min operation and merging: For each element of the matrix merged in step (3), calculate the distance between the element and the cluster center of the subspace where each of its copies is located, record the subspace number of the copy closest to the element in the corresponding position in the index matrix 2, and use the code of the copy closest to the element as the final code of the element. This step is called the "Count-min" operation. Merge the codes of each element to obtain the final quantization result. By dequantizing the encoded data, the approximate original image can be restored, thereby achieving image compression.

[0062] For another example, another embodiment of the present invention provides a large language model (LLM) weight quantization storage method based on sliding barrier product quantization, comprising the following steps:

[0063] 1) Initial segmentation: The weight matrix of the large-scale language model is regarded as an N×D matrix, and each pair of adjacent weights is regarded as an element pair. By comparison, smaller values ​​are placed in the small value matrix, larger values ​​are placed in the large value matrix, and an indicator matrix 1 is generated to record the size relationship.

[0064] 2) Recursive partitioning: Continue to recursively partition the small value matrix and the large value matrix to generate multiple layers of small value matrices and large value matrices, and generate the corresponding indicator matrix 1. The number of layers of recursive partitioning is specified by the user to form a gradually refined tree-like weight representation.

[0065] 3) Divide the subspace by sliding window: merge the recursively segmented matrices to restore them to N×D size, and perform sliding window segmentation on the weight matrix by setting the sliding window size and sliding step size to obtain multiple subspaces. The weight vector in each sliding window is used as a subspace, so as to perform more fine-grained quantization on the weight matrix.

[0066] 4) K-Means clustering: K-means clustering is performed in each subspace divided in the previous step. The cluster center is used as the codebook, and all vectors in the subspace are encoded to obtain the code.

[0067] 5) Count-min operation and merging: For each element of the matrix merged in step (3), calculate the distance between the element and the cluster center of the subspace where each of its copies is located, record the subspace number of the copy closest to the element in the corresponding position in the index matrix 2, and use the code of the copy closest to the element as the final code of the element. This step is called the "Count-min" operation. The codes of each element are merged to obtain the final quantization result of the entire weight matrix, thereby realizing the weight quantization storage of large-scale language models.

[0068] Another embodiment of the present invention provides a sliding barrier product quantization device, comprising:

[0069] An input module is used to obtain N D-dimensional vectors as input;

[0070] An initial segmentation module is used to regard N D-dimensional vectors as N×D matrices and perform initial segmentation on the N×D matrices; the initial segmentation includes: regarding two adjacent elements in the matrix as a pair of elements, placing the element with a smaller value in each pair of elements into a small value matrix, placing the element with a larger value into a large value matrix, and using indicator variables to record the relative size relationship of each pair of elements, and storing the indicator variables in a first indicator matrix;

[0071] A recursive segmentation module is used to recursively segment the small value matrix and the large value matrix after the initial segmentation, and generate a new small value matrix, a large value matrix and a first indicator matrix by continuing to compare and segment element pairs;

[0072] The subspace partitioning module is used to merge the small value matrix and the large value matrix obtained after recursive partitioning into a new N×D matrix, and divide the merged matrix into subspaces by sliding windows;

[0073] The clustering module is used to perform K-means clustering in each divided subspace, and use the cluster centers obtained by clustering as the codebook to encode all vectors in the subspace;

[0074] The count-min operation and merging module is used to obtain the final code of each element in the merged matrix through the count-min operation, and merge the final codes of each element to obtain a quantization result.

[0075] The division of the above modules is only an example. In actual applications, the above functions can be assigned to different functional modules as needed to complete all or part of the functions described in the above method. The specific working process of each module can refer to the corresponding process in the above method embodiment, which will not be repeated here.

[0076] Another embodiment of the present invention provides a computer device (computer, server, smart phone, etc.), which includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the method of the present invention.

[0077] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, magnetic disk, optical disk), wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the steps of the method of the present invention are implemented.

[0078] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. It can be understood by those skilled in the art that various replacements, changes and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the contents disclosed in the embodiments of this specification, and the scope of protection of the present invention shall be subject to the scope defined in the claims.

Claims

1. A sliding barrier product quantization method, characterized in that: The following steps are involved: Get the input N D-dimensional vectors; Considering N D-dimensional vectors as N×D matrices, and performing initial segmentation on the N×D matrices; the initial segmentation includes: considering two adjacent elements in the matrix as a pair of elements, placing the element with a smaller value in each pair of elements into a small value matrix, and placing the element with a larger value in a large value matrix, and using indicator variables to record the relative size relationship of each pair of elements, and storing the indicator variables in a first indicator matrix; Recursively partition the small value matrix and the large value matrix after the initial partition, and generate a new small value matrix, a large value matrix and a first indicator matrix by continuing to compare and partition element pairs; The small value matrix and the large value matrix obtained after recursive segmentation are merged into a new N×D matrix, and the merged matrix is ​​divided into subspaces by sliding windows; K-means clustering is performed in each divided subspace, and the cluster centers obtained by clustering are used as codebooks to encode all vectors in the subspace; The final code of each element in the merged matrix is ​​obtained through the Count-min operation, and the final codes of each element are merged to obtain the quantization result.

2. The method according to claim 1, characterized in that The initial segmentation also includes: assuming that the position of any element in the original matrix is ​​(a, b), then its corresponding position in the small value matrix or the large value matrix is ​​(a, [b / 2]), and the indicator variable is stored in (a, [b / 2]) of the first indicator matrix, where [·] represents a rounding function; traversing the element pairs of the entire matrix in order to generate a small value matrix, a large value matrix and a first indicator matrix of size N×D / 2.

3. The method according to claim 1, characterized in that By specifying the number of recursive partitions, a tree-like partition structure is formed. Assuming the number of recursive partitions is l, the total size of the first indicator matrix is ​​l×N×D / 2.

4. The method according to claim 1, characterized in that: The subspace division of the merged matrix by sliding window includes: setting the window size and the distance of each sliding, taking the vector in each sliding window as a subspace, mapping each element of the merged matrix to a different subspace to form a copy of the element, and each element of the merged matrix has multiple copies.

5. The method according to claim 4, characterized in that The Count-min operation includes: for each element of the merged matrix, calculating the distance between the element and the cluster center of the subspace where each of its copies are located, recording the subspace number of the copy closest to the element in the corresponding position of the second exponential matrix, and using the code of the copy closest to the element as the final code of the element.

6. An image compression method based on sliding barrier product quantization, characterized in that: The following steps are involved: Performing initial segmentation on the original image to be compressed; the initial segmentation includes: regarding the image as a pixel matrix of size N×D, regarding each pair of adjacent pixels as an element pair, storing the smaller value in the element pair into a small value matrix, storing the larger value into a large value matrix, and generating a first indicator matrix for recording the relative size relationship of each pair of pixels; Recursively split the small value matrix and the large value matrix to form a new small value matrix, a large value matrix and a corresponding first indicator matrix; The small value matrix and the large value matrix obtained after recursive segmentation are merged into a new N×D matrix, and the merged matrix is ​​divided into subspaces by sliding windows. The pixel vector in each sliding window is regarded as a subspace, and multiple copies are generated for each pixel. K-means clustering is performed in each divided subspace, and the cluster centers obtained by clustering are used as codebooks to encode all vectors in the subspace; The final code of each element in the merged matrix is ​​obtained through the Count-min operation, and the final codes of each element are merged to obtain the quantization result. The quantization result is dequantized to restore the approximate original image, thereby achieving image compression.

7. A large language model weight quantization storage method based on sliding fence product quantization, characterized in that: The following steps are involved: Performing an initial segmentation on the weight matrix of the large language model; the initial segmentation includes: regarding the weight matrix of the large language model as an N×D matrix, regarding each pair of adjacent weights as an element pair, placing the smaller value in the element pair into a small value matrix, placing the larger value into a large value matrix, and generating a first indicator matrix for recording the size relationship of each pair of weights; Recursively split the small value matrix and the large value matrix to generate new small value matrix and large value matrix, and generate a corresponding first indicator matrix; The small value matrix and the large value matrix obtained after recursive segmentation are merged into a new N×D matrix. The merged weight matrix is ​​divided into subspaces by sliding windows. The weight vector in each sliding window is used as a subspace, and multiple copies are generated for each weight. K-means clustering is performed in each divided subspace, and the cluster centers obtained by clustering are used as codebooks to encode all vectors in the subspace; The final encoding of each element in the merged matrix is ​​obtained through the Count-min operation, and the final encodings of each element are merged to obtain the quantization result of the entire weight matrix, thereby realizing the weight quantization storage of the large language model.

8. A sliding fence product quantization device, characterized in that: include: An input module is used to obtain N D-dimensional vectors as input; An initial segmentation module is used to regard N D-dimensional vectors as N×D matrices and perform initial segmentation on the N×D matrices; the initial segmentation includes: regarding two adjacent elements in the matrix as a pair of elements, placing the element with a smaller value in each pair of elements into a small value matrix, placing the element with a larger value into a large value matrix, and using indicator variables to record the relative size relationship of each pair of elements, and storing the indicator variables in a first indicator matrix; A recursive segmentation module is used to recursively segment the small value matrix and the large value matrix after the initial segmentation, and generate a new small value matrix, a large value matrix and a first indicator matrix by continuing to compare and segment element pairs; The subspace partitioning module is used to merge the small value matrix and the large value matrix obtained after recursive partitioning into a new N×D matrix, and divide the merged matrix into subspaces by sliding windows; The clustering module is used to perform K-means clustering in each divided subspace, and use the cluster centers obtained by clustering as the codebook to encode all vectors in the subspace; The count-min operation and merging module is used to obtain the final code of each element in the merged matrix through the count-min operation, and merge the final codes of each element to obtain a quantization result.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • 3D-HEVC rapid CU segmentation method based on image entropy K-means clustering

    CN111741313A

  • Metagenome species reconstruction method based on pre-training and deep clustering

    CN115579068A

  • Method for constructing prediction and evaluation model based on postherpetic neuralgia

    CN117894477A

  • Low-rank tensor data compression and missing value recovery method and system

    CN117972323A

  • Approximate nearest neighbor search method based on local vector quantization

    CN118013085A