Quick decoding method for Active-Bit codes
By using the Active-Bit encoding method and employing a full subtractor and multiplexer for parallel decoding, the problems of high storage pressure and high computational complexity in LLM calculation are solved, achieving fast decoding and efficient memory usage.
Patent Information
- Application Number
- CN202511252648.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-10-31
AI Technical Summary
The calculation process of LLM generates a large amount of key-value data, resulting in high storage pressure, high computational complexity, and low efficiency of existing decoding methods.
The Active-Bit encoding method is adopted, which uses a full subtractor and a multiplexer for parallel processing. The decoding is calculated by using the difference between the high-active bits and the low-active bits, thus achieving fast decoding.
Decoding is completed within one clock cycle, improving memory usage and computational efficiency during LLM inference while reducing memory consumption and computational complexity.
Smart Images

Figure CN120874907A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data encoding / decoding technology, and in particular to a fast decoding method for Active-Bit encoding. Background Technology
[0002] LLM (Large Language Model) is a complex algorithm model in the field of artificial intelligence based on deep learning technology. With the development of the times, it has been trained on massive amounts of text data and has the function of understanding and generating natural language. It can perform various tasks such as text generation, translation, and question answering.
[0003] The core features of LLM include: the number of model parameters typically reaches hundreds of millions to billions, enabling it to capture subtle patterns and complex rules in language; its training method is based on pre-training on large-scale unlabeled data to obtain general language representations, and fine-tuning for specific tasks to achieve application adaptation; its architecture is advanced, generally adopting the Transformer architecture, using the self-attention mechanism to process sequence data, effectively solving the problem of long-distance dependencies, but generating a large amount of key-value data during the computation process, which puts a great deal of pressure on storage. Summary of the Invention
[0004] The purpose of this invention is to provide a fast decoding method for Active-Bit encoding to solve the problems in the background art.
[0005] To address the aforementioned technical problems, this invention provides a fast decoding method for Active-Bit encoding, comprising: The encoding is compared with a finite number of constant arrays, and the first part of the high-activity bit E is decoded by using a full subtractor and the highest borrow output of each subtractor. Based on the first part of the high-activity bit E variable, select an appropriate difference with the first multiplexer, and then use this difference to select the corresponding low-activity bit from the second multiplexer. Perform bitwise OR logic to complete the decoding to obtain the original data; among them, the smallest active bit 1 is the low-activity bit, and the largest active bit 1 is the high-activity bit.
[0006] In one implementation, the low-activity bits of the original data are obtained through the following steps: The codes are classified into different groups, and a fixed list B is formed by the starting position of the dividing group; A new list B' is obtained by subtracting each value in the fixed list B from the encoded value M; Find the first positive difference from left to right in list B', denoted by j. The j-th bit in the original data is a low-activity bit.
[0007] In one implementation, the highly active bits of the original data are obtained through the following steps: The offset number relative to the leftmost starting position of list B' is obtained based on the location of j in list B', and is called k; The 7th to kth bit of the original data is a highly active bit.
[0008] In one implementation, subtracting each value in the fixed list B from the encoded value M includes: comparing the encoded value M with the values in the fixed list B in parallel, and then selecting the correct output based on the comparison results.
[0009] In one implementation, the condition for finding the first positive number j in list B' from left to right is determined by the most significant bit b of the subtraction result of each number in list B. o Output to complete; put b K o Defined as b corresponding to the most significant bit of the k-th bit of the original data o Output when the encoded value M is less than B k time b K o If it is 0, then b is 0; otherwise, b is 0. K o =1; where B k This is the (k+1)th value from left to right in the fixed list B.
[0010] In one implementation, except for the borrowed data between subtractors, all other parts are processed in parallel, so they can be completed in one clock cycle.
[0011] This invention provides a fast decoding method for Active-Bit encoding. During LLM inference, data bytes stored in the key-value cache often have only two active bits, making AB=2 encoding suitable. This encoding itself can achieve a compression ratio of 62.5%, significantly improving memory usage and computational efficiency during LLM inference. As sequence length increases, the key-value cache grows linearly, leading to significant memory consumption and computational complexity. The decoding method of this invention is simple and fast, completing within one clock cycle. It can be seamlessly embedded into the tensor multiplier pipeline for parallel execution, greatly improving memory usage efficiency and overall computational latency during LLM inference. Attached Figure Description
[0012] Figure 1 This is a diagram showing the correspondence between the original data of this invention and the AB=2 encoding.
[0013] Figure 2 This is a diagram illustrating the definition of high / low active bits, most / least significant bits, and bits using a raw data example.
[0014] Figure 3(a) shows the truth table and logic circuit diagram of a one-bit subtractor based on the first different input conditions.
[0015] Figure 3(b) shows the truth table and logic circuit diagram of a one-bit subtractor based on the second different input condition.
[0016] Figure 3(c) shows the truth table and logic circuit diagram of a one-bit subtractor based on a third different input condition.
[0017] Figure 3(d) shows the truth table and logic circuit diagram of a one-bit subtractor based on the fourth different input condition.
[0018] Figure 4(a) is a schematic diagram showing the relationship between the borrow output of the most significant bit of each subtractor and (1 << highbit).
[0019] Figure 4(b) is a schematic diagram of the logic circuit design for the high-activity bit screen code.
[0020] Figure 5(a) shows the overall architecture of the first type of AB decoding circuit. Figure 5(b) shows the overall architecture of the second type of AB decoding circuit.
[0021] Figure 6(a) is a schematic diagram of the first logic circuit in the AB decoding circuit.
[0022] Figure 6(b) is a schematic diagram of the second type of logic circuit in the AB decoding circuit.
[0023] Figure 6(c) is a schematic diagram of the third logic circuit in the AB decoding circuit.
[0024] Figure 6(d) is a schematic diagram of the fourth logic circuit in the AB decoding circuit.
[0025] Figure 6(e) is a schematic diagram of the fifth logic circuit in the AB decoding circuit.
[0026] Figure 6(f) is a schematic diagram of the sixth logic circuit in the AB decoding circuit.
[0027] Explanation of the reference numerals in the attached diagram: 301, Full Subtractor 1; 302, Full Subtractor 2; 303, Full Subtractor 3; 304, Full Subtractor 4; They all contain the input of the previous borrow requirement and the output of the next borrow requirement. Detailed Implementation
[0028] The following detailed description, in conjunction with the accompanying drawings and specific embodiments, provides a fast decoding method for Active-Bit encoding proposed in this invention. The advantages and features of this invention will become clearer from the following description. It should be noted that the drawings are all in a very simplified form and use non-precise proportions, and are only used to facilitate and clarify the illustration of the embodiments of this invention.
[0029] Because many bytes in the key-value cache matrix contain only a few bits that are "1" and the rest are "0", when a bit in a byte is "1", this bit is referred to as the active bit in this invention. AB encoding treats data as a byte stream and compresses each byte; when each byte of the encoded object has only n active bits, this encoding is simply called AB=n. In large language model (LLM) inference, the data bytes stored in the key-value cache often have only 2 active bits, which is suitable for encoding in the AB=2 manner. This invention provides a circuit and method for fast parallel decoding within a single pipeline clock cycle for AB=2.
[0030] The AB=2 encoding method involves rearranging and numbering the values of the two active bits that may appear in an 8-bit byte to obtain the encoding. One, such as Figure 1 As shown, an 8-bit byte can be represented by 5 bits, increasing the information entropy value of the stored data.
[0031] This encoding method has a characteristic: from the perspective of the original data, the position of the largest value of the most active bit located in the nth position within a byte after encoding rearrangement can be calculated using the following formula; where n ∈ {0, 1, ..., 7}, n=0 refers to the least significant bit (LSB), and n=7 refers to the most significant bit (MSB), such as... Figure 2 As shown.
[0032]
[0033] For example, the original data 001100002 is encoded as 011102. This is because the most active bit of the original data is located in the sixth bit position. Since the digits are counted starting from 0, n = 6 - 1 = 5. The digits of the encoded data are exactly as follows:
[0034] When decoding along this line of thought, the encoding can be categorized into different groups, such as... Figure 1The black and white alternating groups. The starting position of these groups can be obtained by adding 1 to the above formula (1), forming a fixed list B, which is [21, 15, 10, 6, 3, 1, 0]. The decoding algorithm obtains a new list B' by subtracting each value in list B from the encoded value M, and then finds the first positive difference from left to right in list B'. This positive number (difference) is represented by j. The offset number relative to the leftmost starting position of list B' is obtained according to the position of j in list B', which is called k. Then the 7-kth bit of the original data must be a high-activity bit, and the jth bit of the original value must be a low-activity bit.
[0035] Using the previous example, the code M is 011102. Subtracting list B, we get [-7, -1, 4, 8, 11, 13, 14]. The first positive number from left to right is j=4. Its offset from the starting position of list B' is 2, so the corresponding bit of the original data = (7-2) = 5 must be a high-activity bit. The difference j = 4 means that the 4th bit of the original data is a low-activity bit. Therefore, the original data is 001100002. Let's take two more extreme examples, M=0 and M=27. According to the algorithm, 0 - [21, 15, 10, 6, 3, 1, 0] = [-21, -15, -10, -6, -3, -1, 0], so k = 6, j = 0, therefore 7 - k = 1, and the original data is 000000112. And 27 - [21, 15, 10, 6, 3, 1, 0] = [6, 12, 17, 21, 24, 26, 27], so k = 0, j = 6, therefore 7 - k = 7, and the original data is 110000002.
[0036] The following demonstrates an AB decoding algorithm written in Python: Suppose M is the AB code to be decoded, which is the input value of the decoding algorithm and consists of 5 bits.
[0037] B = [21, 15, 10, 6, 3, 1, 0] def decode(m): i = 0 for b in B: j = mb if j >= 0: break i += 1 return (i, j) k, j = decode(M) highbit = 7-k lowbit = j value = (1 << highbit) | (1 << lowbit) However, implementing the decoding circuit according to this algorithm would be very slow because M compares each value in list B sequentially. A more practical approach is to compare M with the values in list B simultaneously in parallel, and then select the correct output based on the comparison results. Because the compared value is a known constant, this invention allows for optimization of the subtractor to reduce the number of logic gates: for example... Figures 3(a) to 3(d) As shown, the borrow is called b, the subtrahend is called x, and the minuend is called y. The subscripts i and o represent the input and output signals, respectively. When y = 0, the full subtractor 301 can be simplified as shown in Figure 3(a). The full subtractor 301 is a full subtractor for xy|y = 0, which saves three logic gates compared to the full subtractor, representing a 60% saving. Similarly, when y = 1, the full subtractor 302 can be simplified as shown in Figure 3(b). The full subtractor 302 is a full subtractor for xy|y = 1. Subtractors: 303, a full subtractor with xy|y=0 and b=0; 304, a full subtractor with xy|y=1 and b=0; The calculation of the least significant bit (LSB) does not require a borrow, so it can be further optimized, as shown in Figure 3(c) as full subtractor 303 and Figure 3(d) as full subtractor 304. Full subtractor 303 is a full subtractor with xy|y=0 and b=0, while full subtractor 304 is a full subtractor with xy|y=1 and b=0.
[0038] Furthermore, when encoding the values in subtraction list B, it's not necessary to use 5 bits for all of them. Except for 21, the other values of 15, 10, 6, 3, and 1 only require 4, 4, 3, 2, and 1-bit subtractors respectively. Although this method slightly reduces the number of logic gates, the circuit design used for determining and selecting high / low active bits becomes more complex and increases the fan-in number of logic gates, ultimately making it potentially less cost-effective. Therefore, this invention only describes a simple and fast logic circuit scheme.
[0039] The conditional check for j ≥ 0 in the algorithm can actually be achieved by using the most significant bit b of the subtraction result of each number in list B. o Output to complete; put b K o Defined as b corresponding to the most significant bit of the k-th bit of the original data o Output when M is less than B k time b K o If it is 0, then b is 0; otherwise, b is 0. K o =1; where B kThis is the value of the (k+1)th element from left to right in fixed list B. Figure 4(a) illustrates each b. K o The resulting value and the relationship between 1 << the high bit, if M is not less than 21, then b 7 o And the b afterwards K o All are 0. If M is not less than 15 but less than 21, b 7 o For 1 and then b K o All are 0, and so on, which leads to the logic circuit design of the high-activity bit screen code in Figure 4(b), and thus [e 7 o e 6 o e 5 o e 4 o e 3 o e 2 o e 1 o e 0 o ] is the result of 1 << the high bit, this e K o The resulting 8-bit number is named E, answering the first question of decoding. The difference corresponding to the Kth bit of the original data in the subtractor, 5 bits in total, is called d. K o The e in the high-activity bit screen code K o This can be viewed as the control signal for a multiplexer (MUX1) to select from different subtractors' d K o The output value is selected according to the corresponding j value, and then j is used as the control signal of the next multiplexer (MUX2) to obtain the 1 << low bit. Figures 5(a) to 5(b) and 6(a) to 6(f) show the overall structure diagram and the optimized subtractor for each bit, respectively.
[0040] The above description is merely a description of preferred embodiments of the present invention and is not intended to limit the scope of the present invention in any way. Any changes or modifications made by those skilled in the art based on the above disclosure shall fall within the protection scope of the claims.
Claims
1. A fast decoding method for Active-Bit encoding, characterized in that, include: The encoding is compared with a finite number of constant arrays, and the first part of the high-activity bit E is decoded by using a full subtractor and the highest borrow output of each subtractor. Based on the first part of the high-activity bit E variable, select an appropriate difference with the first multiplexer, and then use this difference to select the corresponding low-activity bit from the second multiplexer. Perform bitwise OR logic to complete the decoding to obtain the original data; among them, the smallest active bit 1 is the low-activity bit, and the largest active bit 1 is the high-activity bit.
2. The fast decoding method for Active-Bit encoding as described in claim 1, characterized in that, The low-activity bits of the original data were obtained through the following steps: The codes are classified into different groups, and a fixed list B is formed by the starting position of the dividing group; A new list B' is obtained by subtracting each value in the fixed list B from the encoded value M; Find the first positive difference from left to right in list B', denoted by j. The j-th bit in the original data is a low-activity bit.
3. The fast decoding method for Active-Bit encoding as described in claim 2, characterized in that, The highly active bits of the original data are obtained through the following steps: The offset number relative to the leftmost starting position of list B' is obtained based on the location of j in list B', and is called k; The 7th to kth bit of the original data is a highly active bit.
4. The fast decoding method for Active-Bit encoding as described in claim 2, characterized in that, The process of subtracting each value in the fixed list B from the encoded value M includes: comparing the encoded value M with the values in the fixed list B in parallel, and then selecting the correct output based on the comparison results.
5. The fast decoding method for Active-Bit encoding as described in claim 2, characterized in that, The condition for finding the first positive number j from left to right in list B' is determined by the most significant bit b of the subtraction result of each number in list B. o Output to complete; put b K o Defined as b corresponding to the most significant bit of the k-th bit of the original data o Output when the encoded value M is less than B k time b K o If it is 0, then b is 0; otherwise, b is 0. K o =1; where B k This is the (k+1)th value from left to right in the fixed list B.
6. The fast decoding method for Active-Bit encoding as described in claim 1, characterized in that, Except for the borrow data between subtractors, the rest of the process is parallel, so it can be completed in one clock cycle.
Citation Information
Patent Citations
Composition for enhancing anticancer effect of ERK inhibitors comprising aripiprazole as an active ingredient
KR1020230157876A
Apparatus and method for floating point normalization prediction
US4922446A
Priority encoder and floating-point normalization system for IEEE 754 standard
US5187678A
Zero run-length value encoding and decoding methods and video encoding and decoding methods, apparatuses and systems
WO2023168712A1
Coding method, decoding method, code stream, coder, decoder and storage medium
WO2025138048A1