Multiplier coding algorithm, application and storage medium
Generate intermediate encoding and two-bit binary number mapping to low-bit coefficient compression encoding by low-bit coefficient polynomials, which solves the problem of data width increase caused by MBE encoding method, realizes lower interconnection line width and power consumption, and expands the application of multiplier encoding algorithm.
Patent Information
- Application Number
- CN202510392196.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
In the existing multiplier encoding algorithm, the MBE encoding method causes the encoded data width to increase, occupy more chip area and increase power consumption, affecting the area optimization and power consumption control of chip design.
The intermediate encoding is generated using a low-bit width polynomial, and the low bit coefficient compression encoding is mapped by a two-bit binary number to reduce the encoded data width and reduce the interconnection line width and power consumption.
It significantly reduces the encoded data width, reduces the width and power consumption of interconnected lines, and expands the application scenarios and ranges.
Smart Images

Figure CN120255846A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing algorithms, and particularly to a multiplier coding algorithm, application, and storage medium. Background Art
[0002] As the importance of matrix multiplication in artificial intelligence has become increasingly prominent, and due to the performance bottleneck of traditional general-purpose processors in processing large-scale tensor calculations, many research institutions and companies have begun to develop dedicated hardware to accelerate matrix multiplication operations. These dedicated hardware devices have significantly improved the performance and energy efficiency of tensor calculations by optimizing data flow, reducing data transmission overhead, and improving the utilization rate of computing resources.
[0003] In the key operation unit of matrix multiplication - the multiplier, most of its encoders use the MBE (Modified Booth Encoding) coding method for multiplication coding algorithms. For an n-bit binary multiplicand, the MBE coding method will obtain bits of encoded values. This coding method significantly increases the data width after encoding. The increase in the data width after encoding will cause the interconnection lines for transmitting these data to become wider. In chip design, the increase in the width of interconnection lines will occupy more chip area and increase the power consumption of signal transmission, which is not conducive to chip area optimization and power control. Based on the above problems, there is an urgent need to redesign the coding algorithm. Summary of the Invention
[0004] In view of the deficiencies of the prior art, the present invention provides a multiplier coding algorithm, application, and storage medium, which can significantly reduce the data width after encoding, thereby reducing the width and power consumption of interconnection lines.
[0005] The present invention solves the technical problems by adopting the following technical solutions:
[0006] The present invention provides a multiplier coding algorithm, including:
[0007] Obtaining the binary original code of the multiplicand A;
[0008] Generating intermediate coding based on the binary original code using a low-bit-width polynomial;
[0009] Based on the intermediate coding, mapping two-bit binary numbers to low-bit coefficient compressed coding, and the number of bits of the compressed coding is the number of bits of the binary original code plus 1.
[0010] Preferably, the binary original code includes a sign bit sign and a numerical bit a i , where sign identifies the positive or negative of the multiplicand A, and a i is the i-th bit of the binary original code representation of the multiplicand A.
[0011] Preferably, the low-bit-width polynomial satisfies:
[0012]
[0013] where |A| is the unsigned value of the multiplicand A, m is the number of bits after padding the binary encoding of the unsigned value of the multiplicand A, and w i is the value of the i-th bit of the intermediate encoding, and the w i is configured to generate four different values through a recursive expression and a carry-chain encoding.
[0014] Preferably, the m is the number of bits after padding the number of bits m1 of the value of the multiplicand A, specifically including:
[0015] If m1 is even, then m = m1, and the binary original code of the multiplicand A is the binary encoding of the unsigned value of the multiplicand A;
[0016] If m1 is odd, then m = m1 + 1, and the binary original code of the unsigned value of the multiplicand A is obtained by padding 0 in front of the highest bit of the binary encoding of the unsigned value of the multiplicand A.
[0017] Preferably, the w i is configured to generate through a recursive expression and a carry-chain encoding, specifically including:
[0018] Generate a carry symbol C i using a carry-chain encoding, and generate w i using a recursive expression. The logical calculation expression is:
[0019] c i+1 = (a 2i+1 & a 2i ) | (a 2i+1 & C i )
[0020] where c0 = 0, and a 2i+1 , a 2i , a 2i+1 are all the values of the multiplicand A, and i ≥ 0;
[0021] w i = [a 2i+1 a 2i 10 + C i
[0022] [a 2i+1 a 2i 10 represents the decimal number of the 2-bit binary a 2i+1 a 2i , and w i ∈ {3, 0, 1, 2}.
[0023] Preferably, the intermediate encoding is mapped to a low-bit coefficient compression encoding using two-bit binary numbers, specifically as follows:
[0024] The values {0, 1, 2, 3} of the intermediate encoding are represented using the 2-bit binary numbers {00, 01, 10, 11} respectively;
[0025] Then, the 2-bit binary values {00, 01, 10, 11} after representation are mapped to the values {0, 1, 2, -1} of each bit K of the compression encoding to generate the compression encoding. i
[0026] Preferably, the bit weight of the compression encoding bit K i is 2 2i .
[0027] Preferably, the algorithm further includes:
[0028] Identifying the sign value sign of the multiplicand A:
[0029] If the sign value represents a negative number, then invert the sign bit of the multiplier B, otherwise keep the sign bit of the multiplier B unchanged;
[0030] Associating the bit weight 2 i with the value of the compression encoding bit K 2i and then multiplying each by the multiplier to obtain partial products;
[0031] Accumulating the partial products to obtain the final multiplication operation result;
[0032] Among them, a shift circuit is used to implement the operation of the partial products, and a register circuit and a full adder circuit are used to implement the accumulation of the partial products.
[0033] The present invention also provides an electronic device, including:
[0034] At least one processor; and
[0035] A memory communicatively connected to the at least one processor; wherein,
[0036] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the foregoing multiplier encoding algorithm.
[0037] The present invention also provides a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause the computer to execute the foregoing multiplier encoding algorithm.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] The multiplier coding algorithm of the present invention decomposes the multiplicand into a polynomial form with a low bit width, avoiding the need for control signals with a high bit width in traditional MBE, reducing the transmission complexity of data within the chip, and significantly reducing the data width after coding. In addition, through recursive expressions and carry chain coding, this algorithm realizes the compressed expression of low-bit coefficients, effectively expanding the application scenarios and scope.
[0040] Regarding the present invention compared with the prior art, other outstanding substantive features and remarkable progress are further described in detail in the embodiment section. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] By reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:
[0042] Figure 1 It is a schematic diagram of the connection structure and signal processing of the encoder and selector in the prior art;
[0043] Figure 2 It is a schematic diagram of the MBE coding algorithm logic in the prior art;
[0044] Figure 3 It is a schematic diagram of the hardware structure mapped by the multiplier coding algorithm in Embodiment 1;
[0045] Figure 4 It is a schematic diagram of the coding algorithm when the multiplicand A is 91 in Embodiment 1;
[0046] Figure 5 It is a schematic diagram of the coding algorithm when the multiplicand A is 124 in Embodiment 1;
[0047] Figure 6 It is a comparison schematic diagram of the coding algorithm of the present invention and the prior art algorithm when the multiplicand A is 34 in Embodiment 1;
[0048] Figure 7 It is a comparison schematic diagram of the coding algorithm of the present invention and the prior art algorithm when the multiplicand A is 50 in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0050] It should be noted that certain names are used in the description and claims to refer to specific components. It should be understood that those of ordinary skill in the art may use different names to refer to the same component. The description and claims of this application do not use the difference in names as a way to distinguish components, but use the substantial difference in function of components as the criterion for distinguishing components. As used in the description and claims of this application, "comprising" or "including" is an open-ended term and should be interpreted as "comprising but not limited to" or "including but not limited to". The embodiments described in the specific implementation section are the preferred embodiments of the present invention and are not intended to limit the scope of the present invention.
[0051] In addition, those skilled in the art know that various aspects of the present invention can be implemented as a system, method, or computer program product. Therefore, various aspects of the present invention can be specifically implemented in the form of a combination of software and hardware, which can be collectively referred to as "circuit", "module", or "system" here. In addition, in some embodiments, various aspects of the present invention can also be implemented in the form of a computer program product in one or more microcontroller-readable media, which contain program codes readable by the microcontroller.
[0052] Before introducing the technical solution of the present invention, in order to compare the prominent substantial features and remarkable progress of the present invention relative to the prior art, the prior art is further explained and described.
[0053] Most encoders in modern multipliers use the MBE coding method. For an n-bit operand A(a n-1 a n-2 …a0), where a i is the i-th bit of the two's complement representation of A. The MBE encoder represents A in the following form:
[0054]
[0055] where a -1 =0, m i ∈{-2, -1, 0, 1, 2}. Therefore, when A×B is executed inside the multiplier, as Figure 1 shown, the selector after the encoder will select -2B, -B, 0, B, and 2B to the subsequent adder unit according to the signals NEG, CE, and SE encoded for A, where the logical expressions of NEG, CE, and SE are as follows:
[0056]
[0057] For example, when A = 34, as Figure 2 shown, after the MBE coding logic, four different bit weights {2 6, 2 4 , 2 2 , 2 0}, multiply the coding coefficients of different weights by B and then sum them to get the multiplication result of A×B. This process only requires hardware circuits of shift and addition to realize multiplication calculation. For an n-bit binary complement number, the MBE coding method will get the bit coding value ( Figure 2 The MBE encoding bit width of INT8 is 12 bits), so a higher encoding bit width will increase the number of wire nets inside the PE and have a negative impact on area, power consumption, etc.
[0058] The reason is that MBE coding essentially requires encoding every 2 bits of the multiplier into a 3-bit control signal (NEG, SE, CE). For the multiplication of two n-bit numbers, the multiplicand needs to be encoded as This encoding method significantly increases the width of the encoded data. The increase in the width of the encoded data will cause the interconnection lines that transmit this data to become wider. In chip design, the increase in the width of the interconnection lines will occupy more chip area and increase the power consumption of signal transmission. This is not conducive to chip area optimization and power control.
[0059] Therefore, the present invention proposes a new encoding algorithm based on the design of the existing encoder: for the multiplication of two n-digit numbers, the multiplicand needs to be encoded as n+1 bits, compared to The new design can significantly reduce the width of the encoded data, thereby reducing the width and power consumption of the interconnection line. The specific content of the present invention is introduced in detail in the following by way of specific embodiments.
[0060] Example 1
[0061] A multiplier encoding algorithm in this embodiment includes:
[0062] Get the binary code of the multiplicand A, which includes the sign bit sign and the value bit a i , where sign indicates the positive or negative of the multiplicand A, a i It is the i-th bit of the binary code representation of the multiplicand A.
[0063] like Figure 3 As shown, the intermediate code is generated based on the binary original code using a low bit width polynomial; the low bit width polynomial in this embodiment satisfies:
[0064]
[0065] Where |A| is the unsigned value of the multiplicand A, m is the number of bits after the binary code of the unsigned value of the multiplicand A is matched with m1, and w iis the value of the i-th bit of the intermediate encoding, and the w i is configured to generate four different values through a recursive expression and a carry chain encoding.
[0066] In this embodiment, m is the number of bits after the spouse of the number of bits m1 of the multiplicand A, specifically including:
[0067] If m1 is an even number, then m = m1, and the binary original code of the multiplicand A is the binary encoding of the unsigned value of the multiplicand A;
[0068] If m1 is an odd number, then m = m1 + 1, and the binary original code of the multiplicand A is obtained by adding 0 in front of the highest bit of the binary encoding of the unsigned value of the multiplicand A.
[0069] The w in this embodiment i is configured to generate specifically including:
[0070] Generate a carry symbol C using a carry chain encoding i , and generate w using a recursive expression i , and the logical calculation expression is:
[0071] c i+1 =(a 2i+1 &a 2i )|(a 2i+1 &C i )
[0072] where c0 = 0, and a 2i+1 , a 2i , a 2i+1 are all the values of the multiplicand A, i≥0;
[0073] w i =[a 2i+1 a 2i 10 +C i
[0074] [a 2i+1 a 2i 10 represents the decimal number of the 2-bit binary a 2i+1 a 2i , and w i ∈{3,0,1,2}.
[0075] Based on the intermediate encoding, a two-bit binary number is mapped to a low-bit coefficient compression encoding, and the number of bits of the compression encoding is the number of bits of the binary original code plus 1, specifically:
[0076] Use the 2-bit binary {00, 01, 10, 11} to represent the values {0, 1, 2, 3} of the intermediate encoding respectively;
[0077] Then map the represented 2-bit binary values {00, 01, 10, 11} to the values {0, 1, 2, -1} of each bit K of the compressed code to generate the compressed code; i
[0078] In this embodiment, the bit weight of the compressed code bit K i is 2 2i .
[0079] The multiplier coding algorithm in this embodiment further includes:
[0080] Identify the sign bit sign value of the multiplicand A:
[0081] If the sign value represents a negative number, then invert the sign bit of the multiplier B, otherwise keep the sign bit of the multiplier B unchanged;
[0082] Multiply the values of the compressed code bit K i associated with the bit weight 2 2i by the multiplier respectively to obtain partial products;
[0083] Accumulate the partial products to obtain the final multiplication operation result;
[0084] Among them, a shift circuit is used to implement the operation of the partial product, and a register circuit and a full adder circuit are used to implement the accumulation of the partial products.
[0085] To further clarify and verify the principle and performance of the multiplier coding algorithm of the present invention, the following is an example:
[0086] Such as Figure 4 shown:
[0087] When A = 91 (0, 1011011), the sign flag is 0, indicating that the value is positive;
[0088] After padding 0 at the high position, its binary original code is (01011011);
[0089] Based on the binary original code (01011011), according to the calculation expression of w i the intermediate coding value of A is represented as {0, 1, 2, 3, 3};
[0090] Use 2-bit binary {00, 01, 10, 11} to represent the intermediate coding value {0, 1, 2, 3, 3} respectively, and get {00, 01, 10, 11, 11};
[0091] Based on the represented {00, 01, 10, 11, 11}, map it to the compressed code {1, 2, -1, -1};
[0092] The corresponding bit weights are {26 , 2 4 , 2 2 , 2 0};
[0093] Associate the weighted value 2 of the numerical value of the compressed encoded bit K i and then multiply by the multiplier respectively to obtain the partial product, that is, A×B = 64B + 32B - 4B - B = 91B. 2i Again, as shown in
[0094] For another example Figure 5 as shown:
[0095] When A = 124(0,1111100), according to the calculation expression of w i the encoded value of A is represented as {0, 2, 0, -1, 0}, so A×B = 128B - 4B = 124B.
[0096] In order to further compare the differences between the present invention and the prior art, continue to compare as follows
[0097] Such as Figure 6 shown:
[0098] When A = 34(00100010), after the MBE encoding logic, as shown in Figure 6 (left), the encoded value of A is {1, -2, 1, -2}. Since the coefficient m of MBE i ∈{-2, -1, 0, 1, 2}, so 3 bits are required to represent each encoded bit. Therefore, a total of 12 bits of encoded bits are required to represent the encoded value of A for the INT8 operand. Finally, the calculation formula of MBE is A×B = 64B - 32B + 4B - 2B = 34B. The encoding method proposed in this embodiment is as shown in Figure 6 (right). According to the calculation expression of w i the encoded value of A is {0, 0, 2, 0, 2}, and the calculation formula A×B = 32B + 2B = 34B, and only 9 bits of encoded bits are required to represent the encoded value of A.
[0099] When A = 50(00110010), after the MBE encoding logic, as shown in Figure 7 (left), the encoded value of A is {1, -1, 1, -2}. Finally, the calculation formula of MBE is A×B = 64B - 16B + 4B - 2B = 50B. The encoding method proposed in this embodiment, as shown in Figure 7 (right). According to the calculation expression of w i the encoded value of A is {0, 1, -1, 0, 2}, and the calculation formula A×B = 64B - 16B + 2B = 50B. Since w i∈{-1, 0, 1, 2}, so only 2 bits can be used to represent one of the encoded bits compared to MBE.
[0100] Therefore, it can be undoubtedly concluded that for the multiplication of two n-digit numbers, the multiplicand in this embodiment needs to be encoded into n + 1 bits. Compared with the MBE method with bits, the new design can significantly reduce the width of the encoded data, thereby reducing the width and power consumption of the interconnection lines.
[0101] Embodiment 2
[0102] This embodiment provides an electronic device, including:
[0103] at least one processor; and
[0104] a memory communicatively connected to the at least one processor; wherein,
[0105] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the multiplier encoding algorithm described in Embodiment 1.
[0106] Embodiment 3
[0107] This embodiment provides a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause the computer to execute the multiplier encoding algorithm as described in Embodiment 1.
[0108] The various embodiments of the systems and techniques described in the above embodiments can be implemented in digital electronic circuit systems, integrated circuit systems, dedicated ASICs (Application Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs, and the one or more computer programs can be executed and / or interpreted on a programmable system including at least one programmable processor. The programmable processor can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0109] These computing procedures (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0110] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0111] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0112] A computer system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0113] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
[0114] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A multiplier coding algorithm, characterized in that, including: obtaining the binary original code of the multiplicand A; generating an intermediate code based on the binary original code by using a low-bit-width polynomial; mapping the intermediate code to a low-bit coefficient compressed code by using two-bit binary numbers, and the number of bits of the compressed code is the number of bits of the binary original code plus 1.
2. The multiplier coding algorithm according to claim 1, characterized in that, The binary original code includes a sign bit sign and a numerical bit a i , where sign indicates the positive or negative of the multiplicand A, and a i is the i-th bit of the binary original code representation of the multiplicand A.
3. The multiplier encoding algorithm according to claim 2, wherein The low-bit-width polynomial satisfies: Where |A| is the unsigned value of the multiplicand A, m is the number of bits after mating the binary encoding bits m1 of the unsigned value of the multiplicand A, and w i is the value of the i-th bit of the intermediate code, and the w i is configured with four different values generated by a recursive expression and a carry chain encoding.
4. The multiplier coding algorithm according to claim 3, wherein The m is the number of bits after the number of bits m1 of the multiplicand A is paired, specifically including: If m1 is an even number, then m = m1, and the binary original code of the multiplicand A is the binary code of the unsigned value of the multiplicand A; If m1 is an odd number, then m = m1 + 1, and the binary original code of the multiplicand A is obtained by adding 0 in front of the highest bit of the binary code of the unsigned value of the multiplicand A.
5. The multiplier coding algorithm according to claim 4, wherein The said w i configured to generate through recursive expressions and carry chain encoding specifically including: Generate carry symbol C using carry chain encoding i , generate w using recursive expression i , the logical calculation expression is: c i+1 = (a 2i+1 & a 2i ) | (a 2i+1 & C i ) where c0 = 0, a 2i+1 , a 2i , a 2i+1 are all the values of the multiplicand A, and i ≥ 0; w i = [a 2i+1 a 2i 10 + C i [a 2i+1 a 2i 10 Represents the 2-bit binary number of a 2i+1 a 2i in decimal, w i ∈ {3, 0, 1, 2}. 6. A multiplier encoding algorithm according to claim 5, wherein The mapping of the intermediate code to a low-bit coefficient compressed code by using two-bit binary numbers is specifically: using 2-bit binary numbers {00, 01, 10, 11} to represent the values {0, 1, 2, 3} of the intermediate code respectively; Then, map the represented 2-bit binary values {00, 01, 10, 11} to the values {0, 1, 2, -1} of each bit K of the compression code to generate the compression code. i 7. A multiplier coding algorithm according to claim 6, characterized in that The compressed coding bit K i has a bit weight of 2 2i .
8. The multiplier encoding algorithm according to claim 7, wherein The algorithm further includes: identifying the sign bit sign value of the multiplicand A: if the sign value represents a negative number, then taking the inverse of the sign bit of the multiplier B, otherwise keeping the sign bit of the multiplier B unchanged; Associate the compression-encoded bit K i with the numerical associated bit weight 2 2i and then multiply each by a multiplier to obtain partial products; accumulating the partial products to obtain the final multiplication operation result; wherein the operation of the partial product is implemented by using a shift circuit, and the accumulation of the partial products is implemented by using a register circuit and a full adder circuit.
9. An electronic device, characterized in that, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the multiplier encoding algorithm according to any one of claims 1-8.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the multiplier encoding algorithm according to any one of claims 1-8.