Error self-compensated approximate adder tree for in-memory computing and construction method thereof

CN122672749APending Publication Date: 2026-09-01SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610859142.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-15
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提供一种面向数字存内计算的误差自补偿近似加法树及其构建方法,用于解决现有数字存内计算部分积归约路径中加法树功耗和面积开销较高、传统近似加法方案容易产生方向性偏置误差并影响计算结果数值稳定性的技术问题

Benefits of technology

[0023](1)本发明提供了一种面向数字存内计算的误差自补偿近似加法树,应用于数字存内计算阵列的部分积累加路径中。该近似加法树将低位权权重切片对应的一组或多组通道级归约路径设置为通道级近似加法树,将高位权权重切片对应的通道级归约路径设置为通道级精确加法树,并使权重级加法树和特征值级加法树保持精确归约结构。由于低位权部分对最终乘累加结果的数值贡献相对较小,在该部分引入选择性近似处理能够在控制误差影响的同时,减少部分积归约路径中的电路开销和切换活动。进一步地,通道级近似加法树可根据目标精度、目标功耗或目标硬件开销,在一个或多个预设计算层级中设置近似4-2压缩模块,从而实现计算精度与硬件开销之间的折中。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122672749A_ABST
    Figure CN122672749A_ABST
Patent Text Reader

Abstract

This invention discloses an error-compensating approximate addition tree for in-memory digital computation and its construction method. The approximate addition tree is applied to the partial accumulation path of a digital in-memory computation array. The partial accumulation path includes a channel-level addition tree, a weight-level addition tree, and an eigenvalue-level addition tree. The channel-level approximate addition tree adopts a hybrid reduction structure combining approximate compression and exact reduction, and sets an approximate 4-2 compression module in one or more preset computation levels. The approximate 4-2 compression module includes a bit-wise approximate 4-2 compressor. The approximate 4-2 compressor generates the current bit-weight result and the next higher bit-weight result according to a preset output mapping relationship, so that it produces positive error, negative error, or zero error under different input modes. This invention can reduce the circuit complexity and switching activities of the low-weight partial product reduction path, improve the energy efficiency of the digital in-memory computation circuit, and reduce the impact of directional bias error accumulation on the multiplication-accumulation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to digital in-memory computing, approximate computing circuits, and low-power digital integrated circuit design technology, and particularly to an error-compensating approximate addition tree for digital in-memory computing and its construction method, belonging to the technical field of calculation, deduction, or counting. Background Technology

[0002] With the development of artificial intelligence inference and edge intelligence applications, computing systems are increasingly demanding lower power consumption, higher energy efficiency, and real-time performance. Digital Computing-In-Memory (DCIM) can perform core operations such as multiplication and accumulation near the memory array, reducing the frequent data movement between storage and computing units, thus becoming an important technical route in high-efficiency AI computing architectures. However, after the DCIM array completes partial multiplication or bitwise multiplication, the partial product output by the array still needs to be reduced through multiple levels. Therefore, the addition tree becomes a crucial component affecting the power consumption, area, and computational latency of DCIM circuits.

[0003] In existing digital in-memory computing architectures, multiplication and accumulation operations are typically performed using a bit-serial approach. Input feature values ​​and weight data participate in the computation according to bit slices or bit sequences. The in-memory computing array outputs the corresponding partial product results, which are then reduced through a multi-level addition tree to obtain the final computation result. Depending on the reduction object and the reduction level, the multi-level addition tree typically includes a channel-level addition tree, a weight-level addition tree, and a feature-level addition tree. Specifically, the channel-level addition tree is used to reduce the partial product along the input channel direction; the weight-level addition tree is used to reduce the intermediate results corresponding to different weight slices by shifting them according to their bit weights; and the feature-level addition tree is used to further reduce the intermediate results corresponding to different input feature value slices. Because the channel-level addition tree is located at the front end of the partial product reduction chain and involves a large number of input channels, its preceding reduction levels typically have a large input scale and high switching activity, making it an important target for low-power optimization.

[0004] Existing digital in-memory computing systems typically employ precise addition units to implement the aforementioned multi-level addition tree to ensure the accuracy of the multiplication-accumulation result. While this fully precise reduction method maintains high computational accuracy, it requires numerous precise addition or compression units when dealing with a large number of input channels, large partial product sizes, and deep reduction levels. This leads to high hardware complexity, switching activity, and power consumption, hindering the implementation of energy-efficient digital in-memory computing systems. To reduce the power consumption and circuit complexity of the addition tree, some existing technologies attempt to introduce approximate computation concepts into the addition circuit design, sacrificing some computational accuracy for higher energy efficiency. However, ordinary approximate adders or approximate compressors are prone to unidirectional bias errors under certain input modes. These directional bias errors may continue to propagate during multi-level partial product reduction and accumulate in large-scale matrix multiplication-accumulation operations, thus affecting the numerical stability of the final calculation result and making it difficult to simultaneously achieve power reduction and computational reliability.

[0005] Therefore, to address the issues of high power consumption and area overhead in existing digital in-memory computing, and the tendency of traditional approximate addition schemes to generate directional bias errors that affect numerical stability, it is necessary to propose a new approximate addition tree structure and its construction method. On the one hand, selective approximation processing is introduced in the low-weight reduction path where the error impact is relatively small. On the other hand, based on the target accuracy, target power consumption, and target hardware overhead, approximate compression modules are set in one or more preset computation levels of the channel-level addition tree. Simultaneously, approximate compression units capable of generating positive, negative, or zero errors are constructed, so that errors under different input modes exhibit a statistically significant mutual compensation trend during the overall reduction process. This reduces hardware overhead and switching activities while minimizing the impact of directional error accumulation on the multiplication-accumulation result. Summary of the Invention

[0006] The purpose of this invention is to provide an error-compensating approximate addition tree and its construction method for in-memory digital computing, which addresses the technical problems of high power consumption and area overhead of addition trees in existing in-memory digital computing partial product reduction paths, and the tendency of traditional approximate addition schemes to generate directional bias errors that affect the numerical stability of calculation results. The approximate addition tree is applied to the partial accumulation path of an in-memory digital computing array. By introducing selective approximation processing into one or more channel-level reduction paths corresponding to low-weight slices, and maintaining the accurate reduction structure of the channel-level reduction paths corresponding to high-weight slices, as well as the weight-level and eigenvalue-level addition trees, the hardware overhead and switching activities of the addition tree are reduced, while minimizing the impact of directional bias error accumulation on the multiplication-accumulation results.

[0007] The present invention adopts the following technical solution:

[0008] An error self-compensation approximate adder tree for digital in-memory computing, applied to the partial product accumulation path in a digital in-memory computing array. Said partial product accumulation path includes a channel-level adder tree, a weight-level adder tree and an eigenvalue-level adder tree; wherein the channel-level adder tree is configured to reduce partial products in the direction of input channels, the weight-level adder tree is configured to reduce intermediate results corresponding to different weight slices after bit-weight shifting, and the eigenvalue-level adder tree is configured to further reduce intermediate results corresponding to different input eigenvalue slices.

[0009] Said channel-level adder tree includes a channel-level approximate adder tree and a channel-level exact adder tree. Wherein, one or more groups of channel-level reduction paths corresponding to low-weight weight slices are configured as the channel-level approximate adder tree, and the channel-level reduction path corresponding to high-weight weight slices is configured as the channel-level exact adder tree. Said weight-level adder tree and said eigenvalue-level adder tree maintain an exact reduction structure. Accordingly, the present invention introduces selective approximation processing into the low-weight reduction paths where the impact of error is relatively small, and keeps the high-weight reduction paths and subsequent reduction levels maintain exact reduction.

[0010] As a further optimization solution, when the bit width of weight data is W bit and the slice calculation is performed in accordance with B bit, said weight data forms Q weight slices, wherein , both W and B are positive integers, represents the ceiling operation. The channel-level reduction paths corresponding to R weight slices starting from the least significant bit are configured as the channel-level approximate adder tree, and the channel-level reduction paths corresponding to the remaining Q-R higher-weight weight slices are configured as the channel-level exact adder tree, wherein 1≤R<Q; when W cannot be divided exactly by B, the actual bit width of the highest weight slice is less than or equal to B bit.

[0011] As a further optimization solution, said channel-level approximate adder tree adopts a hybrid reduction structure combining approximate compression and exact reduction. Said channel-level approximate adder tree includes a plurality of calculation levels, wherein one or more preset calculation levels are provided with approximate 4-2 compression modules, and the calculation levels not provided with said approximate 4-2 compression modules adopt an exact reduction structure. Preferably, said preset calculation levels include the first calculation level of said channel-level approximate adder tree; in other embodiments, the approximate 4-2 compression modules can also be arranged in the second calculation level or a plurality of calculation levels according to target accuracy, target power consumption or target hardware overhead.

[0012] As a further optimization, the approximate 4-2 compression module groups four input channels together, performs bitwise approximate compression on the equal-weighted partial products within the same group of input channels, and outputs the corresponding intermediate compression results. For a partial product with a bit width of M, the approximate 4-2 compression module includes M bitwise-configured approximate 4-2 compressors, each corresponding to one of the M bits of the partial product. The two compressed bits output by the approximate 4-2 compressors are used as the current bit-weighted result and the next higher bit-weighted result, respectively, in subsequent reduction.

[0013] As a further optimization scheme, the approximate 4-2 compressor receives four equally weighted input bits IN0, IN1, IN2, and IN3, and outputs a low-order result OUT[0] and a high-order result OUT[1]. OUT[0] is fixed at logic 1. IN0 and IN1 form the first input group, and IN2 and IN3 form the second input group. OUT[1] is obtained by performing an OR operation between the all-1 detection results of the first input group and the all-1 detection results of the second input group, and OUT[1] = (IN0 AND IN1) OR (IN2 AND IN3). The approximate 4-2 compressor generates the current bit-weighted result and the higher-order bit-weighted result based on the above output mapping relationship, causing it to produce positive, negative, or zero errors under different input modes.

[0014] This invention also provides a method for constructing an error-compensating approximate addition tree for digital in-memory computation, comprising the following steps:

[0015] Step 1: Determine a partial accumulation path in the digital in-memory computing array, wherein the partial accumulation path includes a channel-level addition tree, a weight-level addition tree, and an eigenvalue-level addition tree;

[0016] Step 2: Based on the position weight of the weight slice, determine one or more sets of approximate processing objects from the channel-level reduction path corresponding to the low position weight slice, and determine the channel-level reduction path corresponding to the high position weight slice, as well as the weight-level addition tree and the eigenvalue-level addition tree as the exact reduction objects.

[0017] Step 3: Based on the target accuracy, target power consumption, or target hardware overhead, determine one or more approximate insertion levels from multiple computational levels of the channel-level addition tree corresponding to the approximate processing object, and determine the computational levels that are not determined as approximate insertion levels as exact reduction levels;

[0018] Step 4: Construct an approximate 4-2 compression module in one or more approximate insertion levels, so that the approximate 4-2 compression module performs bit-by-bit approximate compression on the equally weighted partial product bits from multiple input channels and generates intermediate compression results;

[0019] Step 5: Construct the approximate 4-2 compressor in the approximate 4-2 compression module, divide the four equally weighted input bits into two input groups, fix the low-order result to logic 1, and generate the high-order result based on whether any input group satisfies the all-1 condition;

[0020] Step 6: The current position weight result and the next higher position weight result output by the approximate 4-2 compressor are accumulated and reduced according to their position weights. This allows the positive and negative errors generated by multiple approximate 4-2 compressors to enter the overall reduction process along with the intermediate compression results, and they exhibit a statistically significant mutual compensation trend during the overall reduction process.

[0021] When performing partial product reduction based on the aforementioned error-compensated approximate adder tree, the digital in-memory computing array performs multiplication and accumulation operations in a bit-serial manner, dividing the input feature values ​​and weight data into slices according to a preset slice width to generate corresponding partial product data. The partial product data first enters the channel-level adder tree for reduction, where the partial products corresponding to the low-weight slices are mixed-reduced by the channel-level approximate adder tree, and the partial products corresponding to the high-weight slices are precisely reduced by the channel-level exact adder tree. Subsequently, the output results of the channel-level approximate adder tree and the output results of the channel-level exact adder tree are shifted and aligned according to the bit weights of the corresponding weight slices and then jointly sent to the weight-level adder tree for reduction. The output results of the weight-level adder tree are then sent to the feature-value-level adder tree for further reduction to obtain the multiplication and accumulation result.

[0022] Beneficial effects:

[0023] (1) This invention provides an error-compensating approximate addition tree for in-memory computing, applied to the partial accumulation path of an in-memory computing array. This approximate addition tree sets one or more channel-level reduction paths corresponding to low-weight slices as channel-level approximate addition trees, and sets the channel-level reduction paths corresponding to high-weight slices as channel-level exact addition trees, while maintaining the exact reduction structure of the weight-level addition tree and the eigenvalue-level addition tree. Since the low-weight part contributes relatively little to the final multiplication-accumulation result, introducing selective approximation processing in this part can reduce circuit overhead and switching activities in the partial product reduction path while controlling the impact of errors. Furthermore, the channel-level approximate addition tree can set an approximate 4-2 compression module in one or more preset computing levels according to the target accuracy, target power consumption, or target hardware overhead, thereby achieving a trade-off between computing accuracy and hardware overhead.

[0024] (2) This invention employs an approximate 4-2 compressor constructed based on a preset output mapping relationship. The approximate 4-2 compressor generates positive, negative, or zero errors under different input modes. After the output results of multiple approximate 4-2 compressors enter the subsequent reduction process, the positive and negative errors can participate in the overall reduction along with the intermediate compression results. In the overall reduction process involving multiple bits, multiple input channels, and multiple approximate 4-2 compression modules, they exhibit a statistically significant mutual compensation trend, thereby reducing the impact of directional bias error accumulation on the numerical stability of the multiplication-accumulation result. At the same time, the approximate 4-2 compressor generates compression results based on a simplified output mapping relationship, has a simple logical structure, and is suitable for repeated deployment in large-scale partial product reduction paths of digital in-memory computing arrays. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of the application structure of the error self-compensating approximate addition tree described in this invention in digital memory computation;

[0026] Figure 2 This is a schematic diagram of the construction method of the error self-compensation approximate addition tree described in this invention;

[0027] Figure 3 This is a schematic diagram of the hybrid reduction structure of the channel-level approximate addition tree described in this invention;

[0028] Figure 4 This is a schematic diagram of bitwise compression and output bit weight alignment of the approximate 4-2 compression module described in this invention;

[0029] Figure 5 This is a schematic diagram of the output mapping relationship and error distribution of the approximate 4-2 compressor described in this invention;

[0030] Figure 6 This is a schematic diagram of the gate-level circuit of the approximate 4-2 compressor described in this invention. Detailed Implementation

[0031] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings.

[0032] like Figure 1As shown, the present invention provides an error self-compensating approximate adder tree for digital in-memory computing, which is applied to a partial product accumulation path of a digital in-memory computing array. The input eigenvalues correspond to a plurality of input channels, which are denoted as eigenvalue channels 0 to N-1. The eigenvalues corresponding to each input channel are input into the in-memory computing array, participate in calculation with weight data in a slicing manner, and generate corresponding partial product data. According to different bit weights of the weight slices, the partial product data respectively enter a high bit weight reduction path and a low bit weight reduction path. The high bit weight reduction path includes an in-memory computing array and a channel-level precise adder tree, wherein the channel-level precise adder tree is configured to precisely reduce partial products of different input channels corresponding to high bit weight weight slices, and output a precise reduction result. The low bit weight reduction path includes an in-memory computing array and a channel-level approximate adder tree, wherein the channel-level approximate adder tree is configured to approximately reduce partial products of different input channels corresponding to low bit weight weight slices, and output an approximate reduction result. Wherein, Q represents the total number of weight slices, and R represents the number of low bit weight approximate slices, Figure 1 wherein the Q-R groups represent the number of weight slices corresponding to the high bit weight precise reduction path, and the R groups represent the number of weight slices corresponding to the low bit weight approximate reduction path.

[0033] The precise reduction result output by the channel-level precise adder tree and the approximate reduction result output by the channel-level approximate adder tree are jointly sent to a weight slice bit weight shift alignment unit. The weight slice bit weight shift alignment unit is configured to perform shift alignment on the channel-level reduction result according to the bit weight corresponding to each weight slice. The result after shift alignment is sent to a weight-level adder tree for reduction. The result output by the weight-level adder tree is then sent to an eigenvalue slice bit weight shift alignment unit, and shift alignment is performed according to the bit weight corresponding to the input eigenvalue slice. The result after shift alignment is sent to an eigenvalue-level adder tree for further reduction, and a multiply-accumulate result is finally obtained.

[0034] Specifically, when the bit width of the weight data is W bits and slice calculation is performed according to B bits, the weight data forms Q weight slices, wherein , represents a ceiling operation. R represents the number of low bit weight weight slices set as approximate processing objects starting from the least significant bit, and 1≤R<Q. Correspondingly, R groups of low bit weight approximate reduction paths and Q-R groups of high bit weight precise reduction paths are arranged in the partial product accumulation path. The channel-level adder tree in the low bit weight approximate reduction path is configured as a channel-level approximate adder tree, the channel-level adder tree in the high bit weight precise reduction path is configured as a channel-level precise adder tree, and the weight-level adder tree and the eigenvalue-level adder tree maintain a precise reduction structure.

[0035] As a concrete example, when both the input feature values ​​and weight data are 8 bits, and both are sliced ​​into 2-bit segments, W=8, B=2, and Q=4. If the three low-weight slices corresponding to the lower 6 bits are set as approximate processing objects, then R=3. The partial accumulation path includes three low-weight approximate reduction paths and one high-weight exact reduction path. Specifically, the channel-level approximate addition tree is used to approximate the partial product corresponding to the lower 6-bit weight slices, and the channel-level exact addition tree is used to accurately reduce the partial product corresponding to the highest 2-bit weight slices.

[0036] like Figure 2 As shown, this invention also provides a method for constructing an error-compensating approximate addition tree for digital in-memory computation. This construction method is based on... Figure 1 The partial accumulation structure shown first determines the addition tree object that needs to be approximated and then determines the low-weight approximation processing object based on the weight slice position weight.

[0037] Furthermore, based on the target accuracy, target power consumption, or target hardware overhead, one or more approximation insertion levels are determined from the channel-level addition tree corresponding to the low-weight approximation processing object, and an approximation 4-2 compression module is constructed in the approximation insertion level to perform bit-by-bit approximation compression on the equally weighted partial product bits from multiple input channels.

[0038] Furthermore, the output truth table of the approximate 4-2 compressor is adjusted to form an approximate output mapping with positive and negative error distributions, so that the approximate 4-2 compressor outputs the current position weight result and the next higher position weight result. After the approximate compression results participate in the cumulative reduction according to their position weights, the positive and negative errors in multiple approximate compression results show a statistically significant mutual compensation trend in the overall reduction process, thereby forming an error self-compensation structure.

[0039] like Figure 3 As shown, in this embodiment, the channel-level approximate addition tree is used to reduce the weighted partial products generated by different input channels. The channel-level approximate addition tree adopts a multi-level tree-shaped reduction structure, including multiple computational levels; wherein, the preset computational level uses an approximate 4-2 compression module for approximate compression, and other computational levels use an accurate adder or an accurate compressor for accurate reduction. Figure 3 In the embodiment shown, the preset computation level is the first computation level, and the second computation level and subsequent computation levels adopt an exact reduction structure.

[0040] Specifically, in the first computational level, the equally weighted partial products of the inputs are grouped into groups of four input channels, with each group corresponding to an approximate 4-2 compression module. Figure 3Taking the example shown, the weighted inputs corresponding to partial product channels #0 to #3 form one group, the weighted inputs corresponding to partial product channels #4 to #7 form another group, and so on for the remaining input channels. Each approximate 4-2 compression module receives the weighted partial product bits corresponding to the four input channels and outputs the intermediate compression result. The intermediate compression result is then fed into the second computational level and subsequent computational levels for precise reduction until the reduction result of the channel-level approximate addition tree is obtained.

[0041] In a tree-structured reduction method with a log2(N) level, N represents the number of input channels involved in the reduction, and N is a power of 2. If a channel-level precise addition tree is constructed using two-input precise addition units, the number of precise addition units corresponding to the first computation level is N / 2, and the sum of the number of precise addition units corresponding to the second and subsequent computation levels is N / 4 + N / 8 + ... + 1. Therefore, the first computation level has a large reduction input scale and high hardware overhead. In this embodiment, an approximate 4-2 compression module is placed in the first computation level to perform column-direction approximate compression on the equally weighted partial product bits of the four input channels. Subsequent computation levels continue to complete the precise reduction, thereby reducing the hardware overhead and switching activities of the preceding reduction stage while maintaining the accuracy of the subsequent reduction process.

[0042] In other implementations, the second computational level or multiple computational levels can be set as approximate insertion levels based on target accuracy, target power consumption, or target hardware overhead. For cases where the number of input channels is not an integer multiple of 4, the remaining input channels can participate in subsequent reduction using zero-padding or exact reduction methods.

[0043] like Figure 4As shown, the approximate 4-2 compression module is used to perform bitwise compression on the partial products corresponding to a set of input channels in a channel-level approximate addition tree. In the figure, the vertical direction represents the input channel direction, MSB represents the high-order bit of the partial product, and LSB represents the low-order bit of the partial product. For a partial product with a bit width of M, an approximate 4-2 compression module receives the partial products corresponding to four input channels as inputs, denoted as P0[M-1:0], P1[M-1:0], P2[M-1:0], and P3[M-1:0], where M represents the bit width of a single partial product. The approximate 4-2 compression module internally includes M bitwise-configured approximate 4-2 compressors, each corresponding to one of the M bits of the partial product. For any bit j of the partial product, the weighted input bits P0[j], P1[j], P2[j], and P3[j] from the four input channels are used as the input to the j-th approximate 4-2 compressor. The j-th approximate 4-2 compressor compresses the four equally weighted input bits and outputs two compressed bits, namely the sum bit S[j] and the carry C[j]. The sum bit S[j] corresponds to the j-th bit weight, and the carry C[j] corresponds to the (j+1)-th bit weight. After subsequent bit weight alignment, these bits are fed into the second computational level and subsequent precise reduction levels for reduction.

[0044] Specifically, the carry C[0] generated by the 0th bit approximation 4-2 compressor and the sum S[1] generated by the 1st bit approximation 4-2 compressor have the same bit weight and participate together in the subsequent exact reduction; similarly, the carry C[j] generated by the jth bit approximation 4-2 compressor and the sum S[j+1] generated by the (j+1)th bit approximation 4-2 compressor have the same bit weight. When j=M-1, the carry C[M-1] generated by the (M-1)th bit approximation 4-2 compressor corresponds to the Mth bit weight and participates in the subsequent exact reduction as the highest extended bit weight result in the output of the approximation 4-2 compression module. Thus, the approximation 4-2 compression module compresses the 4 M-bit partial products into intermediate compressed results aligned by bit weight through M bit-by-bit approximation 4-2 compressors and sends them to the second calculation level of the channel-level approximation addition tree and the subsequent exact reduction level for further processing.

[0045] In this embodiment, when both the input feature value slice and the weight slice are 2 bits, each input channel generates a 4-bit partial product, where M=4. The approximate 4-2 compression module internally includes four bit-wise approximate 4-2 compressors. The sum and carry outputs of each approximate 4-2 compressor together constitute the intermediate compression result of the approximate 4-2 compression module, and are sent to the second computational level of the channel-level approximate addition tree and subsequent precise reduction levels for further reduction.

[0046] Figure 5The output mapping relationship and error example of the approximate 4-2 compressor are given. The approximate 4-2 compressor in this invention refers to an approximate circuit that compresses four equally weighted input bits into a low-order result and a high-order result, where the low-order result corresponds to the current bit-weight result, and the high-order result corresponds to the bit-weighted result of the next higher bit. The low-order result is fixed as logic 1, and the high-order result is generated based on the result of detecting all 1s after grouping the input bits.

[0047] In this embodiment, the four input bits IN0 to IN3 are divided into two groups, where the first group includes IN0 and IN1, and the second group includes IN2 and IN3. When two input bits in any input group are simultaneously logic 1, the high-order bit result is output as logic 1; when neither group of input bits satisfies the condition of all bits being 1, the high-order bit result is output as logic 0. Figure 5 It can be seen that when all input bits are 0, the output result is 01, which produces an error of +1 compared to the exact count result; when there is only a single significant bit or three significant bits, the output result is consistent with the exact count result, and the error is 0; when the two significant bits are distributed in different input groups, the output result is 01, which produces an error of -1; when the two significant bits are distributed in the same input group, the output result is 11, which produces an error of +1; when all input bits are 1, the output result is 11, which produces an error of -1.

[0048] Therefore, the approximate 4-2 compressor produces positive error, negative error, or zero error under different input modes. Figure 3 In the illustrated embodiment, multiple approximate 4-2 compressors are configured at the first computational level of the channel-level approximate addition tree. The intermediate compression results output by each approximate 4-2 compressor are further fed into the second computational level and subsequent computational levels for precise reduction. Positive and negative errors generated under different input modes participate in subsequent reduction along with the intermediate compression results, and exhibit a statistically significant mutual compensation trend in the overall reduction process involving multiple bits, multiple input channels, and multiple approximate 4-2 compression modules.

[0049] like Figure 6 As shown in the figure, a gate-level implementation structure of the approximate 4-2 compressor is given. For the approximate 4-2 compressor corresponding to the j-th bit of the partial product, OUT[0] corresponds to the sum bit S[j], and OUT[1] corresponds to the carry C[j]. Among them, OUT[0] is connected to logic 1; OUT[1] is generated by combinational logic consisting of a first NAND gate, a second NAND gate, and a third NAND gate. The first NAND gate receives IN0 and IN1, the second NAND gate receives IN2 and IN3, the outputs of the first NAND gate and the second NAND gate are input to the third NAND gate, and the output of the third NAND gate is used as OUT[1]. Thus, OUT[1] can be expressed as: OUT[1] = (IN0 AND IN1) OR (IN2 AND IN3).

[0050] In this embodiment, the approximate 4-2 compressor is implemented using three NAND gates; when implemented using CMOS static logic, each 2-input NAND gate is implemented using four transistors, and the approximate 4-2 compressor uses a total of 12 transistors.

[0051] This invention is not limited to the embodiments described above. Without departing from the concept of this invention, the number of input channels, input bit width, slice bit width, number of approximate channel-level adder trees, number of levels of approximate 4-2 compression modules, and gate-level implementation can all be adjusted according to actual application requirements, and all should fall within the protection scope of this invention.

Claims

1. An error-compensating approximate addition tree for in-memory digital computation, characterized in that, A partial accumulation path is applied in a digital in-memory computing array. This partial accumulation path includes a channel-level addition tree, a weight-level addition tree, and an eigenvalue-level addition tree. The channel-level addition tree includes a channel-level approximate addition tree and a channel-level exact addition tree. One or more channel-level reduction paths corresponding to lower weights are set as channel-level approximate addition trees, and channel-level reduction paths corresponding to higher weights are set as channel-level exact addition trees. The weight-level addition tree and the eigenvalue-level addition tree maintain an exact reduction structure. The channel-level approximate addition tree adopts a hybrid reduction structure combining approximate compression and exact reduction. The channel-level approximate addition tree includes multiple computational levels, where one or more preset computational levels are equipped with an approximate 4-2 compression module. Computational levels without the approximate 4-2 compression module adopt an exact reduction structure. The approximate 4-2 compression module includes an approximate 4-2 compressor and is used to approximate compress the equally weighted partial product bits from multiple input channels to generate an intermediate compression result. This intermediate compression result participates in subsequent reductions in the channel-level approximate addition tree.

2. The error-compensating approximate addition tree for digital in-memory computation according to claim 1, characterized in that, When the bit width of weight data is W bit and the fragment calculation is performed according to B bit, the weight data forms Q weight slices, wherein , both W and B are positive integers, represents the rounding-up operation; the channel-level reduction paths corresponding to R weight slices starting from the least significant bit are set as channel-level approximate addition trees, and the channel-level reduction paths corresponding to the remaining Q-R higher weight slices are set as channel-level exact addition trees, wherein 1≤R<Q; when W is not divisible by B, the actual bit width of the highest weight slice is less than or equal to B bit.

3. The error-compensating approximate addition tree for digital in-memory computation according to claim 2, characterized in that, The channel-level approximate addition tree includes one or more preset computation levels comprising multiple approximate 4-2 compression modules. Each approximate 4-2 compression module groups four input channels together and performs bitwise approximate compression on the equal-weighted partial products in the same group of input channels. For a partial product with a bit width of M, the approximate 4-2 compression module includes M bitwise-configured approximate 4-2 compressors, each corresponding to one of the M bits of the partial product. The two compressed bits output by the approximate 4-2 compressors are used as the current bit-weighted result and the higher bit-weighted result, respectively, for subsequent reduction.

4. The error-compensating approximate addition tree for digital in-memory computation according to claim 3, characterized in that, The approximate 4-2 compressor receives four equally weighted input bits IN0, IN1, IN2 and IN3, and outputs a low-order result OUT[0] and a high-order result OUT[1]. The low-order result OUT[0] is fixed to logic 1. The four equally weighted input bits are divided into two input groups. The first input group includes IN0 and IN1, and the second input group includes IN2 and IN3. When two input bits in any input group are both logic 1, the high-order result OUT[1] outputs logic 1. When neither input group satisfies the condition of all 1s, the high-order result OUT[1] outputs logic 0. The high-order result OUT[1] satisfies the following logical relationship: OUT[1] = (IN0 AND IN1) OR (IN2 AND IN3).

5. A method for constructing an error-compensating approximate addition tree for digital in-memory computation as described in claim 1, 2, 3, or 4, characterized in that, Includes the following steps: Step 1: Determine a partial accumulation path in the digital in-memory computing array, wherein the partial accumulation path includes a channel-level addition tree, a weight-level addition tree, and an eigenvalue-level addition tree; Step 2: Based on the position weight of the weight slice, determine one or more sets of approximate processing objects from the channel-level reduction path corresponding to the low position weight slice, and determine the channel-level reduction path corresponding to the high position weight slice, as well as the weight-level addition tree and the eigenvalue-level addition tree as the exact reduction objects. Step 3: Based on the target accuracy, target power consumption, or target hardware overhead, determine one or more approximate insertion levels from multiple computational levels of the channel-level addition tree corresponding to the approximate processing object, and determine the computational levels that are not determined as approximate insertion levels as exact reduction levels; Step 4: Construct an approximate 4-2 compression module in one or more approximate insertion levels, so that the approximate 4-2 compression module performs bit-by-bit approximate compression on the equally weighted partial product bits from multiple input channels and generates intermediate compression results; Step 5: Construct the approximate 4-2 compressor in the approximate 4-2 compression module, divide the four equally weighted input bits into two input groups, fix the low-order result to logic 1, and generate the high-order result based on whether any input group satisfies the all-1 condition; Step 6: The current position weight result and the next higher position weight result output by the approximate 4-2 compressor are accumulated and reduced according to their position weights. This allows the positive and negative errors generated by multiple approximate 4-2 compressors to enter the overall reduction process along with the intermediate compression results, and they exhibit a statistically significant mutual compensation trend during the overall reduction process.