Multiple reference voltage based difference input circuit and in-memory computing chip architecture

By using a differential input circuit based on multiple reference voltages and a Shuffle-CIM architecture, the hardware resource overhead problem of multi-bit input activation and weight calculation in analog domain computing architecture is solved, realizing low-overhead and efficient multi-bit MAC operation, and improving computing throughput and energy efficiency.

CN122195390APending Publication Date: 2026-06-12EVERY MOMENT THINKING INTELLIGENT TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
EVERY MOMENT THINKING INTELLIGENT TECH (BEIJING) CO LTD
Filing Date
2026-05-13
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

In existing technologies, analog domain computing architectures suffer from hardware resource overhead that grows exponentially with bit width when activating multiple inputs and calculating weights. In particular, a large amount of overhead is introduced during analog-to-digital conversion and digital domain accumulation, which limits energy efficiency and computing throughput.

Method used

It adopts a differential input circuit based on multiple reference voltages and a full analog domain in-memory computing chip architecture. It realizes the analog domain mapping of multi-bit digital input through multiple reference voltage sources, voltage switching circuits and decoders. Combined with the Shuffle-CIM architecture, it completes multiplication and accumulation operations in the analog domain, eliminating the accumulation requirement in the digital domain, and integrates batch normalization and nonlinear activation operations in the analog domain.

Benefits of technology

It achieves low-overhead multi-bit activation and weighted MAC operations, improves computing throughput, reduces power consumption, breaks through the bottleneck of traditional analog domain CIM architecture, and improves energy efficiency and computing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122195390A_ABST
    Figure CN122195390A_ABST
Patent Text Reader

Abstract

The application provides a difference input circuit based on multiple reference voltages and an in-memory computing chip architecture, wherein the difference input circuit based on multiple reference voltages comprises a first multiplexer, an input end of the first multiplexer is connected to a reference voltage corresponding to a low bit part in a multiple reference voltage source, a control end of the first multiplexer receives a first control signal, and an output end of the first multiplexer is connected to a global voltage node through a first switch; a second multiplexer, an input end of the second multiplexer is connected to a reference voltage corresponding to a high bit part in the multiple reference voltage source, a control end of the second multiplexer receives a second control signal, and an output end of the second multiplexer is connected to the global voltage node through a second switch; and a switch circuit is used for selectively applying a voltage of the global voltage node to an upper plate of a storage capacitor according to weight data stored in the storage unit, and a lower plate of the storage capacitor is used for coupling out an accumulated voltage. The purpose of reducing the hardware resource overhead of the multiple reference voltage type is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of in-memory computing technology, and in particular relates to a differential input circuit based on multiple reference voltages and an in-memory computing chip architecture. Background Technology

[0002] In the era of artificial intelligence, neural networks are widely deployed in edge devices and applied to fields such as image classification, object detection, and natural language processing. The multiply-accumulate (MAC) operation between input activation values ​​and weights is the core operation of neural network inference, and this operation generates a large amount of data.

[0003] In the traditional von Neumann architecture, instructions and data are first written to memory. Then, the control unit obtains instructions through the control bus, sends the data in memory to the computing unit through the data bus, and finally writes the calculation results back to memory through the data bus. Since the storage unit and the computing unit are physically isolated, this method requires frequent data transfer between the storage unit and the computing unit, which limits the speed of data processing. At the same time, the data transfer process also causes a lot of power consumption.

[0004] Computing-in-memory (CIM) has emerged as a promising paradigm for energy-efficient acceleration of neural networks, particularly for models involving repetitive and data-intensive operations, such as MAC computations in convolutional neural networks (CNNs). By integrating memory and computation, the CIM architecture eliminates frequent data transfers between memory and processing units, overcoming the inherent memory bottleneck of traditional von Neumann architectures and reducing latency while improving energy efficiency. CIM architectures can be broadly categorized into digital CIM (DCIM) and analog CIM (ACIM). DCIM performs MAC operations in the digital domain, offering excellent flexibility, scalability, and support for high-precision calculations such as floating-point (FP) computations. However, digital circuitry itself incurs significant area and power costs, and as computational load increases, the resource overhead of digital MAC in DCIM increases dramatically, ultimately limiting energy efficiency. In contrast, ACIM utilizes physical laws such as Kirchhoff's laws, Ohm's law, and charge conservation to perform MAC operations directly in the analog domain, achieving superior computational parallelism and energy efficiency. While ACIMs are more susceptible to noise and variations in process, voltage, temperature, and PVT (pV / V), typically limiting accuracy to binary or two bits, charge-domain ACIMs mitigate these effects by storing computation as charge on capacitors. Capacitors are less affected by power supply fluctuations and device mismatches compared to active devices, thus achieving higher accuracy, such as INT4. Meanwhile, recent neural network research has shown a continued trend from FP32 to low-precision (INT4) computation, demonstrating practical applications in lightweight computing. Despite these advantages, achieving accurate and efficient multi-bit MAC computation remains a fundamental bottleneck for ACIMs. Unlike binary operations, multi-bit processing requires compact analog-to-digital conversion and efficient accumulation while minimizing analog-to-digital conversion. These limitations hinder the scalability of high-precision ACIMs. In recent years, many researchers have focused on designing high-precision ACIM architectures. Through appropriate circuit and architecture-level co-design, analog computation can effectively support multi-bit MAC operations while maintaining excellent energy efficiency. Extensive research has explored multi-bit implementations in ACIMs, but significant limitations remain. For multi-bit input activation, existing work has employed serial bit-by-bit input, voltage conversion schemes based on digital-to-analog converters (DACs), or methods based on MUXs. However, these methods suffer from problems such as low bit-width scalability, low throughput, frequent analog-to-digital conversion, considerable area overhead, or high wiring complexity.For multi-bit weighting, traditional methods rely on digital domain shifting and addition operations or proportional sampling, which cannot effectively handle sign weights or maintain analog domain efficiency. Furthermore, for bit-serial data streams, a high-precision ADC is needed to minimize quantization noise in order to reduce errors during the accumulation process. Summary of the Invention

[0005] To address the problems existing in the prior art, the present invention provides a differential input circuit based on multiple reference voltages and an in-memory computing chip architecture, which at least partially solves the problem that the hardware resource overhead of multiple reference voltages caused by serial input in the prior art increases exponentially with the bit width.

[0006] In a first aspect, embodiments of this disclosure provide a differential input circuit based on multiple reference voltages for an in-memory computing chip, comprising: multiple reference voltage sources, a voltage switching circuit, and a decoder, wherein the in-memory computing chip includes at least one in-memory computing unit; The multi-reference voltage source is used to provide multiple reference voltages at different levels; The input terminal of the decoder is used to receive a multi-digit input, and the decoder is configured to generate a first control signal and a second control signal based on the low-order part and the high-order part of the multi-digit input, respectively. The voltage switching circuit includes a first multiplexer, a second multiplexer, a first switch, and a second switch; The input terminal of the first multiplexer is connected to the reference voltage corresponding to the low-order portion of the multiple reference voltage sources, its control terminal receives the first control signal, and its output terminal is connected to the global voltage node through the first switch. The input of the second multiplexer is connected to the reference voltage corresponding to the high-order portion of the multiple reference voltage sources, its control terminal receives the second control signal, and its output terminal is connected to the global voltage node through the second switch; The in-memory computing unit includes a storage unit, a switching circuit, and a storage capacitor. The input terminal of the switching circuit is connected to the global voltage node, and the control terminal of the switching circuit is connected to the storage unit. The switching circuit is used to selectively apply the voltage of the global voltage node to the upper plate of the storage capacitor according to the weight data stored in the storage unit. The lower plate of the storage capacitor is used to couple out the accumulated voltage.

[0007] Optionally, the circuit includes: In the first stage of the operation, the first switch is closed, the second switch is open, and the decoder selects the first reference voltage to the global voltage node based on the low-order part of the multi-digit input. In the second stage of the operation, the first switch is opened and the second switch is closed. The decoder selects the second reference voltage to the global voltage node based on the high-order part of the multi-digit input. This causes the multi-digit input to be mapped in the analog domain to the difference between the first reference voltage and the second reference voltage, and output by the lower plate of the storage capacitor.

[0008] Optionally, the circuit may further include a sign bit control circuit, which includes a third switch and a fourth switch; The global voltage node includes a positive global voltage node and a negative global voltage node; Wherein, the first switch is connected between the output of the first multiplexer and the positive global voltage node, and the second switch is connected between the output of the second multiplexer and the positive global voltage node; The third switch is connected between the output of the first multiplexer and the negative global voltage node, and the fourth switch is connected between the output of the second multiplexer and the negative global voltage node. The sign bit control circuit further includes a sign bit storage unit and a selection switch. The selection switch is controlled by the sign bit stored in the sign bit storage unit and is used to connect one of the positive global voltage node and the negative global voltage node to the global voltage node.

[0009] Optionally, in the first stage of the operation, the first and fourth switches are closed, and the second and third switches are open, so that the positive global voltage node and the negative global voltage node are respectively connected to the reference voltage selected by the low-order part and the high-order part; In the second stage of the operation, the first and fourth switches are opened, and the second and third switches are closed, so that the positive global voltage node and the negative global voltage node are respectively connected to the reference voltage selected by the high-order part and the low-order part, thereby forming a voltage difference change in opposite directions.

[0010] Optionally, the circuit also includes a charge multiplexing switch connected between the positive global voltage node and the negative global voltage node; Between the first stage and the second stage, there is a charge reuse stage. In the charge reuse stage, the first switch, the second switch, the third switch, the fourth switch, and the lower plate grounding switch connected between the lower plate of the storage capacitor and the ground are all disconnected, and the charge reuse switch is closed, so that the positive global voltage node and the negative global voltage node can share charge to realize energy recovery.

[0011] Secondly, embodiments of this disclosure also provide a fully analog in-memory computing chip architecture, including the input circuit described in any of the first aspects, comprising: Multiple computing unit groups, each computing unit group includes several in-memory computing units, the in-memory computing units are analog domain computing units, and the in-memory computing units are used to perform multiplication and accumulation operations on the input data; The channel mixing module is used to cross-recombine the outputs from each of the computing unit groups in the digital domain to generate the input activation values ​​for the next layer. The in-memory computing unit generates a multiply-accumulate result that does not depend on digital domain accumulation processing based on multi-bit input activation values ​​and multi-bit privileged values ​​within a single computing cycle, and integrates batch normalization and nonlinear activation operations in the analog domain to construct a fully analog computing link.

[0012] Optionally, the architecture also includes: The non-overlapping control signal generation circuit is used to generate control signals for the operation timing of the control chip, including a delay chain and non-overlapping logic to control the signal transition order, so as to ensure the accuracy of capacitor charge distribution.

[0013] Optionally, the fusion batch normalization and nonlinear activation operations within the analog domain are implemented using an analog domain Non-MAC operation ADC, specifically including: The offset parameter unit is set in the same column as the in-memory calculation unit used to perform multiplication. During the multiplication-accumulation operation, the offset is directly accumulated into the multiplication-accumulation result. The dynamic range adjustable capacitive digital-to-analog converter array receives a digital signal representing a scaling factor and dynamically adjusts the quantization full-scale reference voltage of the ADC according to the scaling factor. The successive approximation logic module, which integrates ReLU logic, compares the voltage of the multiplication and accumulation result with the zero voltage before the ADC starts quantization. If it is less than the zero voltage, it directly outputs the quantization result representing the zero value and terminates subsequent quantization. If it is greater than the zero voltage, it starts the complete successive approximation quantization.

[0014] Optionally, the analog domain calculation unit further includes an analog domain weighting circuit; The analog domain weighting circuit is used to perform weighted summation of the analog quantities generated by the amplitude bits of the multi-privileged values ​​within the analog domain.

[0015] Optionally, the weighted summation of the analog quantities generated by the amplitude bits of the multiple privileged values ​​in the analog domain includes: after the multiplication and accumulation operation is completed in each column of the analog domain computing unit, the coupling lines representing different bit weights are physically segmented to form an equivalent capacitance proportional to the bit weight on each of the segmented coupling lines, and then the weighted analog voltage is obtained directly from each coupling line through charge redistribution by closing the shared switch.

[0016] The differential input circuit based on multiple reference voltages provided by this invention adopts a two-stage input mapping logic to map the numerical value to the voltage difference on the upper plate of the capacitor and simultaneously couple it to the lower plate of the capacitor. Only a few reference voltages are needed to achieve numerical mapping equivalent to the traditional scheme with at least 16 reference voltages, thereby reducing the hardware resource overhead of the multi-reference voltage scheme. Attached Figure Description

[0017] The above and other objects, features and advantages of this disclosure will become more apparent from the accompanying drawings, in which like reference numerals generally denote like parts.

[0018] Figure 1 This is a schematic diagram of the simulated domain CIM architecture that uses pSum to implement high-dimensional convolution in the existing technology; Figure 2 A schematic diagram of pSum introduced for weighted multi-privilege value in the prior art; Figure 3 A schematic diagram of the Shuffle-CIM architecture for full analog computation provided in the embodiments of this disclosure; Figure 4 A schematic diagram of a differential input circuit architecture based on multiple reference voltages provided in an embodiment of this disclosure; Figure 5 A schematic diagram of a bidirectional difference input circuit architecture based on weighted symbols provided in an embodiment of this disclosure; Figure 6 A schematic diagram illustrating the sign-magnitude storage format of weights in the CIM array under the bidirectional difference input mode provided in this embodiment of the disclosure; Figure 7a A schematic diagram illustrating the charge change of the upper plate of the capacitor in the first stage during the analysis of charge change on the upper plate of the capacitor in the differential input process provided in this embodiment of the disclosure; Figure 7b A schematic diagram illustrating the charge change of the upper plate of the capacitor in the second stage during the differential input process provided in this embodiment of the disclosure; Figure 8 A schematic diagram of a differential input circuit architecture employing charge multiplexing provided in an embodiment of this disclosure; Figure 9 Interpolation input timing diagram for charge multiplexing mode provided in embodiments of this disclosure; Figure 10 A schematic diagram of a gate circuit based on zero-value decoding provided in an embodiment of this disclosure; Figure 11 A schematic diagram illustrating a weighting method based on bit line segmentation and capacitor reuse provided in an embodiment of this disclosure; Figure 12 A schematic diagram of a non-overlapping control signal generation circuit provided in an embodiment of this disclosure; Figure 13 A schematic diagram of the ANM-ADC logic architecture provided in the embodiments of this disclosure; Figure 14 A schematic diagram of an improved CDAC array provided in an embodiment of this disclosure; Figure 15 A schematic diagram of successive comparisons of an ANM-ADC provided in an embodiment of this disclosure; Figure 16 The timing diagram of the ANM-ADC provided in the embodiments of this disclosure is shown. Detailed Implementation

[0019] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0020] It should be understood that the following specific examples illustrate the implementation of this disclosure, and those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0021] This embodiment discloses an efficient, low-overhead ACIM architecture that can implement MAC for multi-bit activation and multi-bit weights in the simulation domain and execute the entire calculation process.

[0022] The evolution of multi-bit CIM architectures (where both activation values ​​and weights are multi-bit) has encountered new bottlenecks. While existing mixed-signal chip architectures achieve good bit-width expansion flexibility through serial input, they essentially still rely on the shifted addition of partial sums (pSum) in the digital domain. This mixed-signal processing mode introduces high-frequency ADC conversion and data transport overhead, becoming a key bottleneck restricting further improvements in the energy efficiency of analog-domain CIM computing architectures. To address these challenges, this chapter focuses on the research of analog-domain computing circuits based on channel shuffling and bidirectional interpolation, aiming to explore a novel circuit architecture for all-analog-domain computing. First, this embodiment analyzes the generation mechanism of pSum and its profound impact on the system from a mechanistic perspective, and then proposes using the shuffle mechanism to eliminate pSum caused by high-dimensional input, thereby pushing all computation to the analog domain. Based on this, this embodiment further studies and designs an analog-domain CIM architecture based on bidirectional interpolation, achieving efficient integrated computation across the entire analog domain by deeply integrating batch normalization and nonlinear activation functions into the MAC operation and ADC quantization process. The introduction of Shuffle-CIM breaks through the pSum bottleneck in the analog domain CIM architecture, exploring a new path for improving the computing power of sensor nodes under resource constraints.

[0023] In a typical analog domain CIM architecture, MAC results are usually represented as voltage signals. As the number of CIM cells on a single bitline (BL) increases, the dynamic range of the MAC results required per unit voltage interval also expands linearly. However, limited by the supply voltage, the voltage swing on the BL is constant. This leads to a drastic compression of the signal margin (SM) with increasing computational scale, as shown in the following equation: , in, It is the power supply voltage. and These represent the maximum and minimum values ​​within the range of MAC result values, respectively. Reducing the signal margin presents multiple technical challenges. On one hand, a smaller signal margin leads to increased noise sensitivity; various on-chip non-ideal characteristics (such as PVT fluctuations, device mismatch, and coupling noise) are effectively amplified, severely impacting computational accuracy. On the other hand, extracting accurate MAC results from a very small signal margin significantly increases the resolution requirements of the subsequent ADC, resulting in a disproportionate increase in the area and power consumption of the conversion circuit. Due to these physical bottlenecks, the number of CIM units mounted on the BL is typically limited to a few hundred.

[0024] Therefore, for high-dimensional convolution operations with thousands of input channels, it is necessary to break them down into multiple sub-operations to be executed in multiple banks, obtaining pSum from each sub-operation, and then summing them to obtain the complete MAC result. For example... Figure 1 As shown, the analog domain CIM architecture utilizes pSum to implement a typical process for high-dimensional convolution. For an input vector of dimension n, it is split into G groups, each containing k = n / G inputs. Each group of inputs is processed by a Bank with the corresponding dimension. For example, for a convolutional layer with 64 input channels and 64 output channels, the hardware mapping requires 4×4 Banks. From the perspective of the output channel dimension, the four Banks generate corresponding pSums, which are then fed into the digital domain partial summation module and finally merged into a complete channel MAC result. This pSum-based splitting strategy solves the problem caused by reduced signal margin, but it also introduces new problems: The overhead of analog-to-digital conversion increases dramatically: To prevent quantization errors from worsening during multiple accumulations, each bank must be equipped with a high-resolution ADC to cover the dynamic range of pSum. For example, when the MAC result is in the range of [128, 127], an 8-bit ADC is required to achieve accurate analog-to-digital conversion; otherwise, quantization errors will accumulate in the final MAC result, leading to a severe decrease in accuracy. This not only increases the analog-to-digital conversion overhead of a single bank, but also, due to the presence of the group G, the total conversion frequency doubles linearly, directly causing a significant drop in system energy efficiency.

[0025] Memory access and interconnect overhead issues: The introduction of pSum generates a large amount of intermediate data, which is G times the number of output activation values, leading to a significant increase in the frequency of on-chip memory access. At the same time, due to the frequent data interaction between the Bank and the digital domain accumulation module, the large-scale parallel bus interconnects significantly increase the wiring complexity, not only consuming valuable wiring resources but also generating significant wiring parasitic capacitance, resulting in a sharp increase in chip dynamic power consumption and chip area pressure.

[0026] Increased digital domain resource consumption: The size of the adder tree in the partial and accumulator modules expands rapidly with the increase of the number of groups G. Large-scale digital adder tree logic not only significantly increases chip area overhead but also prolongs the delay of the critical path, thus having a serious negative impact on the system's peak throughput and area efficiency.

[0027] The serial input method for multi-bit activation values ​​also introduces pSum. The MAC operation for multi-bit activation values ​​is broken down into multiple 1w1a sub-operations. This process sequentially generates pSum for the corresponding bits in the time dimension. These pSums have different bit significance, so they need to be weighted in the digital domain by shifting and adding to recover the complete range of MAC result values. In this input mode, the impact of pSum on system energy efficiency and hardware performance becomes more significant with increasing bit width.

[0028] The increased overhead of analog-to-digital conversion (ADC) is twofold: First, the ADC conversion frequency increases linearly with the bit width N, directly leading to a significant increase in the power consumption of the CIM module. Second, similar to the pSum effect caused by high input dimensions, to avoid the accumulation of pSum errors in the digital domain, the pSum generated on each bit must undergo high-resolution ADC conversion. This high-frequency, high-precision ADC requirement severely limits further improvements in system energy efficiency.

[0029] Limited computational throughput: Since a single multi-bit MAC operation is split into N computation cycles, at a fixed clock frequency, the effective computational throughput of the system decreases linearly with the increase of the bit width N (or the clock frequency must be increased linearly to maintain the throughput), which greatly limits the computational efficiency at high bit widths.

[0030] In the CIM architecture, multiple privileged values ​​are stored in parallel within CIM cells of multiple adjacent columns, with their position representing their positional weight. During MAC operations, each CIM cell is multiplied by the input activation value, and the results are accumulated and summed in the analog domain through column-level joins to obtain the pSum for the corresponding bit. These pSum values ​​are then weighted and summed in the digital domain using a shift-add module to obtain the final multi-bit MAC result. The calculation process is as follows: Figure 2 As shown. Similar to multi-bit activation values, the splitting of multi-bit privileged values ​​also faces the severe challenges of a dramatic increase in analog-to-digital conversion overhead and limited computational throughput. More seriously, if the CIM architecture simultaneously employs a combination strategy of serial input of N-bit activation values ​​and shift-weighted N-bit privileged values, the activation values ​​and weights are split simultaneously along the bit-weight dimension, resulting in a non-linear deterioration in performance metrics: The ADC conversion overhead grows exponentially: the total conversion frequency of the ADC will increase dramatically with the bit width N in a N² relationship, resulting in an explosive increase in the proportion of analog-to-digital conversion power consumption in the total power consumption of the chip, which seriously deteriorates the system energy efficiency.

[0031] The computational throughput suffers a severe decline: the number of clock cycles required for a single multi-bit MAC operation increases quadratically with the bit width, causing the computational throughput to decline rapidly in an N² trend. This results in a severe "diminishing returns" dilemma for both hardware resource utilization and energy efficiency when traditional CIM architectures handle high-bit-width computational tasks.

[0032] To fundamentally address the significant overhead of pSum on analog-to-digital conversion in the analog domain and storage and accumulation in the digital domain, this paper implements a Shuffle-CIM architecture that deeply integrates the design concept of the lightweight convolutional neural network ShuffleNet and proposes a Shuffle-CIM architecture suitable for analog domain CIM.

[0033] Shuffle-Net's grouping characteristics are highly compatible with the hardware physical constraints caused by signal margin in the analog domain CIM architecture in terms of topology. Based on this insight, this design constructs a system by logically mapping the grouping strategy in the algorithm to the Bank array in the hardware. Figure 3 The Shuffle-CIM architecture shown has the following core features: Bank-level grouping and computational dimensionality reduction (dimension alignment): In analog domain CIM, the number of units mounted on the BL (corresponding to the input dimension) is limited by voltage swing and signal margin. Shuffle-CIM actively aligns the physical dimension of the hardware bank with the grouping dimension k of the algorithm. This design allows each bank to independently process a complete feature group, ensuring that the simulation computation always runs in a more robust low-dimensional range, thus physically guaranteeing the linearity and accuracy of MAC operations.

[0034] Data Stream Reconstruction and pSum Elimination: In Figure 3 In this topology, input features are assigned to G independent hardware banks. Since the computational dimension is decoupled at the algorithm level, the digital output of each bank after ADC conversion represents the complete MAC result for that channel, rather than the pSum that requires further accumulation. This design eliminates the need for cross-bank accumulation at the logical source, overcoming the bottleneck caused by the storage and accumulation of pSum in the digital domain in traditional analog domain CIM architectures.

[0035] The pSum generated by the MAC operation itself, which involves multi-bit activation values ​​and multi-bit privileged values, is eliminated. (1) Parallel input and single-cycle calculation: Multi-bit activation values ​​are input in parallel to ensure that the MAC operation is completed in a single calculation cycle; (2) Analog domain weighting: The ability to weight multi-bit privileged values ​​is realized in the analog domain, generating an analog quantity representing the result of multi-bit MAC in a single calculation cycle, rather than generating multiple pSums and then performing digital weighting. The pSum introduced at the bit weight level due to multi-bit splitting is eliminated from the physical source, and multi-bit MAC can be completed in one calculation (referred to as One-ShotMAC), thereby greatly improving the computing throughput and reducing the power consumption caused by frequent ADC conversion and pSum storage and accumulation.

[0036] Analog Domain Non-MAC Computation: By eliminating the need for caching and accumulation of pSum, non-multiply-accumulate operations (such as BN, activation, etc., hereinafter referred to as Non-MAC) are no longer limited to implementation in the digital domain. Through innovative analog circuit design, Shuffle-CIM directly integrates Non-MAC computation within the analog domain CIM, constructing a fully analog computation chain. This approach allows both MAC and Non-MAC operations to be efficiently completed in the analog domain, while the digital domain only undertakes the tasks of storing and transmitting activation values. Through this division of labor, the low-power computing advantages of the analog domain and the storage and transmission reliability of the digital domain can be synergistically leveraged to maximize the overall system benefits.

[0037] Digital domain shuffle mechanism: The outputs of each bank directly enter a dedicated shuffle buffer. According to the algorithm's preset shuffle rules, the activation values ​​in the buffer are simply cross-reorganized and can be directly used as the input for the next convolutional layer. This mechanism, with its near-zero overhead static topology mapping, solidifies the feature fusion process within the data flow path, further reducing the hardware resource overhead of the digital domain.

[0038] This example proposes a differential input circuit based on multiple reference voltages for use in-memory computing chips. The circuit is characterized by including: multiple reference voltage sources, voltage switching circuits, and decoders. The in-memory computing chip includes at least one in-memory computing unit. The multi-reference voltage source is used to provide multiple reference voltages at different levels; The input terminal of the decoder is used to receive a multi-digit input, and the decoder is configured to generate a first control signal and a second control signal based on the low-order part and the high-order part of the multi-digit input, respectively. The voltage switching circuit includes a first multiplexer, a second multiplexer, a first switch, and a second switch; The input terminal of the first multiplexer is connected to the reference voltage corresponding to the low-order portion of the multiple reference voltage sources, its control terminal receives the first control signal, and its output terminal is connected to the global voltage node through the first switch. The input of the second multiplexer is connected to the reference voltage corresponding to the high-order portion of the multiple reference voltage sources, its control terminal receives the second control signal, and its output terminal is connected to the global voltage node through the second switch; The in-memory computing unit includes a storage unit, a switching circuit, and a storage capacitor. The input terminal of the switching circuit is connected to the global voltage node, and the control terminal of the switching circuit is connected to the storage unit. The switching circuit is used to selectively apply the voltage of the global voltage node to the upper plate of the storage capacitor according to the weight data stored in the storage unit. The lower plate of the storage capacitor is used to couple out the accumulated voltage.

[0039] like Figure 4 As shown, this input circuit uses a two-stage input mapping logic to map the numerical value to the voltage difference on the upper plate of the capacitor and simultaneously couple it to the lower plate of the capacitor. Only 7 reference voltages are needed to achieve the numerical mapping equivalent to the traditional 16 reference voltage scheme, breaking through the architectural bottleneck of the hardware resource overhead of multi-reference voltage DACs increasing exponentially with the bit width.

[0040] The mapping of the 4-bit activation value to the multi-reference voltage DAC in differential input mode is shown in Table 1. The 4-bit activation value is divided into high 2 bits and low 2 bits, corresponding to the 4 Vref reference voltages respectively. Figure 4 Vref[0]~Vref[6] in Table 1 correspond to 0×VDD / 15~15×VDD / 15 respectively. Since the reference voltages corresponding to the lower 2 bits '00' and the higher 2 bits '00' are both 3×VDD / 15, the mapping of the 4-bit activation value only needs to include 7 reference voltages, including VDD (15×VDD / 15) and GND (0×VDD / 15). The dashed area in Table 1 represents the voltage difference between the reference voltages corresponding to the higher 2 bits and the lower 2 bits, that is, the analog mapping corresponding to the 4-bit digital input. For example, for the 4-bit activation value '0110' (6), the corresponding voltage difference is 7×VDD / 15 corresponding to the higher 2 bits '01' minus 1×VDD / 15 corresponding to the lower 2 bits '10', that is, 6×VDD / 15. The specific process of the difference input is divided into the following two stages: Phase 1: Figure 4The control signal Ch1 controls the first switch to close, and the control signal Ch2 controls the second switch to open. The selection signal generated by the decoder from the lower two bits IA[1:0] acts on the multiplexer MUX, selecting the corresponding reference voltage connected to ia_v, and entering the CIM cell (the CIM cell includes a 6T-SRAM, a CMOS switch, and an NMOS switch, using a total of 9 transistors, abbreviated as 9T1C-CIM). At this time, the lower plate cl node of the capacitor in the 9T1C-CIM cell is connected to GND. If the 1-privilege value q stored in the SRAM is '1', the CMOS switch is turned on, and the upper plate of the capacitor is charged to the corresponding reference voltage by ia_v; if the stored 1-privilege value q is '0', the NMOS switch is turned on, and the upper plate of the capacitor is connected to GND.

[0041] The second stage: First, disconnect the lower plate of the capacitor in the 9T1C-CIM cell from GND, leaving it in a floating state. Simultaneously, control signal Ch1 controls the first switch to open, and control signal Ch2 controls the second switch to close. The selection signal generated by the decoder from the high 2 bits of the input IA[3:2] acts on the multiplexer MUX, selecting the corresponding reference voltage connected to ia_v, which then enters the 9T1C-CIM cell. If the 1-privilege value q stored in the SRAM is '1', the CMOS switch is turned on, and the voltage of the upper plate of the capacitor changes along with ia_v. Since the lower plate of the capacitor is in a floating state, according to the law of charge conservation, the voltage of the lower plate will change along with the voltage of the upper plate. Therefore, the voltage difference change of the upper plate is mapped to the lower plate. If the stored 1-privilege value q is '0', the NMOS switch is turned on, and there is no voltage change between the upper and lower plates of the capacitor.

[0042] Through the above two stages, the result of multiplying the 4-bit digital input activation value and the 1-privilege value is mapped to the lower plate of the capacitor in the 9T1C-CIM cell, and is expressed as the difference in voltage change of the lower plate.

[0043] Table 1. Multi-reference voltage DAC mapping table in differential input mode (4 bits): like Figure 4 The circuit shown can perform multiplication of unsigned activation values ​​and unsigned weights, but it lacks the ability to handle the sign bit of the weights. However, in neural network inference, signed weights are crucial for maintaining model accuracy. Therefore, this embodiment further proposes a bidirectional difference input circuit architecture based on the sign bit of the weights. Figure 5 As shown, the input current also includes a sign bit control circuit, which includes a third switch and a fourth switch; The global voltage node includes a positive global voltage node and a negative global voltage node; Wherein, the first switch is connected between the output of the first multiplexer and the positive global voltage node, and the second switch is connected between the output of the second multiplexer and the positive global voltage node; The third switch is connected between the output of the first multiplexer and the negative global voltage node, and the fourth switch is connected between the output of the second multiplexer and the negative global voltage node. The sign bit control circuit further includes a sign bit storage unit and a selection switch. The selection switch is controlled by the sign bit stored in the sign bit storage unit and is used to connect one of the positive global voltage node and the negative global voltage node to the global voltage node. By controlling the non-overlapping switches via control signals Ch1 and Ch2, a positive global voltage node ia is generated in both stages of the differential input. pos With negative global voltage node ia neg Two opposite voltage differences: Forward global voltage node ia pos In the first stage, control signal Ch1 controls the corresponding switches (first and fourth switches) to close, and control signal Ch2 controls the corresponding switches (second and third switches) to open. pos First, the reference voltage selected by the lower two bits of input IA[1:0] is connected; after entering the second stage, control signal Ch1 controls the corresponding switches (first and fourth switches) to open, and control signal Ch2 controls the corresponding switches (second and third switches) to close. pos Switch to the reference voltage selected by the high 2-bit input IA[3:2]. The two voltage switchings produce a positive voltage change.

[0044] Negative global voltage node ia neg Its temporal logic and ia pos The exact opposite is true. In the first stage, a reference voltage selected by the high two bits of the input is applied, and in the second stage, the input is switched to a reference voltage selected by the low two bits. This switching from high to low level produces a negative voltage change.

[0045] The sign bit of the weight w sign This determines the direction of the voltage difference connected to the 9T1C-CIM unit. If w sign If it is '0', then ia pos If selected, then ia neg It is selected. Therefore, ia pos with ia negThis mechanism can generate positive and negative voltage changes on the upper plate of a capacitor, which are then coupled to the lower plate. The direction of the voltage change represents the sign of the weight; a positive voltage change represents a positive weight, and a negative voltage change represents a negative weight. Through this bidirectional differential input mechanism, the result of multiplying an unsigned activation value with a signed weight is mapped to the difference in voltage changes on the lower plate of the capacitor. Under this mechanism, the encoding of the multiplier privilege value must adopt a 'sign-amplitude' format, with the sign bit and amplitude bit stored in the sign bit unit and the 9T1C-CIM unit, respectively. Figure 6 As shown. The sign bit is only responsible for strobing ia. pos with ia neg The multiplication operation is performed within the amplitude bit CIM cell and coupled to its corresponding column cl line for further accumulation.

[0046] like Figure 7a and Figure 7b The diagram shows the changes in the upper plate of the capacitor within the 9T1C-CIM unit during the MAC operation. Specifically, the upper plate of the capacitor in the 9T1C-CIM unit storing the positive weight value '1' (hereinafter referred to as the positive weight unit) is connected to ia. pos The capacitor upper-level board in the 9T1C-CIM unit (hereinafter referred to as the negative weight unit) storing the negative weight value '1' is connected to ia neg During the two stages of differential input, the voltage across the capacitor's upper plates exhibits a regular fluctuation. This embodiment analyzes this pattern in depth and subsequently proposes a charge reuse technology based on the analysis results, aiming to utilize the potential difference between the plates to achieve energy recovery and secondary utilization.

[0047] Analysis of the charging and discharging characteristics of the upper plate of the capacitor: The first stage involves the discharge of the upper plate of the capacitor within the positive weight unit. For example... Figure 7a As shown, during the process of returning from the second stage to the first stage, the voltage on the upper plate of the capacitor in the positive weight unit changes with ia. pos From high level back to low level, the lower plate from V MAC (V) MAC This means that the value after accumulating the multiplication results in a column of CIM cells is changed to GND.

[0048] In the first and second stages, the voltage difference between the upper and lower capacitor plates after stabilization is as follows: , ; in, It is the first stage of ia pos The reference voltage for gating, It is the second stage ia pos The reference voltage for gating. The voltage change across the capacitor from the second stage to the first stage is as follows: ; If ∆V > 0, the upper plate of the capacitor needs to be charged in the first stage; if ∆V < 0, the upper plate of the capacitor discharges outward in the first stage. Only when the 4-bit input activation value is '0000'... = All other cases are < and - ≤-1×VDD / 15. For the MAC operation results of unsigned activation values ​​and signed weights, the values ​​are usually normally distributed about the y-axis, and σ is relatively small. That is, most MAC results fall near 0, corresponding to the simulation domain as... The voltage magnitude is close to GND. Therefore, in most cases, ∆V<0 holds true, and the capacitor in the positive weight cell discharges outward.

[0049] In the first stage, the upper plate of the capacitor in the negative weight unit is charged. For example... Figure 7a As shown, during the process of returning from the second stage to the first stage, the voltage on the upper plate of the capacitor in the negative weight unit changes with ia. neg From low level to high level, the lower plate changes from V MAC Change to GND.

[0050] In the first and second stages, the voltage difference between the upper and lower plates of the capacitor after stabilization is as follows: , ; in, It is the first stage of ia neg The reference voltage for gating, It is the second stage ia neg The reference voltage for gating. The voltage change across the capacitor from the second stage to the first stage is as follows: ; If ∆V > 0, the upper plate of the capacitor needs to be charged in the first stage; if ∆V < 0, the upper plate of the capacitor discharges outward in the first stage. Only when the 4-bit input activation value is '0000'... = All other cases are > and - ≥1×VDD / 15. Because The voltage magnitude is close to GND. Therefore, in most cases, ∆V>0 holds true, and the capacitor in the negative weight cell is charged.

[0051] The second stage involves charging the upper plate of the capacitor within the positive weight unit. For example... Figure 7b As shown, during the transition from the first stage to the second stage, the voltage on the upper plate of the capacitor within the positive weight unit changes with ia. pos As the voltage level changes from low to high, the lower plate is disconnected from GND and left floating. The charge coupled to the lower plate is redistributed to form a charge from V... MAC .

[0052] In the first and second stages, the voltage difference between the upper and lower plates of the capacitor after stabilization is as follows: , ; From the first stage to the second stage, the voltage across the capacitor changes as follows: ; If ∆V > 0, the upper plate of the capacitor needs to be charged in the second stage; if ∆V < 0, the upper plate of the capacitor discharges in the second stage. Only when the 4-bit input activation value is '0000'... = All other cases are < and - ≥1×VDD / 15. Because The voltage magnitude is close to GND. Therefore, in most cases, ∆V>0 holds true, and the capacitor in the positive weight cell is charged.

[0053] The second stage involves the discharge of the upper plate of the capacitor within the negative weight unit. For example... Figure 7b As shown, during the transition from the first stage to the second stage, the voltage on the upper plate of the capacitor within the negative weight unit changes with ia. neg As the voltage drops from high to low, the lower plate disconnects from GND and becomes floating. The charge coupled to the lower plate is redistributed to form a charge from V... MAC .

[0054] In the first and second stages, the voltage difference between the upper and lower plates of the capacitor after stabilization is as follows: , ; From the first stage to the second stage, the voltage across the capacitor changes as follows: ; If ∆V > 0, the upper plate of the capacitor needs to be charged in the second stage; if ∆V < 0, the upper plate of the capacitor discharges in the second stage. Only when the 4-bit input activation value is '0000'... = All other cases are > and - ≤-1×VDD / 15. Since the V_MAC voltage is close to GND, in most cases, ∆V<0 holds true, and the capacitor in the negative weight cell discharges outward.

[0055] When the capacitor in the 9T1C-CIM unit discharges, the voltage connected to its upper plate is always the reference voltage {Vref[0], Vref[1], Vref[2], Vref[3]} corresponding to the low-order part; when the capacitor in the 9T1C-CIM unit is charged, the voltage connected to its upper plate is always the reference voltage {Vref[3], Vref[4], Vref[5], Vref[6]} corresponding to the high-order part. Throughout the entire calculation cycle of the 9T1C-CIM unit, the reference voltage {Vref[3], Vref[4], Vref[5], Vref[6]} corresponding to the high-order part of the multi-digit input is constantly charging the capacitor, while the capacitor is also constantly discharging to the reference voltage {Vref[0], Vref[1], Vref[2], Vref[3]} corresponding to the low-order part of the multi-digit input.

[0056] like Figure 8 As shown, the input circuit disclosed in this embodiment also includes a charge multiplexing switch, which is connected between the positive global voltage node and the negative global voltage node; Between the first and second stages, there is a charge reuse stage. During this stage, the first, second, third, and fourth switches, as well as the lower plate grounding switch connecting the lower plate of the storage capacitor to ground, are all open. The charge reuse switch is closed, allowing charge sharing between the positive and negative global voltage nodes, thus achieving energy recovery. The charge reuse switch is controlled by a signal. and control signals control.

[0057] like Figure 8 As shown, through ia pos with ia neg A set of switches is added between them, and the control signals are controlled by the inverse signals of control signal Ch1 and control signal Ch2. and control signals To control, thereby in ia pos with ia neg A loop is formed between them. During the differential input process, the control timing of the switch is as follows: Figure 9 As shown. Between the first and second stage transition, a charge energy reuse stage is added to perform the following operations: Disconnect the drive. Control signals Ch1 and Ch2, as well as the lower plate grounding switch Cl_gnd, are all disconnected, causing ia to...pos ia neg The lower plate of the capacitor is completely suspended.

[0058] Charge sharing. Control signal at this time. and control signals The control switch is closed, ia pos with ia neg Short-circuiting means that the upper plates of the positive and negative weight cells are short-circuited. Since the potential of the upper plate of the capacitor in the negative weight cell is higher than that in the positive weight cell, the charge spontaneously moves from the negative weight cell to the positive weight cell and eventually reaches equilibrium, realizing the in-situ reuse of energy in the CIM cell array.

[0059] For the positive weight unit, during the transition from a low level in the first stage to a high level in the second stage, the upper plate of its capacitor is pre-raised to an intermediate level through the charge reuse stage. This 'pre-charging' process significantly reduces the amount of charge injected into the capacitor by the reference voltage in the second stage, thereby significantly reducing the power consumption of the reference voltage source, which optimizes the energy consumption of the activation value input from the source. It should be emphasized that the balance voltage generated in the charge energy reuse stage is only an intermediate state, and the voltage difference between the upper plate and the lower plate of the capacitor is still determined by the reference voltages of the second and first stages. Therefore, this mechanism optimizes the dynamic power consumption of the calculation without affecting the accuracy of MAC operation. From a deeper physical perspective, this technical solution constructs a closed-loop cycle of 'energy-information-energy' in the underlying circuit. During circuit operation, energy first drives the hardware components as the physical carrier of information mapping, accurately representing specific numerical information through the high and low levels of the voltage. When the first stage sampling ends, the voltage difference between the upper and lower plates of the capacitor in the 9T1C-CIM unit is completely determined, and the lower plate changes from grounded to floating. At this point, the information has been successfully converted into static charge storage between the plates, and the physical-level information transfer is complete. At this node, the charge that originally drove the first-stage mapping, no longer responsible for information mapping at the current moment, is often considered 'redundant energy' and directly discharged to the ground or power source in traditional architectures. However, this design, through an innovative charge-sharing mechanism, recovers this energy at the end of its mission and cleverly transforms it into 'initial power' for the next information representation (i.e., the second-stage mapping), thereby maximizing energy utilization within the closed loop of information flow.

[0060] Building upon the reduction of reference voltage power consumption in the bidirectional differential input circuit through charge reuse, this embodiment further proposes an input gating technique based on zero-value decoding, specifically addressing the distribution characteristics of the input activation values. This aims to eliminate power consumption from ineffective switching in certain situations. In the bidirectional differential input circuit, the 4-bit input activation value is divided into high 2 bits and low 2 bits. A 2-to-4 decoder generates a strobe signal to drive the MUX to select the corresponding reference voltage V. ref Subsequently, the reference voltage enters the CIM unit via the non-overlapping control signals Ch1 and Ch2 to control the switching group. For this part of the circuit, its dynamic power consumption mainly consists of the frequent switching of the switching group and the toggling of the decoding gating logic. (Referring to the activation values ​​and reference voltage V in Table 1...) ref The mapping relationship reveals that when the activation value is 0 ('0000'), the reference voltage V corresponding to the high 2 bits and low 2 bits is... ref All are 3 × VDD / 15. In this state, the switching of the control switch group by control signals Ch1 and Ch2 will not cause a change in the reference voltage level entering the CIM unit. That is, the switching action at this time contributes nothing to the calculation results and is considered redundant power consumption. Considering that the proportion of zero-value activations in a CNN model is typically close to 50%, this ineffective switching leads to unnecessary energy loss in high-concurrency computing arrays. For sensor nodes with extremely limited resources, it is necessary to save on static or dynamic power consumption.

[0061] Based on the above analysis, this embodiment proposes the following: Figure 10 The diagram shows a gated circuit based on zero-value decoding. When the input activation value IA[3:0] = 0000, the high 2 bits and low 2 bits of the 2-to-4 decoder simultaneously convert En_V ref[3] The enable signal is pulled high. These two En_V signals... ref[3] The enable signal then generates a low-level active all-zero flag signal through a NAND gate. This all-zero flag signal is further NAND-NOT-ed with the inverted signals Ch1_b and Ch2_b of control signals Ch1 and Ch2, forcing the outputs Ch1_En and Ch2_En to remain high (while their inverted signals Ch1_Enb and Ch2_Enb remain low). In this state, the control signals Ch1 and Ch2 control the switch group, which is forcibly locked in the ON state by the outputs Ch1_En and Ch2_En, thus eliminating the dynamic switching power consumption caused by invalid flips. Simulation results show that, with this design, the dynamic power consumption of the bidirectional differential input circuit is reduced by 31.8%, significantly improving the energy efficiency of the input driver stage.

[0062] In addition, such as Figure 3As shown, this embodiment discloses a fully analog in-memory computing chip architecture, including an input circuit, comprising: Multiple computing unit groups, each computing unit group includes several in-memory computing units, the in-memory computing units are analog domain computing units, and the in-memory computing units are used to perform multiplication and accumulation operations on the input data; The channel mixing module is used to cross-recombine the outputs from each of the computing unit groups in the digital domain to generate the input activation values ​​for the next layer. The in-memory computing unit generates a multiply-accumulate result that does not depend on digital domain accumulation processing based on multi-bit input activation values ​​and multi-bit privileged values ​​within a single computing cycle, and integrates batch normalization and nonlinear activation operations in the analog domain to construct a fully analog computing link.

[0063] like Figure 11 As shown, the analog domain computation unit also includes an analog domain weighting circuit; The analog domain weighting circuit is used to perform weighted summation of the analog quantities generated by the amplitude bits of the multi-privileged values ​​within the analog domain.

[0064] The weighted summation of the analog quantities generated by the amplitude bits of the multiple privileged values ​​in the analog domain includes: after the multiplication and accumulation operation is completed in each column of the analog domain computing unit, the coupling lines representing different bit weights are physically segmented to form an equivalent capacitance proportional to the bit weight on each of the segmented coupling lines, and then the weighted analog voltage is obtained directly by closing the shared switch through charge redistribution.

[0065] The MAC operator circuit in this technical solution is a charge-coupled modular induction (CIM) circuit. The multiplication result is capacitively coupled to the lower plate for accumulation. After accumulation, the upper plate of the capacitor is connected to a fixed potential (reference voltage) and remains unchanged, while the lower plate is in a floating state. After the charge spontaneously balances and stabilizes, the final accumulated voltage V is formed. MAC In this state, the capacitors of all CIM cells in the same column are shorted through the lower plate, effectively forming an equivalent large parallel capacitor.

[0066] Since the sign bit only serves to select ia pos or ia neg Since they do not directly participate in the calculation, the weighting of the multi-bit privileged values ​​only needs to be performed between the couple lines (hereinafter referred to as cl) corresponding to the amplitude bits. Figure 11As shown, the three cls corresponding to the amplitude bits in the 4 privileged values ​​are cl[2], cl[1], and cl[0], with a bit weight ratio of 4:2:1. To achieve weighting between cls in the analog domain, the core is to construct a capacitor ratio difference that matches the bit weight. After the MAC operation is completed, the cls are segmented by controlling the Seg switch; the number of units connected to each segment is precisely allocated, and an equivalent capacitor ratio of 4n:2n:n is formed at the bottom of the three cls, so that the required bit weight difference is accurately corresponded in the hardware structure. After the segmentation is completed, the Share switch is closed, so that the lower plate of the equivalent parallel large capacitor on the three cls is short-circuited. The charge flows spontaneously between the lower plates under the drive of the potential difference and eventually reaches equilibrium. The voltage at this time represents the result of weighting pSum represented by the three columns of cls according to the bit weight.

[0067] The weighted design based on bit line segmentation and capacitor reuse has the following significant advantages: (1) Capacitor reuse and area optimization. This scheme directly reuses the computational capacitors in the 9T1C-CIM cell as weighting proportional capacitors, without the need to introduce additional C-2C capacitor ladder networks or dedicated proportional capacitor arrays, which greatly saves hardware area. Since the proportional capacitors used for weighting reuse the computational capacitors in the CIM cell, area is saved.

[0068] (2) Suppressing the effect of charge injection. During the weighting stage, the equivalent parallel capacitance formed on each cl is relatively large. Therefore, the charge injection effect introduced when the Seg and Share switches switch is switched is significantly weakened under the large load capacitance, thus maintaining the linearity and accuracy of the accumulation result.

[0069] like Figure 12 As shown, the fully analog in-memory computing chip architecture also includes: a non-overlapping control signal generation circuit, which generates control signals for controlling the timing of chip operations, including a delay chain and non-overlapping logic to control the signal transition order, so as to ensure the accuracy of capacitor charge distribution.

[0070] Specifically, the Cl_gnd signal is used to control the connection between the coupling line and ground (i.e., to reset the coupling line). Its timing is consistent with the control signal Ch1 during the calculation process, and therefore it can be directly derived from Ch1. However, considering the charge injection effect caused by the switch switching, the connection between the lower plate of the capacitor and ground must be disconnected before the control signals Ch1 and Ch2 control the switch group switching, leaving it in a floating state. At this time, the injected charge introduced by the control signals Ch1 and Ch2 controlling the switch group switching occurs during the transition of the upper plate of the capacitor from the first stage to the second stage, thus not affecting the final difference calculation result. Therefore, the transition of Cl_gnd must precede that of signal Ch1_b. Since signal Ch1_b needs to drive each row in the CIM array, its load is much higher than that of Cl_gnd, naturally possessing a longer switching delay. Therefore, in this embodiment, only a single-stage inverter is used to distinguish the phase, ensuring that Cl_gnd transitions before Ch1_b. Meanwhile, to prevent the initial voltage of the upper plate of the capacitor from changing due to the simultaneous conduction of the control switches by control signals Ch1 and Ch2, the change of Ch1_b is strictly limited to after Ch2_b by NAND gate logic. Ch1_b and Ch2_b are the inverted signals of control signals Ch1 and Ch2, respectively.

[0071] This embodiment derives the Seg and Share signals from a single original trigger signal, thereby minimizing global signal resources while ensuring computational throughput. To meet the timing constraints between signals, a controlled delay is generated by inserting a 9-stage inverter chain between the Seg and Share signal paths. This physical delay ensures that the charge-sharing action is performed only after the segmentation is completely completed, thus achieving strictly non-overlapping weighted timing within a single cycle.

[0072] like Figure 13 The aforementioned fusion batch normalization and nonlinear activation operations within the analog domain are implemented through an analog domain Non-MAC operation ADC, specifically including: The offset parameter unit is set in the same column as the in-memory calculation unit used to perform multiplication. During the multiplication-accumulation operation, the offset is directly accumulated into the multiplication-accumulation result. The dynamic range adjustable capacitive digital-to-analog converter array receives a digital signal representing a scaling factor and dynamically adjusts the quantization full-scale reference voltage of the ADC according to the scaling factor. The successive approximation logic module, which integrates ReLU logic, compares the voltage of the multiplication and accumulation result with the zero voltage before the ADC starts quantization. If it is less than the zero voltage, it directly outputs the quantization result representing the zero value and terminates subsequent quantization. If it is greater than the zero voltage, it starts the complete successive approximation quantization.

[0073] By employing bidirectional difference input based on sign bits and weighted summation based on bit line segments, the CIM array has acquired the core capability to implement One-Shot MAC. In traditional analog-domain CIM architectures, the data flow typically follows a 'analog-domain MAC + digital-domain post-processing' pattern: multi-bit MAC results need to be converted to the digital domain via an ADC, followed by a series of Non-MAC operations (such as BN, activation, etc.), ultimately completing the entire convolutional layer computation and outputting feature maps (activation values). To optimize the implementation of BN operations in the digital domain, it is simplified to a combination of multiplication and addition. However, this strategy faces significant challenges in the analog domain because of the limitations of V... MAC Multiplicative scaling typically requires a gain-configurable operational amplifier, which incurs significant resource overhead and is accompanied by gain error. Therefore, this implementation first optimizes the computational principles by deeply integrating Batch Normalization (BN), activation, and quantization, resulting in the Non-MAC operator shown in the following equation: , Where x represents the MAC result, b is the offset, and s is the scaling factor, where b and s are parameters learned channel-by-channel during network training; ReLU (Rectified Linear Unit) is the activation function, implementing non-negative selection of activation values; clip filters the range of output activation values ​​to narrow the ADC quantization range and resolution requirements; Quant corresponds to the ADC quantization. This operator breaks away from the traditional analog domain CIM architecture's data flow of quantization, then BN, and finally activation, constructing an integrated post-processing mechanism that combines quantization, BN, and activation.

[0074] The offset b is directly added to the result x. At the circuit level, this technical solution adds some CIM units as offset units to each column, using them to store the offset and participate in the calculation synchronously with the convolution units. The number of offset units can be determined as an adjustable parameter during network training.

[0075] This technical solution achieves seamless integration of offset b and MAC operation. This approach enables V MAC Bias information is included in the signal formation stage, thus avoiding additional circuit overhead. In contrast, the fusion of operators such as scaling, activation, and quantization in the analog domain depends on the dynamic configuration of the ADC quantization range.

[0076] like Figure 14As shown, the circuit maps the scaling factor s to the change in the full-scale voltage (dynamic range) of the ADC through the charge redistribution mechanism. The physical essence of this operation is to transform the amplitude transformation (scaling) of the measurement object into the scaling adjustment of the measurement scale, thus synchronously completing the scaling logic in the Non-MAC operator during the quantization process. In the initialization stage, the scaling factor is written into the register. Through switch control, the dynamic range adjustment capacitor array (with a radix of 3C) selectively participates in the charge redistribution, as Figure 14 shown in the pseudo-differential CDAC (capacitive digital-to-analog converter) topology. This structure forms two groups of 4:2:1:1 quantization capacitor arrays through a symmetric switched-capacitor architecture, thus having the ability to support the'single-step switching' logic. As Figure 15 shown, during the successive approximation process, the CDAC only triggers a single switch to switch, effectively suppressing the generation of voltage glitches.

[0077] The ANM-ADC directly embeds the ReLU logic into the SAR control process to achieve conditional comparison and conversion, and its timing is as Figure 16 shown: (1) First positive and negative number judgment. After the SOC signal is set low, the ADC (analog-to-digital converter) first starts the first comparison cycle and compares V MAC with the analog voltage GND corresponding to the value 0.

[0078] (2) ReLU negative value mode. If the output OUTN of the differential comparator is low level and OUTP is high level, it means that V MAC < GND. At this time, the ACT_EN signal remains low, the ADC conversion end signal EOC signal is directly pulled high, and the ADC will directly turn off the subsequent 4 quantization comparison processes and output the result "0000", thus saving the invalid conversion overhead.

[0079] (3) ReLU positive value mode. If the output of the comparator OUTN is high level and OUTP is low level, it means that the comparator judges V MAC > > GND, and the random ACT_EN signal is pulled high to enable the subsequent 4 comparisons to complete the whole process of quantization. The quantized activation value enters the digital domain for storage and transmission.

[0080] In this disclosure, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The block diagrams of devices, apparatuses, devices, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as "comprising," "including," "having," etc., are open-ended terms meaning "including but not limited to," and are used interchangeably with them. The terms "or" and "and" as used herein refer to the terms "and / or," and are used interchangeably with them unless the context clearly indicates otherwise. The term "such as" as used herein refers to the phrase "such as but not limited to," and is used interchangeably with it.

[0081] Additionally, as used herein, the "or" used in a list of items beginning with "at least one" indicates a separate list, such that a list of, for example, "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not imply that the described example is preferred or better than other examples.

[0082] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A differential input circuit based on multiple reference voltages for use in-memory computing chips, characterized in that, include: The chip includes a multi-reference voltage source, a voltage switching circuit, and a decoder, and the in-memory computing chip includes at least one in-memory computing unit. The multi-reference voltage source is used to provide multiple reference voltages at different levels; The input terminal of the decoder is used to receive a multi-digit input, and the decoder is configured to generate a first control signal and a second control signal based on the low-order part and the high-order part of the multi-digit input, respectively. The voltage switching circuit includes a first multiplexer, a second multiplexer, a first switch, and a second switch; The input terminal of the first multiplexer is connected to the reference voltage corresponding to the low-order portion of the multiple reference voltage sources, its control terminal receives the first control signal, and its output terminal is connected to the global voltage node through the first switch. The input of the second multiplexer is connected to the reference voltage corresponding to the high-order portion of the multiple reference voltage sources, its control terminal receives the second control signal, and its output terminal is connected to the global voltage node through the second switch; The in-memory computing unit includes a storage unit, a switching circuit, and a storage capacitor. The input terminal of the switching circuit is connected to the global voltage node, and the control terminal of the switching circuit is connected to the storage unit. The switching circuit is used to selectively apply the voltage of the global voltage node to the upper plate of the storage capacitor according to the weight data stored in the storage unit. The lower plate of the storage capacitor is used to couple out the accumulated voltage.

2. The differential input circuit based on multiple reference voltages according to claim 1, characterized in that, include: In the first stage of the operation, the first switch is closed, the second switch is open, and the decoder selects the first reference voltage to the global voltage node based on the low-order part of the multi-digit input. In the second stage of the operation, the first switch is opened and the second switch is closed. The decoder selects the second reference voltage to the global voltage node based on the high-order part of the multi-digit input. This causes the multi-digit input to be mapped in the analog domain to the difference between the first reference voltage and the second reference voltage, and output by the lower plate of the storage capacitor.

3. The differential input circuit based on multiple reference voltages according to claim 1, characterized in that, It also includes a sign bit control circuit, which includes a third switch and a fourth switch; The global voltage node includes a positive global voltage node and a negative global voltage node; Wherein, the first switch is connected between the output of the first multiplexer and the positive global voltage node, and the second switch is connected between the output of the second multiplexer and the positive global voltage node; The third switch is connected between the output of the first multiplexer and the negative global voltage node, and the fourth switch is connected between the output of the second multiplexer and the negative global voltage node. The sign bit control circuit further includes a sign bit storage unit and a selection switch. The selection switch is controlled by the sign bit stored in the sign bit storage unit and is used to connect one of the positive global voltage node and the negative global voltage node to the global voltage node.

4. The differential input circuit based on multiple reference voltages according to claim 3, characterized in that, In the first stage of the operation, the first and fourth switches are closed, and the second and third switches are open, so that the positive global voltage node and the negative global voltage node are respectively connected to the reference voltage selected by the low-order part and the high-order part; In the second stage of the operation, the first and fourth switches are opened, and the second and third switches are closed, so that the positive global voltage node and the negative global voltage node are respectively connected to the reference voltage selected by the high-order part and the low-order part, thereby forming a voltage difference change in opposite directions.

5. The differential input circuit based on multiple reference voltages according to claim 3, characterized in that, It also includes a charge multiplexing switch, which is connected between the positive global voltage node and the negative global voltage node; Between the first stage and the second stage, there is a charge reuse stage. In the charge reuse stage, the first switch, the second switch, the third switch, the fourth switch, and the lower plate grounding switch connected between the lower plate of the storage capacitor and the ground are all disconnected, and the charge reuse switch is closed, so that the positive global voltage node and the negative global voltage node can share charge to realize energy recovery.

6. A fully analog in-memory computing chip architecture, comprising the input circuit as described in any one of claims 1 to 5, characterized in that, include: Multiple computing unit groups, each computing unit group includes several in-memory computing units, the in-memory computing units are analog domain computing units, and the in-memory computing units are used to perform multiplication and accumulation operations on the input data; The channel mixing module is used to cross-recombine the outputs from each of the computing unit groups in the digital domain to generate the input activation values ​​for the next layer. The in-memory computing unit generates a multiply-accumulate result that does not depend on digital domain accumulation processing based on multi-bit input activation values ​​and multi-bit privileged values ​​within a single computing cycle, and integrates batch normalization and nonlinear activation operations in the analog domain to construct a fully analog computing link.

7. The all-analog domain in-memory computing chip architecture according to claim 6, characterized in that, Also includes: The non-overlapping control signal generation circuit is used to generate control signals for the operation timing of the control chip, including a delay chain and non-overlapping logic to control the signal transition order, so as to ensure the accuracy of capacitor charge distribution.

8. The all-analog domain in-memory computing chip architecture according to claim 6, characterized in that, The simulated domain fusion batch normalization and nonlinear activation operations are implemented through a simulated domain Non-MAC operation ADC, specifically including: The offset parameter unit is set in the same column as the in-memory calculation unit used to perform multiplication. During the multiplication-accumulation operation, the offset is directly accumulated into the multiplication-accumulation result. The dynamic range adjustable capacitive digital-to-analog converter array receives a digital signal representing a scaling factor and dynamically adjusts the quantization full-scale reference voltage of the ADC according to the scaling factor. The successive approximation logic module, which integrates ReLU logic, compares the voltage of the multiplication and accumulation result with the zero voltage before the ADC starts quantization. If it is less than the zero voltage, it directly outputs the quantization result representing the zero value and terminates subsequent quantization. If it is greater than the zero voltage, it starts the complete successive approximation quantization.

9. The all-analog domain in-memory computing chip architecture according to claim 6, characterized in that, The analog domain calculation unit also includes an analog domain weighting circuit; The analog domain weighting circuit is used to perform weighted summation of the analog quantities generated by the amplitude bits of the multi-privileged values ​​within the analog domain.

10. The all-analog domain in-memory computing chip architecture according to claim 9, characterized in that, The weighted summation of the analog quantities generated by the amplitude bits of the multiple privileged values ​​in the analog domain includes: after the multiplication and accumulation operation is completed in each column of the analog domain computing unit, the coupling lines representing different bit weights are physically segmented to form an equivalent capacitance proportional to the bit weight on each of the segmented coupling lines, and then the weighted analog voltage is obtained directly by closing the shared switch through charge redistribution.