Systems and methods for current-based analog near memory computing

WO2026163247A1PCT designated stage Publication Date: 2026-08-06INDIAN INSTITUTE OF TECHNOLOGY BOMBAY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
INDIAN INSTITUTE OF TECHNOLOGY BOMBAY
Filing Date
2026-01-30
Publication Date
2026-08-06

Smart Images

  • Figure IN2026050167_06082026_PF_FP_ABST
    Figure IN2026050167_06082026_PF_FP_ABST
Patent Text Reader

Abstract

An analog compute system (100) and a method of performing analog computation are disclosed for performing current-based computations in a scalable near-memory architecture with standard memory block compatibility while mitigating analog non-idealities including device mismatch, process variation, and operating-condition drift. The system (100) employs dynamic current mirrors operated using an iterative time-multiplexed dynamic current mirroring (TMDCM) scheme to achieve current equalization across compute arrays. The system (100) includes an input driver block configured to receive and condition input data, one or more standard memory blocks configured to store weight data, activation data, and compute data, and a multiplexing and demultiplexing block (106) configured to manage data exchange. A compute block (108) comprising arrays of dynamic current mirrors performs analog computations using the dynamic current mirror output currents equalized through the TMDCM process. The architecture supports operations including multiply-and-accumulate, and includes a readout block configured to output computation results.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR CURRENT-BASED ANALOG NEAR¬ MEMORY COMPUTING CROSS REFERENCE TO RELATED APPLICATIONThis application is based on and derives the benefit of Indian Provisional Application IN 202421064741, the contents of which are incorporated herein by reference.TECHNICAL FIELD

[0001] Various embodiments of the present disclosure relate to an analog compute system. More particularly, the present disclosure relates to an analog compute architecture with standard RAM compatibility that utilizes dynamic current mirrors with iterative time-multiplexed equalization to achieve accurate and stable analog computations.BACKGROUND

[0002] In recent years, analog in-memory computing (IMC) architectures have emerged as a promising alternative computing paradigm directed toward mitigating the long-standing von Neumann bottleneck. The von Neumann bottleneck arises from the physical and architectural separation of memory and processing units in conventional computing systems, which necessitate frequent and energy-intensive data transfers between processors and memory. This limitation becomes particularly pronounced in data-intensive workloads such as artificial intelligence (Al), machine learning, signal processing, and data analytics, where large volumes of data must be repeatedly accessed, moved, and processed.

[0003] As dataset sizes and model complexities continue to scale, the energy, latency, and bandwidth penalties associated with data movement increasingly dominate system performance. In-memory computing architectures seek to alleviate these constraints by enabling computation to be performed within memory arrays, thereby reducing data transfer overhead and improving energy efficiency.

[0004] Analog implementations of in-memory computing further promise significant gains in parallelism and compute density by leveraging the intrinsic physical properties of analog circuits, such as current summation and charge accumulation. However, despite these theoretical advantages, practical realization of analog IMC systems introduces substantial challenges relatedto precision, robustness, and scalability. Analog circuits are inherently sensitive to process, voltage, and temperature (PVT) variations, device mismatch, aging effects, and environmental noise. These non-idealities cause deviations in transistor characteristics, threshold voltages, and current mirroring behavior, resulting in reduced computational accuracy and limited repeatability across time and operating conditions.

[0005] In particular, current-based analog compute architectures rely on the assumption that nominally identical current sources or memory cells produce substantially equal currents under identical biasing conditions. In practice, however, mismatch between devices, voltage variations, temperature variations, parasitic interconnect effects, resistance-capacitance (RC) delays, and current-resistance (IR) drops introduce significant current non-uniformity across memory arrays or compute crossbars. Such non-uniformity directly degrades the accuracy of analog operations such as weighted summation, dot-product evaluation, and multiply-and-accumulate (MAC) computations, which form the computational backbone of neural network inference and other signal-processing workloads.

[0006] Existing approaches have attempted to address these challenges through techniques such as static current mirroring, one-time or periodic calibration, digital post-correction, redundancy-based averaging, or operation at elevated current levels for marginally reducing the impact of the analog imperfections. However, static mirroring techniques fail to compensate for device-to-device variations especially in large arrays, while calibration-based methods introduce substantial system overhead, require frequent refresh cycles, and are sensitive to aging and temperature variation. Digital correction and redundancy mechanisms increase area, power consumption, and system complexity, undermining the energy-efficiency benefits of analog computation. Moreover, many prior solutions depend on custom or modified memory cells, limiting compatibility with standard memory technologies and impeding scalability and manufacturability.

[0007] Additionally, conventional analog in-memory architectures often lack mechanisms to manage the disproportionate needs with respect to compute and memory of Al hardware systems. Typically, the die sizes of these Al systems are dominated by on-chip memory rather than compute while analog in-memory compute systems demand both to be equal. Also, these systemsrequire a custom memory unit cell and are thereby not compatible with optimized standard memory blocks.

[0008] Accordingly, there remains a need for an analog near-memory computing architecture that is compatible with standard memory blocks (standard RAM), and can provide the requisite memory to compute ratio while ensuring robust, repeatable, and uniform current behavior across large-scale compute arrays.OBJECTS

[0009] The present disclosure relates to an analog compute system configured to perform current-based computations in a scalable, near-memory architecture enabling standard RAM compatibility while mitigating the effects of analog imperfections such as device mismatch, process variation, and operating-condition drift. In various embodiments, the system employs dynamic current mirrors (DCMs) operated under an iterative time-multiplexed dynamic current mirroring (TMDCM) scheme to achieve robust and repeatable current equalization across compute arrays. By leveraging disjoint memory blocks and corresponding analog compute blocks, featuring recursive current equalization, the near memory architecture enables accurate analog computation using standard memory technologies.

[0010] In accordance with one or more embodiments, the analog compute system comprises an input driver block having buffering and wave shaping capability for receiving and conditioning input data, and at least one memory block configured to store one or more of weight data, activation data, and compute data. The memory block may be implemented using standard memory technologies and interfaces, thereby supporting near-memory computing.

[0011] The system further includes a multiplexing and demultiplexing block configured to manage buffering, routing, synchronizing and exchange of input data, weight data, activation data, and compute data among system components. The multiplexing and demultiplexing block enables efficient data sharing and time-multiplexed utilization of compute resources across the analog compute system.

[0012] The computational core of the system includes at least one compute block (CB) comprising one or more arrays of dynamic current mirrors. Each dynamic current mirror includes a bias generator and a capacitive element, which together enable dynamic storage and reproductionof reference currents using bias voltages. The CB performs iterative TMDCM operations, during which starting with at least one reference current, several accurate copies of the reference current are generated in each iteration of the TMDCM process. With sufficient number of iterations, the TMDCM process generates the required number of equalized currents which are made available for compute operation.

[0013] Using the equalized currents, the compute block performs analog computational operations based on at least one of the input data, weight data, activation data, and compute data. Such operations may include, by way of example, multiply-and-accumulate operations, matrixvector multiplications, additions, and subtractions. The analog computational results are provided to a readout block, which is configured to output one or more computation results, and may include analog-to-digital conversion circuitry for interfacing with downstream digital processing components.

[0014] In accordance with one or more embodiments, the present disclosure further provides a method of performing analog computation using an analog compute system. The method comprises providing input data to the analog compute system, storing one or more of weight data, activation data, and compute data in at least one memory block, and buffering, synchronizing, multiplexing and / or demultiplexing the data using at least one multiplexing and demultiplexing block. The method further comprises receiving the data at one or more compute blocks comprising arrays of dynamic current mirrors, and performing iterative time-multiplexed dynamic current mirroring to equalize currents by storing bias voltages on capacitive elements and reproducing currents during compute phases. Using the equalized currents, the compute block performs analog computational operations, and at least one computational result is output by a readout block. In various embodiments, the method supports recursive current equalization, multibit and signed analog computation, charge injection mitigation, auxiliary standby current sourcing, and tile-based scalable operation, while maintaining compatibility with standard memory technologies.BRIEF DESCRIPTION OF FIGURES

[0015] FIG. 1 is a diagram that illustrates an overall architecture of an analog compute system, in accordance with one or more embodiments of the present disclosure.

[0016] FIG. 2 is a diagram that illustrates an architecture configured to handle downtime in DCMs, in accordance with one or more embodiments of the present disclosure.

[0017] FIG. 3 is a diagram that illustrates a circuit-level realization of a Peripheral NearMemory Computing Dynamic Current Mirror (PNM DCM), including both a block representation and a detailed schematic, in accordance with an embodiment of the present disclosure.

[0018] FIG. 4 is a diagram that illustrates an architecture for generating binary weighted current sources within the analog compute system, in accordance with an embodiment of the present disclosure.

[0019] FIG. 5 is a diagram that illustrates a circuit realization of a signed data compatible Peripheral Near-Memory Computing Dynamic Current Mirroring Module (PNM_DCM), in accordance with one or more embodiments of the present disclosure.

[0020] FIG. 6 is a diagram that illustrates a realization of a PNM DCM implemented as a combination of modified RAM and temporary memory-based DCM block at an array and / or subarray level, in accordance with one or more embodiments of the present disclosure.

[0021] FIG. 7 is a diagram that illustrates an example architecture of the near memory analog compute system with multiplexing and demultiplexing block serving as the input driver block, in accordance with one or more embodiments of the present disclosure.

[0022] FIG. 8 is a diagram that illustrates an example of implicit interfacing architecture of the near memory analog compute system with digital input data or digital activation data, in accordance with one or more embodiments of the present disclosure.

[0023] FIG. 9 is a diagram that illustrates a flowchart 900 of a method for performing analog computation using an analog compute system, in accordance with an embodiment of the present disclosure.DETAILED DESCRIPTION

[0024] The present disclosure relates to an analog compute system configured to perform current-based computations in a scalable, near-memory architecture while mitigating the effects of analog imperfections such as device mismatch, process variation, voltage variation, temperature variation and operating-condition drift. In various embodiments, the system employs DCMs operated under an iterative TMDCM scheme to achieve robust and repeatable current equalizationacross compute arrays. By leveraging the equalized currents, the architecture enables accurate analog computation using standard memory technologies.

[0025] In accordance with one or more embodiments, the analog compute system comprises an input driver block having buffering and wave shaping capabilities for receiving and conditioning input data, and at least one memory block configured to store one or more of weight data, activation data, and compute data. The memory block may be implemented using standard memory technologies and interfaces, thereby supporting near-memory computing.

[0026] The system further includes a multiplexing and demultiplexing block configured to manage buffering, routing, synchronizing and exchange of input data, weight data, activation data, and compute data among system components. The multiplexing and demultiplexing block enables efficient data sharing and time-multiplexed utilization of compute resources across the analog compute system.

[0027] The computational core of the system includes at least one CB comprising one or more arrays of dynamic current mirrors. Each dynamic current mirror includes a bias generator and a capacitive element, which together enable dynamic storage and reproduction of applied reference currents through the bias voltage generated by the bias generator. The CB performs iterative TMDCM operations, during which starting with at least one reference current, several accurate copies of the reference current are generated in each iteration of the TMDCM process. With sufficient number of iterations, the TMDCM process generates the required number of equalized currents which are made available for compute operation. Through this process, substantially uniform currents are generated across the DCM arrays, independent of inherent analog imperfections.

[0028] Using the equalized currents, the CB performs analog computational operations based on at least one of the input data, weight data, activation data, and compute data. Such operations may include, by way of example, multiply-and-accumulate operations, matrix-vector multiplications, additions, and subtractions. The analog computational results are provided to a readout block, which is configured to output one or more computation results, and may include analog-to-digital conversion circuitry for interfacing with downstream digital processing components.

[0029] To facilitate understanding of the present disclosure, certain components and subsystems that are conventionally known in the art are briefly described herein. It will be appreciated that these components, such as input buffers, memory arrays, and multiplexing or demultiplexing circuits, are not themselves the focus of the inventive concept but are described to provide contextual clarity and to illustrate the interaction among various elements of the proposed analog compute system.

[0030] In one or more embodiments, an input driver block of the analog compute system comprises one or more input buffers, registers and / or shift registers which are analog, mixed-signal, or digital interface circuits configured to temporarily store and condition incoming data prior to its utilization by the compute system. The input driver block is designed to receive analog or digitized representations of input data, such as voltage or current signals corresponding to activation values, sensor readings, partial computational outputs.

[0031] In such embodiments, the input buffers forming part of the input driver block provide signal integrity management to ensure accurate and stable transfer of data to subsequent processing stages of the analog compute system. The input buffers may comprise one or more of drivers, latches, registers, shift registers, flip-flops, counters, pulse-width modulation waveform generators, digital-to-time converters, digital-to-analog converters, pre-amplifiers, low-pass filters, or sample-and-hold circuits, configured to stabilize, align, and maintain the input signal during one or more computation cycles.

[0032] In one or more embodiments, a memory block of the analog compute system comprises one or more memory arrays configured to retain one or more categories of data utilized by the system, including weight data, activation data, and compute data. The memory block and its associated memory arrays may be implemented using standard, commercially available memory technologies without requiring any hardware modification for integration within the analog compute architecture.

[0033] In some non-limiting embodiments, the memory arrays may be realized using Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Resistive Random Access Memory (RRAM), Flash memory, Ferroelectric RAM (FeRAM), Magnetoresistive RAM (MRAM), Ferroelectric Field Effect Transistor (FeFET) memory, High Bandwidth Memories (HBMs), or Phase Change Memory (PCM), spin-transfer torque MRAM(STT-MRAM), spin-orbit torque MRAM (SOT-MRAM), electrically erasable programmable read-only memory (EEPROM), or hybrid memory structures combining two or more memory technologies. The selection of memory technology may depend on desired trade-offs among access speed, data retention capability, area efficiency, power efficiency, and fabrication compatibility.

[0034] The memory arrays are configured to store data representations that serve as inputs, parameters, or intermediate results during computation. The stored data may include, by way of example, pre-trained model weights, dynamically updated activation values, or partial computation outcomes generated by the compute block.

[0035] In one or more embodiments, a multiplexing and demultiplexing block refers to a data management, routing and buffering module configured to control the selective transmission, buffering, multiplexing, and demultiplexing of signals or data streams among various functional components of the analog compute system. The multiplexing and demultiplexing block may comprise one or more electronic elements such as switches, multiplexers, demultiplexers, registers, shift registers and temporary data storage elements, thereby enabling efficient data flow and time-multiplexed utilization of shared compute and communication resources.

[0036] In operation, the multiplexing functionality of the multiplexing and demultiplexing block enables multiple data types, including input data, weight data, activation data, and compute data, to be transmitted over one or more shared communication interfaces in a time-multiplexed, event-driven, or scheduled manner, thereby reducing interconnect complexity and improving bandwidth utilization.

[0037] In one or more embodiments, the demultiplexing functionality of the multiplexing and demultiplexing block is configured to separate, decode, or route multiplexed data streams into individual output channels for subsequent processing. The demultiplexing functionality operates in coordination with the multiplexing functionality to maintain orderly, synchronized, and phase-aligned communication between the memory block, and compute block. Alternatively, the demultiplexing block may also buffer the data using registers and shift registers.

[0038] In such embodiments, the demultiplexing functionality distributes received data and control signals to designated destinations, including specific memory arrays, subarrays, compute block rows or columns, or individual DCMs within the compute architecture, based oncontrol signals, addressing information, or timing relationships associated with the received and transmitted data.

[0039] In one or more embodiments, a DCM refers to an analog circuit element configured to replicate or “mirror” an input reference current dynamically over time across one or more output branches, while mitigating analog imperfections such as threshold voltage variation and process mismatches that typically affect conventional static current mirrors.

[0040] Each DCM generally comprises a bias generator and a capacitive element. The bias generator is configured to generate a bias voltage corresponding to an applied reference current during a write or refresh phase, while the capacitive element is configured to store this bias voltage. Together they help reproduce the same reference current as an output current during a compute phase. This two-phase operation allows each DCM to effectively function as a temporary analog memory that retains the current characteristics even after the reference source is disconnected.

[0041] FIG. 1 is a diagram that illustrates an overall architecture of a near memory analog compute system 100, in accordance with one or more embodiments of the present disclosure, wherein the system is implemented as an Analog near memory compute architecture comprising a combination of standard RAM and dynamic current mirroring block at the array level to enable near-memory computing. The system, in one or more embodiments, leverages standard conventional RAM with standard input output blocks like readout blocks with sense amplifiers to enable near-memory computing and processing thereby providing a scalable platform for currentbased analog computation. As shown in FIG. 1 , the architecture includes an input driver block 102, a memory block 104, a multiplexing and demultiplexing block 106, a compute block 108, and a readout block 110, which collectively facilitate high-throughput analog operations while maintaining compatibility with conventional standard memory interfaces.

[0042] The system 100 is configured to perform current-based analog computations in a near-memory architecture with standard memory block compatibility while mitigating analog imperfections such as device mismatch, process variation, and operating-condition drift.Input Driver Block 102

[0043] The input driver block 102 is configured to receive and process a plurality of input data streams, which may comprise digital signals, analog signals, or a combination thereof,originating from external off-chip devices, sensors, or on-chip modules, and to prepare the received data for subsequent analog computation within the analog compute system 100.

[0044] The input driver block 102 includes one or more input buffers, which may comprise analog, mixed-signal, or digital circuit elements such as latches, registers, shift registers, analog mux, analog buffer, pre-amplifiers, low-pass filters, or pulse-width modulation(PWM) waveform generators, digital-to-time converters, digital-to-analog converters, the buffers being operable to temporarily store, condition, stabilize, and synchronize the incoming data to ensure accurate and reliable transfer to downstream functional components, including the multiplexing and demultiplexing block 106 and the compute block 108.

[0045] In certain embodiments, the input driver block includes R rows of PWM components, where each PWM component receives a digital input signal D(in)i through D(in)_R and converts it to an analog signal suitable for current-based computation.

[0046] The outputs of the PWM array are processed through one or more signal conditioning stages, which may include analog filtering, integration, amplification, or other conditioning circuits, to convert the pulse-width-modulated signals into stable, continuous analog current or voltage representations. These conditioned analog signals are then routed to the compute block CB 108, thereby ensuring high-fidelity signal transfer and temporal coherence across the analog compute system 100.Memory Block 104

[0047] The memory block 104 is configured to store one or more categories of data utilized by the analog compute system 100, including, but not limited to, weight data, activation data, and compute data. The memory block 104 serves as a central storage and retrieval subsystem interfacing with the input driver block 102, the multiplexing and demultiplexing block 106, and the compute block 108, thereby facilitating high-throughput and low-latency data movement across the various computational stages of the system.

[0048] The weight data may represent, without limitation, neural network model weights, synaptic weights, or scaling coefficients corresponding to parametric relationships used during computation. Activation data corresponds to intermediate signal values, such as neuron activations, node outputs, or other transient computational states generated during forward orbackward computational passes. Compute data comprises intermediate or resultant digital and analog values generated within the compute block 108, including partial sums, multiply -and-accumulate (MAC) results, or other current-based or voltage-based representations of mathematical operations.

[0049] The memory arrays of the memory block 104 may be implemented using standard commercially available memory technologies, including, but not limited to, Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Resistive Random Access Memory (RRAM), Flash memory, Ferroelectric RAM (FeRAM), Magnetoresistive RAM (MRAM), Ferroelectric Field-Effect Transistor (FeFET) memory, High Bandwidth Memory (HBM), or Phase Change Memory (PCM). Integration of these memory arrays within the analog compute system 100 does not require any hardware modification of the memory arrays, thereby preserving compatibility with conventional memory architectures while enabling near-memory analog computation.Multiplexing and Demultiplexing Block 106

[0050] The multiplexing and demultiplexing block 106 is configured to manage the selective transmission, routing, buffering, and coordination of input data, weight data, activation data, and compute data among the various functional components of the analog compute system 100. The multiplexing and demultiplexing block 106 ensures efficient utilization of shared data paths and supports time-multiplexed operation across multiple compute resources.

[0051] In various embodiments, the multiplexing and demultiplexing block 106 may comprise one or more circuit elements, including, without limitation, analog or digital switches, multiplexers, demultiplexers, registers, shift registers and temporary data storage elements. The circuit elements collectively facilitate sequential or parallel transmission of multiple data streams across the system while preserving signal fidelity.

[0052] The multiplexing and demultiplexing block 106 may further incorporate buffering and timing synchronization circuitry, which is configured to compensate for phase or clock mismatches between system components operating at different analog or digital frequencies. Such circuitry ensures data integrity, mitigates timing-induced errors, and maintains consistent temporal alignment of signals during high-speed analog computation within the system.Compute Block (CB) 108

[0053] In one or more embodiments, the CB 108 of the analog compute system 100 is configured to receive input data from the input driver block 102 and to receive at one or more of weight data, activation data, and compute data from the multiplexing and demultiplexing block 106. The CB 108 serves as the primary computational unit of the analog compute system, performing current-based operations utilizing arrays of DCMs.

[0054] In one or more embodiments, the CB 108 constitutes the computational core of the analog compute system 100 and is referred to as the compute block. The CB 108 is configured to perform analog computations with high precision by dynamically equalizing currents across the DCM arrays against analog imperfections such as threshold voltage variation, process mismatch, and channel-length modulation.

[0055] In one or more embodiments, the CB 108 performs various computational operations using the equalized currents based on the received data inputs. Such computational operations may include, but are not limited to, MAC operations, matrix-vector multiplications, additions, or subtractions, and other arithmetic or signal-processing functions, depending on the functional configuration of the analog compute system 100.

[0056] In one or more embodiments, the CB 108 may comprise one or more DCM arrays organized in a hierarchical and / or tiled structure, enabling parallel computation of multiple currentbased operations. Each DCM array employs a time-multiplexed dynamic current mirroring (TMDCM) scheme to achieve uniform current equalization across the entire compute block, thereby enhancing the accuracy of analog computations independent of process, voltage, and temperature (PVT) variations.

[0057] In one or more embodiments, the CB 108 performs iterative TMDCM to achieve current equalization across the at least one DCM array. Each iteration of the TMDCM operation involves sequential application of a reference current to one or more individual DCMs during designated write or refresh phases. During these phases, each DCM generates and stores a corresponding bias voltage that is capable of regenerating the applied reference current. The reference current used for TMDCM may be provided from an on-chip reference source integrated within the analog compute system 100 or alternatively from an external (off-chip) current reference source, depending on system design requirements and desired calibration accuracy.

[0058] In one or more embodiments, the successive iterations of the TMDCM process may be performed recursively across successive DCM arrays or subarrays to further enhance current uniformity across several DCM arrays or subarrays. In such configurations, the currents output from a first DCM array during a compute phase serves as input reference currents to a subsequent DCM arrays during its write or refresh phase. This recursive application continues across multiple arrays or subarrays, enabling seamless propagation of current equalization throughout the compute hierarchy independent of inherent analog imperfections such as device mismatch, process variation, or temperature drift.

[0059] In one or more embodiments, each DCM within the CB 108 comprises two principal components, namely a bias generator and a capacitive element. The bias generator is configured to generate a bias voltage corresponding to an applied reference current during a write or refresh phase. The bias voltage is then stored in the corresponding capacitive element. This bias voltage defines the mirroring behavior of the DCM, thereby establishing a reference condition for subsequent current reproduction. The capacitive element associated with each DCM is configured to store the generated bias voltage during the write phase. During the compute phase, the stored bias voltage is used to reproduce the reference current with high fidelity substantially independent of analog circuit imperfections or mismatch errors.

[0060] In one or more embodiments, each DCM is further configured with dual-port access comprising a first port configured to facilitate write or refresh operations and a second port configured to facilitate compute operations. The first port enables the writing or updating of bias voltages during the refresh phase, ensuring that current equalization remains consistent over time. The second port enables real-time compute operations using the previously stored bias conditions, thereby allowing simultaneous write / refresh and compute functionality. This dual-port architecture decouples computation from bias generation and update operations.

[0061] In one or more embodiments, at least one DCM array within the compute block 108 is configured to generate binary-weighted current sources through the iterative TMDCM process. During this process, TMDCM based current equalization produces current outputs that represent discrete binary weights corresponding to different bit-significance levels. This is achieved by aggregating one or more reproduced reference current outputs from a first DCM array during acompute phase, thereby generating binary-weighted input reference currents for subsequent DCM arrays during their write or refresh phase.

[0062] In one or more embodiments, combinations of these binary-weighted current sources enable multibit analog computation within each DCM array. This capability is achieved by aggregating one or more binary-weighted currents, which may represent multibit weight data, input data or compute data. This approach enables the analog compute system 100 to achieve increased throughput and improved energy efficiency and area efficiency without substantially increasing circuit complexity.

[0063] In one or more embodiments, each DCM further includes a bipolar current routing configuration configured to direct DCM output current toward either a positive bitline or a negative bitline based on a sign bit or polarity control signal. This bipolar configuration enables the representation and processing of signed data values, allowing the analog compute system 100 to perform computations involving both positive and negative operands. This capability improves area efficiency by enabling signed operations in typical DCM unit cells without substantial area overhead. Such a configuration is particularly advantageous for implementing neural network operations and other signal-processing tasks that require signed arithmetic, thereby enhancing the computational flexibility and applicability of the analog compute system 100.

[0064] In one or more embodiments, the analog compute system 100 implements a charge injection cancellation mechanism to mitigate transient errors introduced during switching events between write, refresh, and compute phases of the DCMs. The charge injection cancellation mechanism comprises multiple switching phases and complementary control signals arranged to balance charge transfer at sensitive nodes, thereby reducing offset errors, glitch-induced current spikes, and accumulated computation error during phase transitions.

[0065] In one or more embodiments, the analog compute system 100 comprises one or more auxiliary dynamic current mirrors configured to operate as a standby or substitute current source during write or refresh phases of one or more primary dynamic current mirrors within at least one CB 108. When a primary DCM is temporarily unavailable due to bias update or refresh operations, the auxiliary DCM supplies a substantially equivalent current, thereby maintaining continuity of current availability and preserving uninterrupted analog computation within the CB 108.Readout Block 110

[0066] The readout block 110 is configured to receive and output one or more computational results generated by the CB 108, thereby serving as the interface between the analog computational domain and subsequent digital or mixed-signal processing stages of the analog compute system 100.

[0067] In one or more embodiments, the readout block 110 comprises one or more analog-to-digital converters (ADCs) configured to convert analog current or voltage outputs from the CB 108 into corresponding digital representations with high precision. The readout block 110 may further include signal conditioning circuitry, such as analog or digital MUX, amplifiers, filters, sample-and-hold circuits, comparators or reference stabilization circuits, configured to enhance the fidelity, stability, and integrity of the analog signals prior to digitization.

[0068] The readout block 110 is operatively coupled to downstream system components, including but not limited to digital processors, memory modules, or control units, enabling subsequent processing, storage, or system-level decision-making based on the computational outputs of the analog compute system 100.Tile-Based Architecture

[0069] In one or more embodiments, the analog compute system 100 is implemented using a modular tile-based architecture comprising a plurality of tiles arranged in a one-dimensional or multi-dimensional spatial configuration. Each tile constitutes a semi-autonomous computational unit and includes at least one CB 108 comprising at least one array of DCMs. In addition, each tile further comprises at least one of a memory block 104, a multiplexing and demultiplexing block 106, an input driver block 102, and a readout block 110, thereby enabling localized analog computation within each tile.

[0070] The tile-based architecture enables scalable and modular system expansion by allowing additional tiles to be integrated into the analog compute system 100 without significant modification to existing tiles or core system components. The plurality of tiles may operate independently or cooperatively to perform parallelized analog computations across the multidimensional arrangement, thereby supporting increased computational throughput, flexible workload distribution, and scaled current-based analog processing.

[0071] FIG. 2 is a diagram that illustrates an architecture 200 configured to handle downtime in DCMs, in accordance with one or more embodiments of the present disclosure.

[0072] In one or more embodiments, FIG. 2 illustrates the architecture 200 of the compute block 108 described in FIG. 1, configured for mitigating downtime associated with DCMs. The architecture 200 comprises multiple hierarchical layers of DCMs, including DCM LI corresponding to a first layer and DCM L2 corresponding to a second layer, wherein each layer includes a plurality of DCM cells configured to perform analog computations.

[0073] Each DCM cell within the architecture 200 is associated with a downtime corresponding to its write or refresh phase, during which the DCM is temporarily unavailable to supply current for operation. To address this downtime, the architecture 200 includes at least one auxiliary DCM cell that functions as a standby current source. During the write phase of a primary DCM cell, the auxiliary DCM substitutes for the unavailable primary DCM, thereby ensuring continuous current availability for ongoing operations.

[0074] It is noted that when the auxiliary DCM itself undergoes a write or refresh phase, all other “P” primary DCM cells remain available as active current sources. Consequently, the effective number of DCM cells available at any given time remains constant, ensuring uninterrupted operations across all DCM layers.

[0075] This standby DCM architecture effectively mitigates downtime issues in peripheral DP AMs, including PNM DPAM configurations, thereby enhancing the reliability, precision, and operational continuity of the analog compute system 100 described in FIG. 1. By maintaining a constant set of active DCM current sources, the architecture 200 facilitates seamless analog computation and scalable performance across multiple layers of DCM arrays.

[0076] In one or more embodiments, the auxiliary DCM cell is configured with the same reference current, biasing and capacitive storage mechanisms as the primary DCMs, ensuring that the substituted current closely matches the output currents produced by the primary cells.

[0077] The architecture 200 supports hierarchical organization across multiple DCM layers, such as LI and L2, where each layer may include several groups of primary DCM cells. The standby auxiliary DCM can be dynamically assigned to any group within a layer experiencingwrite or refresh activity, enabling flexible and adaptive current management throughout the system.

[0078] The inclusion of auxiliary DCMs further facilitates the use of TMDCM across layers, allowing iterative current equalization to continue uninterrupted. This enables multi-layer operations and analog computations, including MAC operations or matrix-vector multiplications, are performed with consistent current levels between DCM arrays.

[0079] FIG. 3 is a diagram 300 that illustrates a circuit-level realization of a Peripheral Near-Memory Computing Dynamic Current Mirror Module (PNM_DCM), including both a block representation and a detailed schematic, in accordance with an embodiment of the present disclosure.

[0080] In one or more embodiments, the PNM DCM integrates a standard RAM array 302 with an array of temporary memory-based DCMs 304, wherein the DCMs serve as constant current sources robust against process, voltage, and temperature variations and the RAM provides at least one of a compute data, weight data and activation data.

[0081] The architecture enables near memory computing across the DCM array 304 by employing a time-multiplexing scheme. Specifically, columns of weight data stored in the RAM array 302 are sequentially multiplexed into the DCM array 304, which provide constant current sources robust against process, voltage, and temperature variations. This time-multiplexed interaction between the RAM array 302 and the DCM array 304 enables near-memory computing operations.

[0082] In one or more embodiments, the PNM DCM comprises control circuitry, including multiplexers, demultiplexers, and timing logic, to coordinate the sequential activation of RAM columns and the corresponding DCMs. The DCM array 304 may include auxiliary DCMs that compensate for downtime during write or refresh phases, thereby ensuring continuous availability of current sources during computation.

[0083] The PNM DCM schematic, as shown in FIG. 3, represents one possible circuitlevel implementation of the more general near memory analog compute architecture disclosed in FIG. 1 and FIG. 2. While the RAM array 302 provides storage of weight or data values, the peripheral DCM array 304 serves as constant current sources robust against process, voltage, andtemperature variations. This configuration supports near-memory processing by performing analog computations at or near the memory array without requiring invasive modification of the RAM cells.

[0084] The block representation in FIG. 3 illustrates the functional connectivity between the RAM array 302, DCM array 304, and control circuitry, while the schematic provides as an example, transistor-level detail of the DCM implementation along with interfacing between the RAM and DCM elements. This includes registers, buffers and MUX for interfacing, biasing circuits and capacitive storage elements for current equalization within DCM.

[0085] FIG. 4 is a diagram 400 that illustrates an architecture for generating binary weighted current sources within the analog compute system 100, in accordance with an embodiment of the present disclosure. In one or more embodiments, the architecture employs arrays of DCM 402 to generate discrete binary weighted currents suitable for multibit analog computation.

[0086] During this process, TMDCM based current equalization produces current outputs that represent discrete binary weights corresponding to different bit significance levels. This is achieved by aggregating several of the output currents from a first array of DCMs during a compute phase resulting in binary weighted input reference currents for subsequent arrays of DCMs during their write or refresh phase. For instance, a first layer DCM may generate an output current I, a second layer DCM may generate 2x1 by combining two first layer DCM outputs as the reference for the second layer DCM. Similarly, another second layer DCM may generate 4x1 by combining four first layer DCM outputs as the reference for this second layer DCM, and so forth, thereby creating a set of currents with binary-weighted magnitudes.

[0087] In one or more embodiments, control circuitry, including switches, multiplexers, and timing logic, orchestrates the sequential activation of DCMs to combine the binary weighted currents according to desired multibit analog values. The control circuitry ensures that, at any given time, the cumulative output current represents the summation of the active binary weighted currents, enabling precise analog representation of multibit data.

[0088] The binary weighted current sources may be integrated with memory arrays, such as the RAM array 302 described in FIG. 3, to support multibit near-memory computation operations. By employing dynamic current mirroring techniques, the architecture 400 mitigatesvariations due to PVT while maintaining accurate current scaling across multiple bit significance levels.

[0089] In certain embodiments, auxiliary DCMs may be included to compensate for downtime during write or refresh operations, thereby ensuring continuous availability of reference currents for binary weighting. The architecture is compatible with signed or unsigned data representations, enabling both positive and negative operands for signed analog computation.

[0090] The diagram 400 in FIG. 4 illustrates the functional connectivity between the at least one reference current source 404, the array of DCMs 402, and the control circuitry, highlighting the sequential activation of DCMs and aggregation of binary weighted currents. This architecture forms the basis for multibit computation within the analog compute system 100, supporting high-throughput and energy-efficient processing.

[0091] FIG. 5 is a diagram 500 that illustrates a circuit realization of a signed data compatible Peripheral Near-Memory Computing Dynamic Current Mirroring Module (PNM DCM) 502, in accordance with one or more embodiments of the present disclosure. The PNM DCM 502 is configured to perform current-based signed analog computations using DCMs positioned in the periphery of a memory array, thereby enabling near-memory signed computation while preserving the integrity of the underlying memory core.

[0092] In one or more embodiments, the analog compute system 100, as described with reference to FIG. 1, supports signed weight processing through the PNM_DCM 502 by employing a dual-bitline accumulation technique. In this approach, each computational operation generates current components corresponding to positive and negative data values, which are accumulated on separate bitlines. This separated current representation enables accurate signed arithmetic operations while maintaining robustness against common-mode noise and analog non-idealities.

[0093] The PNM DCM 502 illustrated in FIG. 5 demonstrates a near-memory computing configuration in which temporary memory-based DCMs are arranged adjacent to, but physically distinct from, the RAM array. Weight data stored within the RAM array is selectively accessed and time-multiplexed into the peripheral DCM array, where dynamic current mirroring is employed to generate equalized current outputs for analog computation.

[0094] In one or more embodiments, signed data compatibility within the PNM DCM 502 is achieved by configuring each DCM with a bipolar current steering mechanism. The bipolar configuration directs reproduced currents toward either a positive bitline or a negative bitline based on a sign control signal derived from the product of sign bits of at least two of weight data, activation data, input data and compute data. The magnitude of the current is determined by the absolute value of the product of at least two of weight data, activation data, input data and compute data, while the polarity is determined by the sign control signal, thereby enabling efficient implementation of signed MAC operations in the analog domain.

[0095] Control circuitry associated with the PNM_DCM 502 includes switches, multiplexers, and timing logic configured to coordinate write, refresh, and compute phases of the DCMs, as well as the selective routing of currents to the appropriate bitlines. In certain embodiments, auxiliary or standby DCMs may be incorporated to compensate for downtime during write or refresh phases of primary DCMs, ensuring continuous current availability and uninterrupted signed data computation.

[0096] The circuit realization shown in FIG. 5 further illustrates how the combination of peripheral DCM arrays, dual bitline accumulation, and time-multiplexed access to memory-stored weights enables scalable signed data analog computation with reduced sensitivity to PVT variations. Accordingly, the PNM_DCM 502 provides a robust and energy-efficient near-memory computing solution suitable for neural network inference, signal processing, and other currentbased analog computing applications. Further, in a similar way, the circuit realization shown in FIG. 5 can also be utilized to perform signed MAC computations wherein at least two of weight data, input data, compute data and activation data are signed.

[0097] FIG. 6 is a diagram that illustrates a realization of a PNM DCM implemented as a combination of modified RAM and temporary memory-based DCM block at an array and / or subarray level, thereby enabling near-memory computing and processing. In one or more embodiments, the architecture illustrated in FIG. 6 represents a PNM DCM configuration 600 relative to the architectures described with reference to FIGS. 1-5.

[0098] In one or more embodiments, FIG. 6 presents alternative realizations of a PNM DCM-based near-memory computing system architecture, wherein computational functionality is performed in the peripheral compute block. The architecture enables analogcomputation to be performed at or near the memory array level, thereby reducing data movement overhead between memory and compute blocks and improving overall energy efficiency and throughput.

[0099] In the illustrated embodiments, the structurally modified RAM arrays 602 are used to interface directly with analog compute elements which are DCMs or DCM-based subcircuits. The direct interfacing eliminates the typical memory readout blocks. The DCMs are arranged at the array or subarray level, allowing groups of memory cells or columns to share corresponding DCM resources for current-based computation. Such an arrangement limits the loading during data movement to the compute block from the RAM array thereby improving speed, throughput and energy efficiency.

[0100] In one or more embodiments, weight data stored within the modified RAM array is selectively demultiplexed out at the column or subarray level and provided as input to the associated temporary memory-based DCMs. The DCMs generate equalized analog currents corresponding to the products of applied weight data and activation data or input data, compensating for PVT variations through iterative dynamic current mirroring techniques as previously described. The resulting currents may be accumulated, combined, or further processed locally within the subarray before being routed to downstream compute or readout circuitry.

[0101] The array-level or subarray-level integration of DCM block and modified RAM enables fine-grained near-memory computing, wherein analog operations such as MAC, vectormatrix multiplication, or weighted summation are performed in close physical proximity to the data storage elements.

[0102] In one or more embodiments, control logic and multiplexing circuitry are provided to coordinate the timing of read, write, refresh, and compute phases across the modified RAM and temporary memory-based DCM elements. Such control circuitry may operate at the row, column, or subarray level, enabling flexible scheduling of computation and ensuring that current equalization and data integrity are maintained throughout the near-memory computing process.

[0103] FIG. 7 is a diagram that illustrates an example architecture of the near-memory analog compute system with a multiplexing and demultiplexing block serving as the input driver block, in accordance with one or more embodiments of the present disclosure. In one or more embodiments, the architecture illustrated in FIG. 7 represents a PNM DCM configuration 700relative to the architectures described with reference to FIGS. 1-5. As illustrated, the PNM DCM configuration 700 includes a standard RAM array 1 (702), a multiplexing and demultiplexing block (704), an array of dynamic current mirrors (706), a multiplexing and demultiplexing block (708), a standard RAM array 2 (710), and a readout block (712).

[0104] FIG. 7 illustrates that when the analog compute block (CB) supports direct digital input or activation data, the multiplexing and demultiplexing block can serve as the implicit input driver block. The standard RAM array 1 provides at least one of input data, activation data, and compute data. In this example, the data is expected in digital format within the CB rather than in the PWM waveform format illustrated in FIG. 1 and FIG. 6. Thus, the CB generally supports multiple input, activation, and compute data formats ranging from PWM waveforms to direct digital signals. Furthermore, the input driver block can be realized using a multiplexing and demultiplexing block, particularly when the corresponding data is digital.

[0105] FIG. 8 is a diagram that illustrates an example of implicit interfacing architecture of the near-memory analog compute system with digital input data or digital activation data, in accordance with one or more embodiments of the present disclosure. In one or more embodiments, the architecture illustrated in FIG. 8 represents a PNM DCM configuration 800 relative to the architectures described with reference to FIGS. 1-5. As illustrated, the PNM DCM configuration 800 includes a standard RAM array 1 (802), an array of dynamic current mirrors (804), a standard RAM array 2 (806), and a readout block (808).

[0106] FIG. 8 illustrates that when the analog compute block (CB) supports direct digital input or activation data, the standard input-output block of a standard RAM can serve as the implicit input driver block. The standard RAM array 1 (802) provides at least one of input data, activation data, and compute data. In this example, the data is expected in digital format within the CB rather than in the PWM waveform format illustrated in FIG. 1 and FIG. 6.

[0107] Furthermore, FIG. 8 illustrates that the multiplexing and demultiplexing block shown in FIG. 1 can be implicitly realized using the latches and registers within the input-output block of a standard RAM. For simplicity, FIG. 8 omits illustration of the associated multiplexer in the RAM input-output block. However, each readout block within the input-output block of a standard RAM typically includes multiplexers interfacing the bitlines to sense amplifiers. These multiplexers, along with the latches and registers, can perform the requisite interfacingfunctionality between the memory block and compute block, eliminating the need for an explicit multiplexing and demultiplexing block (as shown in FIG. 1).

[0108] FIG. 9 is a diagram that illustrates a flowchart 900 of a method for performing analog computation using an analog compute system, in accordance with an embodiment of the present disclosure.

[0109] The method illustrated in FIG. 9 may be performed by the analog compute system 100 described with reference to FIGS. 1-5 and utilizes the input driver block 102, memory block 104, multiplexing and demultiplexing block 106, compute block 108 comprising arrays of dynamic current mirrors, and readout block 110.

[0110] It will be appreciated that the steps described herein may be performed sequentially, concurrently, or in a partially overlapping manner, depending on system configuration, timing constraints, and operational requirements.

[0111] At 902, input data is provided to the analog compute system by the input driver block 102. In one or more embodiments, the input data comprises analog signals, digital signals, or mixed-signal representations corresponding to activation values, sensor inputs, or pre-processed partial computational outputs from previous compute operations.

[0112] The input driver block 102 includes one or more input buffers, registers, shift registers, wave shaping circuits or signal-conditioning circuits configured to temporarily store, stabilize, and condition the input data prior to its use by downstream components of the analog compute system 100.

[0113] At 904, at least one of weight data, activation data, and compute data is stored by at least one memory block 104 of the analog compute system 100.

[0114] In one or more embodiments, the memory block 104 comprises standard memory arrays implemented using commercially available memory technologies, including but not limited to SRAM, DRAM, RRAM, Flash memory, FeRAM, MRAM, FeFET memory, HBM, or PCM, without requiring any modification for integration with the analog compute system.

[0115] At 906, at least one multiplexing and demultiplexing block 106 buffers, synchronizes, multiplexes and / or demultiplexes at least one of the input data, weight data, activation data, and compute data.

[0116] The multiplexing and demultiplexing block 106 comprises one or more switches, multiplexers, demultiplexers, registers, shift registers, and temporary data storage elements configured to manage data routing and temporal alignment between the memory block 104, and compute block 108.

[0117] At 908, at least one compute block 108 receives the input data from the input driver block 102 and receives at least one of the weight data, activation data, and compute data from the multiplexing and demultiplexing block 106.

[0118] The compute block 108 comprises one or more arrays of dynamic current mirrors, each dynamic current mirror including a bias generator and a capacitive element, as previously described with reference to FIGS. 1-6.

[0119] At 910, the compute block 108 performs iterative time-multiplexed dynamic current mirroring (TMDCM) using the array of dynamic current mirrors to equalize currents within the compute block.

[0120] In one or more embodiments, TMDCM comprises sequentially applying a reference current to individual dynamic current mirrors or groups of dynamic current mirrors across multiple time intervals during write or refresh phases.

[0121] At 912, during write or refresh phases, bias voltages generated by the bias generators of the dynamic current mirrors are stored on the corresponding capacitive elements.

[0122] The bias voltages are generated in response to the applied reference current during write or refresh phase and the stored bias voltages define the current reproduction characteristics of the dynamic current mirrors.

[0123] At 914, during compute phases, the dynamic current mirrors reproduce currents based on the stored bias voltages.

[0124] The reproduced currents are equalized across the array so as to reduce sensitivity to transistor mismatch, process variations, voltage variations, and temperature variations, thereby enabling robust and repeatable analog computation.

[0125] At 916, the compute block 108 performs analog computational operations using the equalized currents based on at least one of the input data, weight data, activation data, and compute data.

[0126] In one or more embodiments, the analog computational operations include multiply-and-accumulate operations, matrix-vector multiplication, addition, and subtraction. In certain embodiments, binary-weighted current sources and bipolar current routing are employed to enable multibit and signed analog computation.

[0127] At 918, at least one computational result generated by the compute block 108 is output by the readout block 110.

[0128] In one or more embodiments, the readout block 110 comprises one or more analog-to-digital converters configured to convert analog computational results into corresponding digital outputs for further processing or storage.

[0129] In one or more embodiments, the method further comprises mitigating charge injection errors by performing at least two distinct switching phases during transitions between write, refresh, and compute phases, wherein the switching phases are configured to suppress charge redistribution effects at sensitive circuit nodes.

[0130] In one or more embodiments, the method further comprises supplying a standby reference current from at least one auxiliary dynamic current mirror during write or refresh phases of one or more primary dynamic current mirrors, thereby maintaining continuity of current availability during computation.

[0131] In one or more embodiments, the method is performed using a tile-based architecture comprising a plurality of tiles arranged in a multi-dimensional spatial configuration, each tile including at least one compute block comprising an array of dynamic current mirrors and local memory, multiplexing and demultiplexing, input driving, and readout functionality.

[0132] The analog near-memory compute system disclosed herein provides significant technical advantages over conventional digital compute architectures and existing in-memory analog computing solutions. By relocating computation to the periphery of standard memory arrays, the system enables near-memory processing that combines the low-latency and energy efficiency benefits of proximity computation with the structural simplicity, area efficiency and reliability of unmodified memory blocks. This architectural approach allows analog computation to occur adjacent to, rather than within, memory cells, thereby avoiding cell-level redesign and retention issues common in in-memory compute schemes.

[0133] One of the primary technical advantages lies in the system’s ability to achieve compatibility with standard memory technologies while maintaining uniform current behavior across an extended compute array, independent of device-level analog imperfections or process, voltage, and temperature (PVT) variations. Through the implementation of time-multiplexed dynamic current mirroring (TMDCM), the system continuously equalizes currents among multiple dynamic current mirrors (DCM). Each DCM reproduces a stable reference current from stored bias voltages, enabling highly consistent analog operations across the array and minimizing mismatch-induced deviations.

[0134] The architecture achieves improved analog linearity and computational precision by employing bias voltage storage and current regeneration using capacitive elements within each DCM. This design minimizes RC delay effects and mitigates the impact of current-resistance (IR) drops, thereby enhancing current stability during multiply-and-accumulate (MAC) or matrixvector multiplication operations. As a result, the system delivers higher accuracy in analog computations without requiring complex calibration.

[0135] Another key advantage arises from the hierarchical and recursive configuration of the TMDCM framework. Multiple layers of DCMs, organized in a cascaded arrangement, iteratively perform current equalization across successive stages. This recursive behavior allows the system to scale seamlessly for larger compute arrays while preserving current uniformity and reducing cumulative analog drift. The hierarchical organization thus supports modular scalability for advanced near-memory workloads.

[0136] The present disclosure further introduces auxiliary dynamic current mirrors (AUX-DCMs) that act as standby current sources during write and refresh phases. These auxiliary units maintain bias stability and continuity of current flow while the primary DCMs undergo bias programming. This approach eliminates computation downtime and ensures uninterrupted operation of the analog compute system, even during periodic refresh sequences.

[0137] Another notable advantage is the system’s support for multibit and signed analog computation. By generating binary-weighted current sources through iterative TMDCM and incorporating bipolar configurations within the DCMs, the system efficiently represents both magnitude and polarity of analog signals. This enables execution of complex arithmetic operationssuch as signed multiplication, subtraction, and accumulation features critical for implementing analog neural networks and signal-processing pipelines.

[0138] From a hardware integration perspective, the analog compute system provides seamless compatibility with standard memory technologies such as SRAM, DRAM, RRAM, and Flash. The near-memory placement of the compute block eliminates the need for structural modification to the existing memory cells or array topology. This compatibility significantly simplifies manufacturing and integration, allowing near-memory compute capabilities to be embedded within existing optimized memory macros.

[0139] The system also enhances energy and area efficiency by minimizing data transfer between compute and memory components. Since computation occurs near the memory interface, the design effectively mitigates the von Neumann bottleneck and reduces overall data movement energy. This yields lower power consumption and smaller area overhead compared to digital or fully in-memory analog systems, making the architecture suitable for high-throughput, energy-constrained applications.

[0140] Further, the system employs a charge injection cancellation mechanism synchronized with write and compute control signals. This technique mitigates transient charge disturbances during switching transitions, ensuring accurate bias retention and minimizing switching-induced computational errors. Consequently, the compute block maintains stable analog output characteristics across repeated operational cycles.

[0141] Finally, the combination of these innovations including TMDCM-based current equalization, hierarchical layering, auxiliary mirroring, and near-memory deployment enables a robust, scalable, and energy-efficient analog computing platform compatible with standard memory blocks. The system offers improved computational fidelity, reduced sensitivity to analog imperfections, and enhanced operational reliability, all while remaining compatible with existing memory infrastructure. These advantages collectively establish the disclosed near-memory analog compute architecture as a technically superior and practically implementable alternative to conventional in-memory and digital compute approaches.

Claims

STATEMENT OF CLAIMS1. An analog compute system (100), the system (100) comprising:an input driver block (102) comprising buffering and wave shaping capabilities and configured to provide input data to the system (100);at least one memory block (104) configured to store at least one of a weight data, an activation data and a compute data;at least one multiplexing and demultiplexing block (106) configured to buffer, synchronize, demultiplex and / or multiplex at least one of the input data, the weight data, the activation data and the compute data;at least one compute block (CB) (108) comprising at least one array of dynamic current mirrors (DCM), wherein each dynamic current mirror comprises a bias generator and a capacitive element, wherein each CB (108) is configured to:receive the input data from the input driver block (102);receive at least one of the weight data, the activation data and the compute data from the multiplexing and demultiplexing block (106);perform iterative time multiplexed dynamic current mirroring (TMDCM) to equalize currents in the CB (108), wherein each CB achieves current equalization by storing bias voltages from the bias generator in the capacitive elements of the dynamic current mirrors, thereby equalizing currents substantially independent of analog imperfections;perform computational operations using the equalized currents based on at least one of the input data, the weight data, the activation data and the compute data; andat least one readout block (110) configured to output at least one computational result from the at least one CB (108).

2. The system (100) of claim 1, wherein a memory block of the at least one memory block (104) is a standard memory block requiring no modification for integration with the analog compute system (100).

3. The system (100) of claim 2, wherein the standard memory block is selected from the group consisting of Static Random Access Memory (SRAM), Dynamic Random Access Memory(DRAM), Resistive Random Access Memory (RRAM), Flash memory, Ferroelectric RAM, Magnetoresistive RAM (MRAM), FeFET memory, High Bandwidth Memories (HBMs) and Phase Change Memory (PCM).

4. The system (100) of claim 1, wherein each dynamic current mirror further comprises:a bias generator configured to generate a bias voltage using an applied reference current during a write phase; anda capacitive element configured to store the bias voltage and reproduce the reference current during a compute phase.

5. The system (100) of claim 1, wherein each DCM comprises:first port configured for write or refresh operations; anda second port configured for compute operations, wherein the first port and the second port enable simultaneous write or refresh operations and the compute operations.

6. The system (100) of claim 1, wherein TMDCM comprises:applying at least one reference current to an array of dynamic current mirrors sequentially over time during write or refresh phases, to generate and store corresponding bias voltages, wherein the at least one reference current is provided either on-chip of the analog compute system or off-chip of the analog compute system (100); andreproducing the reference current from the corresponding bias voltages during compute phases through the dynamic current mirrors, thereby equalizing output currents of the dynamic current mirrors substantially independent of analog imperfections.

7. The system (100) of claim 6, wherein the iterative TMDCM is performed recursively by applying TMDCM across successive arrays of DCMs, wherein the currents output from a first array of DCMs during a compute phase serve as input reference currents to a subsequent array of DCMs during their write or refresh phase, and this process repeats recursively across successive arrays of DCMs, thereby enabling expansion of the equalized currents across DCM arrays through recursive application of the iterative TMDCM equalization process.

8. The system (100) of claim 7, wherein at least one DCM array generates binary-weighted current sources through the iterative TMDCM.

9. The system (100) of claim 8, wherein combinations of the binary-weighted current sources enable representation of multibit weight data, multibit input data, multibit activation data and multibit compute data within each array of DCMs.

10. The system (100) of claim 1, wherein each DCM includes a bipolar configuration configured to direct current to one of a positive bitline or a negative bitline based on a sign bit polarity enabling signed data computations.

11. The system (100) of claim 1 wherein the system (100) implements a charge injection cancellation mechanism comprising multiple switching phases arranged to mitigate computing errors during transitions between write and compute phases.

12. The system (100) of claim 1, further comprising at least one auxiliary dynamic current mirror configured to serve as a standby current source during write phases of primary dynamic current mirrors within at least one CB.

13. The system (100) of claim 1, wherein the read-out block (110) comprises at least one analog-to-digital converter configured to convert analog computational results to digital outputs.

14. The system (100) of claim 1, wherein the computational operations are selected from the group consisting of multiply and accumulate operations, matrix-vector multiplication, addition, and subtraction.

15. The system (100) of claim 1, wherein a multiplexing and demultiplexing block (106) can comprise one or more elements selected from the group consisting of a switch, a multiplexer, a demultiplexer, a register, a shift register and a data storage element.

16. The system (100) of claim 1, wherein the analog compute system (100) is configured with a tile-based architecture comprising a plurality of tiles arranged in a multi-dimensional spatial configuration, wherein each tile includes at least one compute block (108) comprising at least one array of DCMs, and further comprises at least one of a memory block (104) , a multiplexing and demultiplexing block (106), an input driver block (102),and a readout block (110), enabling modular system expansion and scaled computational operations.

17. A method of performing analog computation using an analog compute system, the method comprising:providing, by an input driver block, input data to the analog compute system;storing, by at least one memory block, at least one of weight data, activation data, and compute data;buffering, synchronizing, multiplexing and / or demultiplexing, by at least one multiplexing and demultiplexing block, at least one of the input data, the weight data, the activation data, and the compute data;receiving, by at least one compute block comprising an array of dynamic current mirrors, the input data from the input driver block and at least one of the weight data, the activation data, and the compute data from the multiplexing and demultiplexing block;performing, by the at least one compute block, iterative time-multiplexed dynamic current mirroring (TMDCM) using the array of dynamic current mirrors to equalize currents within the compute block, wherein each dynamic current mirror comprises a bias generator and a capacitive element;storing, during write or refresh phases, bias voltages generated by the bias generators on the capacitive elements of the dynamic current mirrors;reproducing, during compute phases, currents based on the stored bias voltages to equalize output currents of the dynamic current mirrors substantially independent of analog imperfections;performing, by the at least one compute block, analog computational operations using the equalized currents based on at least one of the input data, the weight data, the activation data, and the compute data; andoutputting, by at least one readout block, at least one computational result from the at least one compute block.

18. The method of claim 17, wherein storing the bias voltages comprises:generating, by the bias generator of each dynamic current mirror, a bias voltage using an applied reference current during a write or refresh phase; andstoring the generated bias voltage on the capacitive element for reproduction during a compute phase.

19. The method of claim 17, wherein performing the iterative TMDCM comprises:operating each dynamic current mirror using a first port configured for write or refresh operations and a second port configured for compute operations, such that write or refresh operations on the first port and compute operations on the second port proceed with temporal overlap without mutual interference.

20. The method of claim 17, wherein performing the TMDCM comprises:sequentially applying at least one reference current to the array of dynamic current mirrors across multiple time intervals during write or refresh phases to generate and store corresponding bias voltages, the at least one reference current being provided on-chip or off- chip of the analog compute system; andreproducing the reference current from the stored bias voltages during compute phases through dynamic current mirrors to equalize output currents of the dynamic current mirrors substantially independent of analog imperfections.

21. The method of claim 20, wherein performing the iterative TMDCM further comprises:recursively applying TMDCM across successive arrays of dynamic current mirrors arranged in cascade, such that the currents output from a first array during a compute phase serve as input reference currents to a subsequent array during a write or refresh phase; and repeating the cascade process across multiple successive arrays, thereby enabling hierarchical expansion of equalized currents in the DCM arrays.

22. The method of claim 21, wherein performing the recursive TMDCM results in generation of binary-weighted current sources using the DCM arrays.

23. The method of claim 22, wherein performing the analog computational operations comprises combining the binary-weighted current sources across at least one array of dynamic current mirrors to represent multibit weight data, multibit input data, and multibit activation data.

24. The method of claim 17, wherein performing the analog computational operations further comprises directing current from each dynamic current mirror to one of a positive bitline or a negative bitline based on a sign bit polarity, thereby enabling computational operations on signed data.

25. The method of claim 17, further comprising mitigating charge injection errors by performing at least two distinct switching phases during transitions between write phases and compute phases, wherein the switching phases are configured to suppress charge redistribution effects at the transition boundary.

26. The method of claim 17, further comprising supplying, during write phases of primary dynamic current mirrors within the at least one compute block, at least one standby current from at least one auxiliary dynamic current mirror, thereby maintaining current source continuity during primary mirror configuration.

27. The method of claim 17, wherein outputting the at least one computational result comprises converting, by at least one analog-to-digital converter of the readout block, the analog computational results to digital outputs.

28. The method of claim 17, wherein performing the analog computation comprises operating the analog compute system using a tile-based architecture, the method further comprising: arranging a plurality of tiles in a multi-dimensional spatial configuration, each tile including at least one compute block comprising at least one array of dynamic current mirrors; and operating the tiles in parallel with local memory, multiplexing and demultiplexing, input driver, and readout functionality within each tile, thereby enabling modular system expansion and scaled computational operations.