3D stacked memristor array design method for high-energy-efficiency AI calculation

Through a layered heterogeneous 3D stacking architecture and dynamic thermal-electric collaborative regulation algorithm, the energy efficiency and reliability bottleneck of memristor arrays in AI computing is solved, and high-energy-efficient AI computing capabilities are achieved, and device stability and computing efficiency are improved.

CN120449799AInactive Publication Date: 2025-08-08ZHONGKE YIXIN MICROELECTRONICS (SUZHOU) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510561172.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing memristor arrays have energy efficiency and reliability bottlenecks in AI computing. The physical separation of the memory unit and the computing unit leads to high power consumption. The sparse calculation and the sparse characteristics of the AI model are mismatched, the signal transmission integrity is insufficient, the thermal-electric coupling effect is not compensated in real time, the control command response delay, and signal distortion caused by vertical interconnection holes.

Method used

The hierarchical heterogeneous 3D stacking architecture is adopted, combined with dynamic thermal-electric collaborative regulation algorithm and sparse-aware adaptive gating, and the vertical interconnection between the storage layer and the computing layer is realized through hybrid bonding technology, hardware acceleration units are deployed for real-time encoding and transmission, a coupling relationship model between temperature field and resistive drift is established, voltage pulse parameters are dynamically optimized, and adaptive gating with sparse threshold criteria and energy efficiency constraints are realized.

Benefits of technology

It solves the additional energy consumption and delay problems caused by cross-layer data handling, improves device stability and life, reduces the redundant power consumption of invalid multiplication and addition operations, ensures fast mode switching capabilities, optimizes data near processing and transmission losses, and improves computing energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120449799A_ABST
    Figure CN120449799A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of memristor array design, and discloses a 3D stacked memristor array design method for high-energy-efficiency AI calculation, and the method comprises the following steps: S1, constructing a layered heterogeneous 3D stacked architecture which comprises the physical isolation and dynamic reconstruction of a storage layer, a calculation layer and a control layer; s2, establishing a coupling relation model of a temperature field and memristor resistance state drift based on a dynamic thermal-electric cooperative regulation and control algorithm, and optimizing voltage pulse parameters in real time; s3, dynamically distributing mixed precision pulses according to the sparse characteristic of the AI model through sparse perception adaptive gating; and S4, deploying a hardware acceleration unit in a control layer, and realizing thermal-electric parameter calculation and real-time coding transmission of a regulation and control instruction. By adopting the technical scheme of layered heterogeneous 3D stacking architecture and hybrid bonding vertical interconnection, the effect of data near processing and transmission loss collaborative optimization is achieved, and the problems of extra energy consumption and delay caused by cross-layer data handling are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of memristor array design, and specifically to a design method for a 3D stacked memristor array for energy-efficient AI computing. Background Art

[0002] In recent years, memristor-based integrated memory and compute architectures have shown significant potential for AI acceleration, but energy efficiency and reliability bottlenecks have hindered practical deployment. Early planar memristor arrays were limited by the data transfer path of the von Neumann architecture, and the physical separation of memory and compute cells resulted in up to 60% inefficient power consumption. As process nodes shrink, the localized heat accumulation caused by dense integration intensifies, making static thermal management strategies unable to suppress the exponential accumulation of resistance drift errors.

[0003] Initial attempts by the industry to mitigate temperature rise by adding heat sinks or reducing operating frequencies sacrificed computational density and throughput. Some solutions introduced sparse computing acceleration units, but the mismatch between the fixed bit width and the dynamic sparsity of AI models still resulted in over 25% ineffective multiplication and addition operations. While 3D stacking technology theoretically breaks through the memory barrier, impedance mismatch in traditional hybrid bonding processes results in signal reflectivity exceeding 10%, degrading the integrity of high-frequency pulse transmission.

[0004] In existing technology systems, thermal-electric coupling effects have long been simplified into linear superposition models, ignoring the nonlinear modulation of resistance drift by temperature gradients. Multi-physics collaborative control algorithms often rely on offline calibration parameters, making it difficult to compensate for thermal shock caused by dynamic workloads in real time. Instruction transmission links generally utilize a serial bus architecture, resulting in timing conflicts between the cache latency of centralized processing units and the parallel computing characteristics of memristor arrays. Control instruction response times exceed the computing cycle by more than 50%.

[0005] At the hardware architecture level, the decoupling of the sparse sensing module from the computing unit results in excessive mask generation latency, creating an inverse relationship between sparse detection accuracy and real-time performance. Parasitic parameters in vertical interconnect vias cause signal waveform distortion, and the fixed LC values of traditional impedance matching networks cannot adapt to temperature-related material property drift. At the process level, existing bonding technologies have a yield rate of less than 90%, and functional unit bypass caused by via failures further reduces effective computing density. Summary of the Invention

[0006] In response to the shortcomings of the existing technology, the present invention provides a design method for a 3D stacked memristor array for high-energy-efficiency AI computing, which solves the problems of the storage wall bottleneck of the memristor architecture, the degradation of computing accuracy caused by thermal runaway, the low utilization of sparse resources, and the insufficient signal integrity between layers.

[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: a design method for a 3D stacked memristor array for energy-efficient AI computing, comprising the following steps:

[0008] S1. Build a layered heterogeneous 3D stacking architecture, including physical isolation and dynamic reconstruction of the storage, computing, and control layers;

[0009] S2. Based on the dynamic thermal-electric coordinated control algorithm, a coupling relationship model between the temperature field and the resistance drift of the memristor is established to optimize the voltage pulse parameters in real time;

[0010] S3, through sparse-aware adaptive gating, dynamically allocates mixed-precision pulses based on the sparse characteristics of the AI model;

[0011] S4. Deploy hardware acceleration units at the control layer to achieve real-time encoding and transmission of thermal-electric parameter calculations and control instructions.

[0012] Preferably, in step S1:

[0013] The storage layer is composed of non-volatile memristor units, supporting multi-valued resistive state storage;

[0014] The computing layer integrates reconfigurable analog computing units to support matrix-vector multiplication operations;

[0015] The control layer includes distributed sensors, adaptive impedance matching networks and sparsity prediction units.

[0016] Preferably, the storage layer and the computing layer are vertically interconnected by hybrid bonding technology, and the impedance value of the adaptive impedance matching network is dynamically adjusted based on the real-time temperature field and voltage gradient.

[0017] Preferably, the coupling relationship model in step S2 satisfies:

[0018]

[0019] Among them, R is the resistance state of the memristor, T is the temperature field, V pulse is the voltage pulse parameter, α, β, γ are material related coefficients.

[0020] Preferably, the asymmetric thermal diffusion term of the temperature field is characterized by a layered thermal conductivity model, specifically:

[0021]

[0022] Among them, κ i (T) is the temperature-dependent thermal conductivity of the i-th layer.

[0023] Preferably, step S3 includes:

[0024] Precompile a sparse template library of typical AI models and load the corresponding sparse pattern according to the input model type;

[0025] The sparsity of the weight matrix is dynamically detected based on a probabilistic sparse sampling strategy, and mixed-precision pulses are allocated.

[0026] Preferably, the bit width allocation rule of the mixed-precision pulse is:

[0027]

[0028] Where s is the real-time sparsity of the weight matrix.

[0029] Preferably, the real-time coding transmission in step S4 adopts a pulse time interval coding technology, and the interval time and the resistance state adjustment amount satisfy a nonlinear logarithmic relationship.

[0030] Preferably, the control and gating steps of step S2 and step S3 and the dynamic reconstruction of the architecture of step S1 form a closed-loop feedback, and the closed-loop feedback collects temperature and voltage gradient data in real time through the distributed sensors of the control layer, and inputs them into the thermal-electric control algorithm and the sparse sensing gating module to form a dynamic optimization loop.

[0031] 3D stacked memristor array design system for energy-efficient AI computing, including:

[0032] Storage module: composed of a multi-layer non-volatile memristor array, supporting multi-valued resistive state storage;

[0033] Computing module: integrated with reconfigurable analog computing unit, supporting matrix operations and activation function processing;

[0034] Control module: including distributed sensors, adaptive impedance matching network, sparsity prediction unit and hybrid computing unit;

[0035] Communication module: uses pulse time coding technology to realize inter-layer instruction transmission.

[0036] This invention provides a design method for a 3D stacked memristor array for energy-efficient AI computing. It has the following beneficial effects:

[0037] 1. This invention utilizes a layered heterogeneous 3D stacking architecture and hybrid bonded vertical interconnect technology to achieve the synergistic optimization of near-field data processing and transmission loss. Compared to the storage wall problem of planar memristor arrays in existing technologies, this solves the additional energy consumption and latency issues caused by cross-layer data transfer, achieving tight physical coupling of storage and computing resources.

[0038] 2. This invention incorporates a dynamic thermal-electric coupling model and a distributed Kalman filter algorithm to form a closed-loop control mechanism for voltage pulse parameters driven by a temperature field gradient. This overcomes the drawback of traditional memristor designs that rely on static voltage drive strategies, suppresses the cumulative effect of resistance drift errors with temperature changes, and improves both device operating stability and lifespan.

[0039] 3. This invention builds a mixed-precision pulse distribution system guided by a sparse template library and establishes adaptive gating rules that link sparsity threshold criteria with energy efficiency constraints. This addresses the current state of AI accelerators, where weight sparsity is underutilized. This eliminates redundant power consumption generated by ineffective multiplication and addition operations, dynamically matching computational energy consumption with the model's sparse characteristics.

[0040] 4. This invention utilizes a hardware acceleration unit and pulse time interval encoding technology to achieve nanosecond-level real-time generation and transmission of control instructions. Compared to the multi-level cache bottleneck of traditional digital signal processing, this solves the mismatch between control latency and computational throughput, ensuring the rapid mode switching capability of the memristor array in multitasking scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 A perspective view of the present invention;

[0042] Figure 2 is a schematic diagram of the present invention; DETAILED DESCRIPTION

[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0044] Example:

[0045] Please see the attached Figure 1 , an embodiment of the present invention provides a design method for a 3D stacked memristor array for energy-efficient AI computing, comprising the following steps:

[0046] S1. Build a layered heterogeneous 3D stacking architecture, including physical isolation and dynamic reconstruction of the storage, computing, and control layers;

[0047] The storage layer is composed of non-volatile memristor cells, exemplarily using transition metal oxide materials as the resistive switching medium. These memristor cells support multi-valued resistive state storage. Resistive state switching is achieved by applying voltage pulses, whose amplitude and duration are correlated to the target resistance state. The computation layer integrates reconfigurable analog computing units, which include transconductance amplifiers and switched capacitor arrays capable of performing matrix-vector multiplication operations. The control layer deploys distributed temperature sensors and voltage gradient detectors, implemented as platinum resistance resistors or thermocouples, to collect real-time temperature field distribution and electrical signal fluctuation data.

[0048] The storage layer and the computing layer are vertically interconnected through hybrid bonding technology. For example, the hybrid bonding adopts a copper-copper direct bonding process, the through-hole diameter ranges from 0.5 to 2 microns, and the through-hole spacing is less than or equal to 1.5 times the through-hole diameter. An insulating dielectric layer is provided at the bonding interface, and the dielectric layer material is preferably silicon dioxide or silicon nitride to prevent interlayer short circuits. The through-hole array layout introduces a redundant design, and the number of redundant through-holes accounts for 10% to 20% of the total number of through-holes to compensate for through-hole failure problems during the manufacturing process.

[0049] An adaptive impedance matching network is deployed between the storage layer and the computing layer. The network consists of adjustable inductors and adjustable capacitors. The objective function of the impedance matching network is to minimize the signal reflection coefficient, which is defined by the following formula:

[0050]

[0051] Among them, Z match is the impedance matching network, Z line The impedance of the matching network is dynamically adjusted according to the real-time temperature field and voltage gradient, and the adjustment rule is related to the effect of temperature on the resistivity of the material and the spectral characteristics of the voltage pulse.

[0052] The dynamic reconfiguration mechanism is implemented through a programmable switch matrix consisting of nonvolatile memory cells and gating transistors. For example, the gating transistors are ferroelectric field-effect transistors (FeFETs) with a hafnium-doped oxide gate dielectric and a programmable threshold voltage. The switch matrix dynamically reconfigures the data paths between the storage and computing layers based on the computing task requirements. For example, in inference mode, some high-resistance cells are bypassed to reduce power consumption.

[0053] The control layer's Sparsity Prediction Unit (SPU) utilizes a hardware accelerator architecture, consisting of a finite state machine and a parallel comparator array. The SPU receives a weight data stream from the computation layer and generates a sparse mask based on a precompiled sparse template library. This sparse template library is constructed through offline analysis of the weight distribution patterns of typical AI models, exemplified by the spatial sparsity patterns of convolutional kernels and the channel sparsity patterns of fully connected layers. The mask instructs the computation layer to skip multiplication-add operations corresponding to zero weights.

[0054] The physical isolation of the storage and computing layers enables near-chip data processing and reduces off-chip data transmission energy consumption. Hybrid bonding technology combined with redundant via design improves vertical interconnect yield. An adaptive impedance matching network dynamically adjusts impedance parameters to suppress energy loss caused by signal reflections. The programmable switch matrix supports dynamic switching of functional modes to adapt to the resource requirements of different AI computing tasks. The sparsity prediction unit reduces inefficient computational operations through hardware-accelerated mask generation.

[0055] S2. Based on the dynamic thermal-electric coordinated control algorithm, a coupling relationship model between the temperature field and the resistance drift of the memristor is established to optimize the voltage pulse parameters in real time;

[0056] The implementation of the dynamic thermal-electric coordinated control algorithm first establishes a coupling relationship model between the temperature field and the resistance state drift of the memristor. This model describes the nonlinear effect of the spatiotemporal temperature distribution on the resistance state change through partial differential equations. The specific expression is:

[0057]

[0058] Among them, R is the resistance state of the memristor, T is the temperature field, V pulse is the applied voltage pulse parameter, α represents the coupling coefficient of temperature change on resistance state drift, β characterizes the spatial diffusion effect of the resistance state, and γ reflects the direct modulation strength of the voltage pulse on the resistance state. The equation is discretized and solved using the separation of variables method, with boundary conditions determined by the thermal conductivity of the device packaging material and the heat dissipation characteristics of the electrodes.

[0059] The asymmetric thermal diffusion behavior of the temperature field is characterized by a layered thermal conductivity model. The model considers the difference in thermal conductivity of different material layers in the 3D stacked structure, and the specific expression is:

[0060]

[0061] Among them, κ i (T) is the temperature-dependent thermal conductivity of the i-th layer, z i The model is discretized using the finite volume method. The spatial grid division is consistent with the physical layout of the memristor array. The time step is preferably in the microsecond range to ensure computational convergence.

[0062]

[0063] Among them, R k is the resistance state vector, u k is the control input (including voltage pulse amplitude, pulse width and frequency), w k and v k are process noise and observation noise respectively. The state transfer matrix A k The observation matrix C is constructed by correlating the thermal diffusion equation coefficient with the material thermal conductivity. k Spatial distribution design based on distributed temperature sensors.

[0064] The distributed Kalman filter algorithm is used to optimize the voltage pulse parameters in real time. First, the thermal-electric coupling model is discretized into a state space equation:

[0065]

[0066] in, is the prediction error covariance matrix, R k is the observation noise covariance matrix. The gain matrix is used to correct the state estimation value and output the optimized voltage pulse parameter V pulse .

[0067] The parameters of the adaptive impedance matching network are adjusted based on real-time temperature field and voltage gradient data. var With adjustable inductor L tune The value of satisfies:

[0068]

[0069] The impedance value dynamically matches the characteristic impedance of the transmission line to minimize the reflection coefficient Γ. The reflection coefficient calculation module is integrated into the mixed signal processing unit of the control layer and extracts the voltage pulse spectrum characteristics through fast Fourier transform.

[0070] The thermal-electric coupling model can characterize the nonlinear drift of the resistive state caused by temperature gradients, overcoming the accuracy limitations of traditional single-physics field models. The distributed Kalman filter algorithm, through multi-node data fusion, can suppress control errors caused by local temperature fluctuations. An adaptive impedance matching network dynamically adjusts transmission line impedance parameters, effectively reducing signal return loss. This collaborative control mechanism, aligned with the physical properties of the layered architecture, enables system-level energy efficiency optimization.

[0071] S3, through sparse-aware adaptive gating, dynamically allocates mixed-precision pulses based on the sparse characteristics of the AI model;

[0072] The implementation of sparse-aware adaptive gating begins with the creation of a library of precompiled sparse templates. This library is generated through offline analysis of the weight distribution characteristics of typical AI models, exemplified by the spatial sparsity patterns of filter weights in convolutional neural networks and the channel sparsity patterns of attention weights in Transformer models. Sparse pattern extraction utilizes a threshold-based approach, marking elements with an absolute weight value less than a set threshold as zero. The threshold is preferably a percentile value of the weight distribution statistic.

[0073] The probabilistic sparse sampling strategy is used to dynamically detect the real-time sparsity of the input weight matrix. First, the weight matrix is divided into multiple sub-blocks, exemplarily using 8×8 or 16×16 rectangular blocks. Then, some sub-blocks are randomly selected to count the non-zero elements. The sparsity s is calculated using the following formula:

[0074]

[0075] Among them, M s is the sparse mask matrix, W sub is the sampling sub-block weight, ⊙ represents the Hadamard product, and ||0 counts the number of non-zero elements. The sampling rate is preferably 5% to 20% to balance detection accuracy and computational overhead.

[0076] The mixed-precision pulse allocation rule dynamically adjusts the bit width according to the real-time sparsity. The bit width allocation function is defined as a piecewise linear relationship:

[0077]

[0078] Among them, s high With s low are high and low sparsity thresholds, respectively, and are preferably set to 0.9 and 0.5. The bit width switching is achieved by a programmable pulse generator, which includes a multi-channel digital-to-analog converter and a pulse width modulation module.

[0079] Energy consistency constraints ensure that the energy consumption of pulses with different precisions is balanced. For bit width b, the pulse amplitude V b satisfy:

[0080]

[0081] Among them, E const The constraint is implemented through a capacitor charge integration circuit, and the integrator output is fed back to the pulse amplitude control module to form a closed-loop regulation.

[0082] The sparse mask hardware acceleration unit utilizes a parallel comparator array. This array contains comparator units of the same dimensions as the weight matrix. Each unit compares the input weight with a threshold voltage and outputs a binary mask signal. The threshold voltage is dynamically configured via a digital-to-analog converter (DAC), with the configured value associated with the sparsity pattern in a precompiled template library. The mask signal is compressed by a priority encoder and transmitted to the computation layer, instructing the multiplication-addition unit to skip zero-weight operations.

[0083] The precompiled sparse template library can capture the inherent sparsity of AI models in advance, reducing the real-time computing load. The probabilistic sparse sampling strategy has the effect of reducing hardware resource usage. The mixed-precision pulse allocation rule dynamically adjusts the calculation accuracy based on the sparsity, achieving a balance between energy efficiency and accuracy. The energy consistency constraint ensures the comparability of energy consumption of pulses of different precisions through a closed-loop feedback mechanism. The sparse mask hardware acceleration unit can reduce the number of invalid computational operations. The gating mechanism matches the computing resource distribution characteristics of the 3D stacking architecture, which can improve the overall system energy efficiency.

[0084] S4. Deploy hardware acceleration units at the control layer to achieve real-time encoding and transmission of thermal-electric parameter calculations and control instructions.

[0085] This step involves the deployment of a control-layer hardware acceleration unit and the implementation of a real-time encoding and transmission mechanism. This hardware acceleration unit forms a data exchange link with the thermal-electrical control algorithm and the sparse-sensing gating module, optimizing command generation efficiency through hierarchical signal processing. Based on dynamic impedance matching and resistance drift compensation, this unit converts control commands into pulse sequences that adapt to the physical constraints of the 3D stacked architecture.

[0086] In some embodiments, the hybrid computing unit uses an analog-digital hybrid circuit to implement thermal-electric parameter calculation. The transimpedance amplifier receives the current signal from the distributed sensor, and the output end is connected to a Gilbert cell multiplier. The multiplier performs the following operations:

[0087] V out =k·(V thermal ×V elec );

[0088] Among them, V thermal is the temperature gradient conversion voltage, V elec is the voltage pulse gradient signal, and k is the transconductance gain coefficient. The calculation result is quantized by the comparator threshold and input into the pulse time encoding module.

[0089] As an option, the pulse time coding technology uses a nonlinear mapping rule. interval The conversion relationship is defined as:

[0090] t interval=τ·ln(1+|ΔR| / R base );

[0091] Where τ is the time constant, R base is the reference resistance value. This logarithmic relationship compresses the encoding length of the high-resistance interval to avoid signal crosstalk caused by dense pulse sequences. The encoded data packet header contains a 4-bit checksum, and the payload uses the Manchester encoding format to improve noise immunity.

[0092] Specifically, a time-to-digital converter (TDC) is used to decode the pulse interval. A ring oscillator generates a high-frequency clock signal, and a phase accumulator records the number of oscillation cycles between rising edges of the pulse. The decoding formula is:

[0093]

[0094] The resolution of the converter is determined by the oscillator frequency. Preferably, the oscillation frequency is not less than 200 MHz to ensure decoding accuracy.

[0095] In one possible implementation, the command transmission channel uses a differential signaling topology. Shielded twisted-pair cables are laid out in an interlayer dielectric, and common-mode chokes are loaded at the impedance matching network terminals. The drive circuit uses current-mode logic (CML), and a slew-rate control module limits the rate of change of signal edges to suppress high-frequency radiated interference.

[0096] In some embodiments, a closed-loop feedback data stream integrates temperature gradients, voltage pulse parameters, and sparsity detection results. The data buffer utilizes a dual-port SRAM structure, with a write port receiving sensor data and a read port connected to the input of the Kalman filter. A priority arbiter dynamically allocates bus bandwidth to ensure real-time transmission of thermal and electrical control commands.

[0097] Typically, the programmable logic portion of the hardware acceleration unit is implemented on an FPGA. A lookup table (LUT) stores thermal-electrical coupling model coefficients and sparse template indices, while an arithmetic logic unit (ALU) performs gradient calculation and mask generation in parallel. Configuration registers support online updates of control algorithm parameters, such as the Kalman filter's noise covariance matrix and impedance matching target values.

[0098] Alternatively, pulse train modulation employs dual-carrier quadrature modulation. Baseband pulse signals are superimposed on 2.4 GHz and 5.8 GHz carriers via a mixer. After synthesis, these signals are radiated and transmitted via an interlayer antenna array. An envelope detector at the receiving end extracts the modulated signal, and an automatic gain control (AGC) circuit compensates for transmission path losses.

[0099] Specifically, the anti-interference mechanism includes forward error correction coding and adaptive equalization. A convolutional encoder adds redundant check bits, and a Viterbi decoder corrects transmission errors. The equalizer tap coefficients are dynamically adjusted based on the channel impulse response, and the minimum mean square error (MMSE) algorithm update step size is preferably in the order of 0.01 to 0.1.

[0100] Please see the attached Figure 2 , an embodiment of the present invention provides a 3D stacked memristor array design system for high-energy-efficiency AI computing, including:

[0101] Storage module: composed of a multi-layer non-volatile memristor array, supporting multi-valued resistive state storage;

[0102] The memory module consists of a vertically stacked multi-layer non-volatile memristor array, preferably using transition metal oxide materials as the resistive dielectric layer. The memristor cells in each layer are three-dimensionally interconnected using a hybrid bonding process, with via diameters preferably ranging from 0.5 to 2 microns. A silicon nitride insulating layer is placed at the bonding interface to prevent interlayer leakage. Resistive state storage is achieved through multi-level voltage pulse modulation, with pulse amplitude and duration correlated to the target resistance value. A resistance drift compensation algorithm is embedded in the control module.

[0103] The non-uniformity of the resistance distribution is alleviated by redundant through-hole design, and the proportion of redundant units is preferably 10% to 20%. The thermal diffusion model is characterized by the layered thermal conductivity equation, and the temperature field evolution of the i-th layer satisfies:

[0104]

[0105] Among them, κ i is the thermal conductivity of each layer of material, Q gen is the Joule heat generation term, which is positively correlated with the resistance state and operating frequency.

[0106] Computing module: integrated with reconfigurable analog computing unit, supporting matrix operations and activation function processing;

[0107] The computation module integrates a reconfigurable analog computing unit, comprising a transconductance amplifier array and a switched-capacitor network. For example, the transconductance amplifier gain is adjustable by programming the floating-gate transistor threshold voltage, with a dynamic range of 60-100 dB. Matrix-vector multiplication is implemented using the charge-sharing principle, with computational errors suppressed using a capacitor mismatch calibration algorithm. The calibration coefficients are stored in specific resistive cells within the memory module.

[0108] The activation function uses a piecewise linear approximation circuit. Preferably, the sigmoid function is implemented using a six-segment broken line fit, with a nonlinear error of ≤1.5%. Operational mode switching is controlled by a ferroelectric transistor (FeFET) switch array, with a switching delay of ≤10ns, synchronized with the resistive read timing of the storage module.

[0109] Control module: including distributed sensors, adaptive impedance matching network, sparsity prediction unit and hybrid computing unit;

[0110] The distributed sensor network consists of temperature sensors and voltage gradient detectors. For example, the temperature sensors are platinum resistors with a sensitivity of 0.385Ω / °C. The voltage gradient detectors are based on differential amplifiers with a common-mode rejection ratio of ≥80dB and a bandwidth covering the 100MHz to 2GHz signal spectrum.

[0111] The adaptive impedance matching network consists of an array of adjustable inductors and capacitors, whose impedance value Z match The dynamic optimization goal is:

[0112]

[0113] Where Zline is the nominal impedance of the transmission line, ΔT is the real-time temperature deviation, and T ref The matching network parameter update cycle matches the thermal diffusion time constant.

[0114] The sparse prediction unit adopts a parallel comparator architecture, the input weight matrix is divided into 8×8 sub-blocks, and the non-zero element detection is realized by the threshold comparison circuit. th Dynamic configuration, where configuration values are associated with sparse pattern features in a precompiled template library. The hybrid computing unit integrates a multiplier-accumulator and nonlinear function hardware, supporting the joint calculation of thermal and electrical parameters. Operation priority is dynamically assigned by the task scheduler.

[0115] Communication module: uses pulse time coding technology to realize inter-layer instruction transmission.

[0116] Pulse time coding technology converts control instructions into non-uniformly spaced pulse sequences. The coding rules meet the following requirements:

[0117]

[0118] Among them, τ is the encoding time constant, ΔR is the resistance adjustment amount, R base The decoding end uses a time-to-digital converter (TDC), the ring oscillator frequency is preferably 200-500MHz, and the decoding error is ≤1%.

[0119] The inter-layer transmission channel adopts a differential signal topology, with shielded twisted pair cables laid in the inter-layer dielectric, and a characteristic impedance matching error of ≤5%. The drive circuit integrates a slew rate control module, limiting the edge change rate to 1-3V / ns to suppress high-frequency crosstalk. The forward error correction coding uses a (15,11) Hamming code, with the check bit embedded in the pulse sequence header, reducing the bit error rate to 10 -9 Magnitude.

[0120] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A design method for a 3D stacked memristor array for energy-efficient AI computing, characterized by: The following steps are involved: S1. Build a layered heterogeneous 3D stacking architecture, including physical isolation and dynamic reconstruction of the storage, computing, and control layers; S2. Based on the dynamic thermal-electric coordinated control algorithm, a coupling relationship model between the temperature field and the resistance drift of the memristor is established to optimize the voltage pulse parameters in real time; S3, through sparse-aware adaptive gating, dynamically allocates mixed-precision pulses based on the sparse characteristics of the AI model; S4. Deploy hardware acceleration units at the control layer to achieve real-time encoding and transmission of thermal-electric parameter calculations and control instructions.

2. The design method of a 3D stacked memristor array for energy-efficient AI computing according to claim 1, characterized in that: In the step S1: The storage layer is composed of non-volatile memristor units, supporting multi-valued resistive state storage; The computing layer integrates reconfigurable analog computing units to support matrix-vector multiplication operations; The control layer includes distributed sensors, adaptive impedance matching networks and sparsity prediction units.

3. The design method of a 3D stacked memristor array for energy-efficient AI computing according to claim 2, wherein: The storage layer and the computing layer are vertically interconnected by hybrid bonding technology, and the impedance value of the adaptive impedance matching network is dynamically adjusted based on the real-time temperature field and voltage gradient.

4. The design method of a 3D stacked memristor array for energy-efficient AI computing according to claim 1, wherein: The coupling relationship model in step S2 satisfies: Among them, R is the resistance state of the memristor, T is the temperature field, V pulse is the voltage pulse parameter, α, β, γ are material related coefficients.

5. The design method of a 3D stacked memristor array for energy-efficient AI computing according to claim 4, wherein: The asymmetric thermal diffusion term of the temperature field is characterized by a layered thermal conductivity model, specifically: Among them, κ i (T) is the temperature-dependent thermal conductivity of the i-th layer.

6. The design method of a 3D stacked memristor array for energy-efficient AI computing according to claim 1, wherein: The step S3 comprises: Precompile a sparse template library of typical AI models and load the corresponding sparse pattern according to the input model type; The sparsity of the weight matrix is dynamically detected based on a probabilistic sparse sampling strategy, and mixed-precision pulses are allocated.

7. The design method of a 3D stacked memristor array for energy-efficient AI computing according to claim 1, wherein: The bit width allocation rule of the mixed precision pulse is: Where s is the real-time sparsity of the weight matrix.

8. The design method of a 3D stacked memristor array for energy-efficient AI computing according to claim 1, wherein: The real-time coding transmission in step S4 adopts a pulse time interval coding technology, and its interval time and the resistance state adjustment amount satisfy a nonlinear logarithmic relationship.

9. The design method of a 3D stacked memristor array for energy-efficient AI computing according to claim 1, wherein: The control and gating steps of step S2 and step S3 and the dynamic reconstruction of the architecture of step S1 form a closed-loop feedback. The closed-loop feedback collects temperature and voltage gradient data in real time through the distributed sensors of the control layer, and inputs them into the thermal-electric control algorithm and the sparse sensing gating module to form a dynamic optimization loop.

10. A 3D stacked memristor array design system for high-energy-efficiency AI computing, according to a 3D stacked memristor array design method for high-energy-efficiency AI computing according to any one of claims 1 to 9, characterized in that: include: Storage module: composed of a multi-layer non-volatile memristor array, supporting multi-valued resistive state storage; Computing module: integrated with reconfigurable analog computing unit, supporting matrix operations and activation function processing; Control module: including distributed sensors, adaptive impedance matching network, sparsity prediction unit and hybrid computing unit; Communication module: uses pulse time coding technology to realize inter-layer instruction transmission.