Neural network parallel computing system suitable for wafer-level stacking

WO2026200660A1PCT designated stage Publication Date: 2026-10-01HUANG ZISHENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/084348
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-10-02
Filing Date
2026-03-18
Publication Date
2026-10-01

Smart Images

  • Figure CN2026084348_01102026_PF_FP_ABST
    Figure CN2026084348_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a system and method based on a dedicated digital and analog hybrid circuit for ultra-large-scale neural network computation in the field of artificial intelligence, and more particularly, to a supercomputing cluster for neural networks based on bonding between silicon wafers or heterogeneous bonding between silicon wafers and thin-film integrated circuits, and a configuration and method for internal interconnection and intercommunication thereof.
Need to check novelty before this filing date? Find Prior Art

Description

A neural network parallel computing system suitable for wafer-level stacking Technical Field

[0001] This invention relates to a system and method for computing ultra-large-scale neural networks based on a hybrid digital and analog circuit in the field of artificial intelligence. More specifically, it relates to the structure and method of a supercomputing cluster based on the combination of silicon wafers or the heterogeneous combination of silicon wafers and thin-film integrated circuits, and the interconnection of the cluster. Background Technology

[0002] With the rapid development of artificial intelligence (AI) technology, neural networks, especially ultra-large-scale neural networks, have demonstrated remarkable capabilities in various fields such as image recognition, natural language processing, and speech recognition. DeepSeek and other technologies have proven that by employing specific techniques and optimization strategies, it is possible to train and apply large-scale models at a relatively low cost. Countries worldwide have invested trillions of dollars in this area in hopes of achieving AGI; however, current implementation methods still have significant limitations.

[0003] I. Solution based on computing chip + GPU:

[0004] First, existing implementations of ultra-large-scale neural networks primarily rely on server clusters and expensive GPUs. Essentially, they involve digital simulations of complex computational processes, leading not only to high hardware costs but also increased system energy consumption and physical size. High energy consumption raises operating costs and negatively impacts the environment; while the massive system size limits their usability in mobile devices or space-constrained applications. Furthermore, because these systems typically require continuous high-performance computing support, their power supply requirements are very stringent, further limiting the flexibility of their application scenarios.

[0005] Secondly, while technological advancements by companies like DeepSeek have significantly reduced the cost of training and deploying large-scale neural networks, their implementation still heavily relies on complex and sophisticated software programming techniques. This means that development teams must possess deep expertise and skills to effectively design, train, and optimize these models. For many small and medium-sized enterprises or research institutions with limited resources, such requirements constitute a significant barrier to entry, limiting the speed of technological innovation and adoption.

[0006] Finally, existing solutions still have several orders of magnitude of room for improvement in performance and energy efficiency compared to the human brain (which consumes approximately 20 watts). In particular, current technology has not yet reached its ideal state in terms of concurrency, processing speed, response time, and energy efficiency. For example, in certain real-time applications, such as autonomous vehicles or instant translation services, rapid and accurate decision-making is crucial, and existing solutions based on servers and computing chips + GPUs often fall short of meeting these requirements.

[0007] II. In-Memory Computing (IMC). As a cutting-edge technology, IMC has demonstrated enormous potential in the field of neural networks. However, despite its advantages and the abundance of solutions available, IMC still faces challenges in practical applications, including durability and reliability issues, low data accuracy, small network size, and incomplete activation functions, making it difficult to effectively support the training and inference of complex models. These problems have hindered the widespread adoption of IMC solutions in large-scale neural network applications.

[0008] III. Analog Circuit Implementation. Using analog circuits similar to operational amplifiers to implement neural networks currently offers no advantages over digital circuits. The design is complex, the network design is inflexible, the data accuracy is low, the network size is small, and the performance is also low.

[0009] IV. In terms of the macroscopic structure of computer clusters, it mainly involves chip wafers based on advanced CMOS processes. After testing and dicing, they are reassembled and packaged. After 3D packaging between the chips, they are interconnected using complex board-level circuits and interconnected at the module level through fiber optic communication, thereby establishing a computing cluster. Advanced CMOS processes are inherently very expensive, and the entire testing, packaging, and assembly process is extremely complex, cumbersome, and costly. The bandwidth, energy efficiency, and performance of the inter-chip interconnect communication during use are also poor.

[0010] V. Thin-film integrated circuits are often used in the manufacture of display panels. Due to their low transistor density and operating frequency, they have not been used in large-scale neural network computing. However, existing thin-film transistor switches have relatively low leakage current. Technical issues

[0011] In summary, current implementation schemes for ultra-large-scale neural networks all suffer from significant drawbacks and challenges. While the computing chip + GPU-based approach offers high computing power, its high cost, high energy consumption, and physical space constraints limit its widespread application. In-memory computing excels in reducing data transmission latency and power consumption, but issues with the durability and reliability of storage media, insufficient data precision, and difficulty in supporting large-scale model training and inference hinder its large-scale adoption. Analog circuit implementations face numerous problems such as design complexity, poor flexibility, low data precision, and poor performance, making it difficult to compete with digital circuits. Furthermore, in terms of network construction, traditional accumulation and transfer methods have failed to adequately integrate with other hardware technologies, resulting in performance and energy efficiency far below expectations. Building computer clusters has also been a disaster. These problems collectively indicate that existing technologies and methods are insufficient to fully meet the practical needs of ultra-large-scale neural networks, necessitating innovative architectural designs and circuit implementations to overcome these limitations. Technical solutions

[0012] This invention proposes a parallel computing system for neural networks suitable for wafer-level stacking. The basic neuron unit includes: a charge adjustment module, a capacitor control module, a capacitor module, and a neuron output module. The capacitor control module and the capacitor modules are differentially configured, each with positive and negative terminals, or a primary and secondary configuration. Each capacitor control module has a circuit for charging and discharging its corresponding capacitor module, capable of independently charging and discharging the corresponding capacitor module to a specific voltage to prepare for simulated neural network detection. The charge adjustment module changes the voltage or capacitance of the capacitor module based on the neuron's weights and inputs, thereby changing its charge before detection. The capacitor control module has a charging and discharging current channel for detection, capable of charging and discharging the corresponding capacitor module. The capacitor control module contains a voltage comparator that, during the detection of capacitor voltage changes, compares the voltage of the capacitor within the capacitor module with a target reference voltage, or compares the voltage of the capacitor within the capacitor module with that of a capacitor of the same symbol. The differential capacitor control module compares the voltage of the capacitors, generates a voltage comparison signal, and terminates the detection based on this signal. Based on the duration or operation of the charging / discharging current channel from conduction to the generation of the voltage comparison signal terminating the detection, and the corresponding bit or ratio, each differential capacitor control module outputs single-symbol calculation information. This single-symbol calculation information corresponds to the duty cycle in pulse-width modulation, the pulse width time in single-pulse pulse-width modulation, the number of switching capacitor operations in switched-capacitor current modulation, and the capacitance value added in a varactor capacitor network discharge operation. The neuron output module generates differential multiplication-addition information based on the differential single-symbol calculation information. The neuron output module generates a symbol level or its digitized signal, i.e., symbol information, based on the comparison of the differential single-symbol calculation information or the voltage comparison of the differential capacitor modules. The output of the basic neuron unit, including the differential multiplication-addition information and the symbol information, is transmitted or stored in its own input or external circuitry.

[0013] This invention also proposes a charge adjustment module comprising an input current modulation control module and a weighted current modulation control module, capable of adjusting the voltage of the capacitor module through modulation current charging and discharging, thereby adjusting the charge amount; it includes one or more reference time source signals; the input current modulation control module generates an input current modulation control signal based on an external signal, or a signal from the neuron output module, or based on the reference time source signal and the value of its loaded input current fast access unit; the weighted current modulation control module further selects, trims, or inserts delays based on the value of its own weighted fast access unit and the corresponding input current modulation control signal to generate... The system includes a control signal for modulating the current, or adjustment of the equivalent resistance of the current channel; multiple current channels linked by a corresponding neural network, each capable of charging and discharging a capacitor module with corresponding attributes according to a set setting; each current channel has a switch for generating the modulated current, controlled by the control signal for modulating the current, generating a modulated current calculated by an analog neural network for charging and discharging the capacitor module; the signs of both the input item fast access unit and the weight item fast access unit jointly control the selection of positive / negative or primary / secondary capacitor modules for charging and discharging; after the modulated current calculated by the analog neural network charges and discharges the capacitor module, the equivalent resistance is detected to continue charging and discharging the capacitor module.

[0014] The present invention also proposes that the charge adjustment module has a circuit for adjusting the capacitance value of the main capacitor of the capacitor module, including a parallel switch and a parallel capacitor; the signs of the input item fast access unit and the weight item fast access unit jointly control the selection of positive / negative or main / auxiliary capacitor modules for capacitance adjustment; after the capacitance value of the main capacitor of the capacitor module is adjusted, its voltage is simultaneously adjusted to a specific voltage, and then the detection stage is entered.

[0015] The present invention also proposes that the capacitor module is a combination of capacitors and switches; the capacitor module contains a main capacitor and a detection capacitor; in the constant capacity mode, that is, the charge adjustment module only adjusts the voltage of the corresponding capacitor module, wherein a certain main capacitor and a certain detection capacitor are charged and discharged together through a switch combination; the remaining disconnected or uncombined detection capacitors are charged and discharged independently or their voltages are compared.

[0016] This invention also proposes a circuit with multiple layers for detecting charge and discharge. The first layer detects charge and discharge to generate differential multiply-add information to control the charge and discharge of the detection capacitor in the next layer. The analog layer selects the charge and discharge mode and charge and discharge layer depth of the detection capacitor according to the settings to generate a linear or nonlinear activated output. The output is subsequently used as input to its own layer or other layers.

[0017] This invention also proposes a multi-layer neural network, in which the middle layer is an analog layer and the beginning and end are digital layers; the output of the analog layer uses equivalent single-pulse pulse width modulation; the equivalent single-pulse pulse width modulation means that the output of the neuron is a pulse signal of one or more channels, and the output of a single channel is one or more independent single-pulse signals.

[0018] The present invention also proposes a multi-layer neural network, wherein each layer of the neural network contains multiple basic neuron units; the differential multiplication-addition information output by the basic neuron units of the last layer of the multi-layer neural network is expressed in binary as the value of each bit, i.e., digitized; the output data of the multi-layer neural network is transmitted using the digitized differential multiplication-addition information and the symbolic information.

[0019] The present invention also proposes an array of basic neuron units in the dies within the wafer, corresponding to the multilayer neural network; adjacent dies have serial or parallel communication lines for transmitting their input or output data; the input or output data of the multilayer neural network includes the digitized differential multiply-add information and the symbol information; the die array or division within the wafer, and adjacent dies within the region use ring-shaped or bilinear independent parallel communication.

[0020] The present invention also proposes that the dies within the wafer have bonding pads; the wafers are stacked together, and the stacks are connected by the bonding pads; the bonding pads include independent upper and lower communication channels for the dies in the stack; the dies within the wafer have independent communication channels connecting adjacent dies; the independent upper and lower communication channels and the independent communication channels of adjacent dies can be combined into a transfer channel, so that data between the dies in the stack and data between adjacent dies can be interconnected.

[0021] The present invention also proposes that there are damaged dies in the die array within the stacked wafer; the damaged dies have independent communication channels above and below and independent communication channels for the normal operation of adjacent dies; the working dies of the connection between the array coordinates corresponding to the damaged dies of two wafers in the stack pass through the transfer channel, so that data transmission avoids the damaged dies.

[0022] This invention also proposes current modulation using the operation of switched capacitors; the capacitor module has a main capacitor; the switched capacitors are divided into a pre-connection group and a post-connection group; the pre-connection group is connected to the main capacitor of the capacitor module before each current modulation of the switched capacitors occurs; the post-connection group is connected to the main capacitor of the capacitor module after each current modulation of the switched capacitors occurs; the capacitors of the pre-connection group and the post-connection group operate synchronously to correct the calculation deviation caused by the parallel discharge of the switched capacitors; while the capacitors of the pre-connection group are disconnected from the main capacitor and grounded or connected to other voltage sources, the capacitors of the post-connection group are disconnected from ground or other voltage sources and connected to the main capacitor, and the original connection state is synchronously restored after the voltage stabilizes.

[0023] The present invention also proposes that the switched capacitors be divided into groups with equal pre-link and post-link numbers, and the smallest unit of a group is a pre-link capacitor and a post-link capacitor; the switching capacitors of each group are operated in a staggered timing sequence.

[0024] This invention also proposes that the detection equivalent resistance be a switched capacitor network or a resistor network; a circuit for detecting accelerated discharge during the charging and discharging process is used, and the size of the detection equivalent resistance is determined based on the voltage of the main capacitor to be detected and the target reference voltage of the detection stage, that is, the size of the modulation current of the charging and discharging; through digital control logic, the operation results and the size of the equivalent resistance value of the operation are combined to achieve detection acceleration and digitization.

[0025] This invention also proposes a wafer-level stacked parallel computing system for neural networks, whose basic neuron unit has two differential capacitors, characterized by a computational method for neurons:

[0026] Step 1: Obtain the values ​​of the input items and the weight items;

[0027] Based on the sign of the product of the input item's value and the weight item's value, if it is a capacitor voltage adjustment mode, open the current channel of the corresponding capacitor; or if it is a capacitor capacitance adjustment mode, open the variable capacitance channel of the corresponding capacitor.

[0028] Step 2, in parallel, the two differential capacitors are connected to a specific voltage source and pre-charged and discharged to a specific voltage;

[0029] Step 3, or if it is the mode of adjusting the capacitance value, turn on the adjustment switch of the two differential capacitors; if it is the mode of adjusting the capacitor voltage, in the corresponding current channel of the neuron link, convert the value of the input item and the value of the weight item into the charging and discharging modulation current of the capacitor, that is, the product of the equivalent conductance and the current conduction time generates the modulation current to charge and discharge the capacitor.

[0030] Step 4: After the capacitor charge adjustment is completed, the detection process begins, and charging and discharging continues until the threshold voltage is reached, generating two differential single-symbol calculation information.

[0031] The single-symbol calculation information corresponds to the duty cycle in pulse width current modulation, the pulse width time in single-pulse pulse width modulation, the number of switching capacitor operations in switched capacitor current modulation, and the capacitance value of the capacitor added in the discharge operation of the varactor capacitor network.

[0032] Step 5: Based on the two single-symbol calculation information of the difference, generate the difference multiplication and addition information and symbol, or its digitized value, as the output of the basic neuron unit.

[0033] This invention also proposes a method for generating activation function output through a multi-layer detection charging and discharging circuit: Step 1, the first-layer charging and discharging circuit is set to a linear mode and generates a linear output of the differential multiply-accumulate information; Step 2, the digital logic circuit, based on the maximum value of the segmented linear segment of the activation function, if it exceeds the maximum value, uses the remaining time of the first-layer linear output as the input of the second-layer charging and discharging circuit; Step 3, the second-layer charging and discharging circuit is set to a non-linear mode and generates a non-linear output of the differential multiply-accumulate information; Step 4, the digital logic circuit merges the outputs of the differential multiply-accumulate information from the first and second layers; Step 5, the merged differential multiply-accumulate information is used as input to its own or an external circuit.

[0034] This invention also proposes a method for correcting the discrete error in parallel calculations of switched capacitors when using switched capacitors to modulate current: During the parallel operation, the operation of the two types of capacitors is synchronized, that is, the operation is completed in approximately the same time period; Step 1, before the parallel operation of the switched capacitors, one of the pre-connected capacitors is connected to the main capacitor and disconnected from the charging / discharging voltage source, and the other is disconnected from the main capacitor and connected to the charging / discharging voltage source; Step 2, during the parallel operation of the switched capacitors, the pre-connected capacitor is disconnected from the main capacitor and connected to the charging / discharging voltage source, and the latter-connected capacitor is connected to the main capacitor and disconnected from the charging / discharging voltage source; Step 3, after the voltage of the switched capacitors stabilizes, the original on and off states are restored; Step 4, after the voltage of the main capacitor stabilizes, one round of parallel charging / discharging operation of the switched capacitors is completed.

[0035] This invention also proposes a method for bit-by-bit segmented calculation: Step 1, the input value and weight value of the neuron are each divided into segments with fewer bits according to binary bits; Step 2, the input segment value and weight segment value of the neuron are selected according to the factorization rule; Step 3, the selected input segment value and weight segment value are used to charge and discharge the capacitor module through current modulation; Step 4, after the charging and discharging of the capacitor is completed, the subsequent detection operation is performed, and the segmented difference multiplication and addition information is obtained; Step 5, the operation is repeated in step 2 until the calculation of each term of the factorization is completed; Step 6, the final equivalent difference multiplication and addition information is obtained, and the digital circuit completes the summation of each term of the factorization through shift addition.

[0036] This invention also proposes a method for accelerating the detection of discharge, obtaining the differential multiplication and addition information, and digitizing it: Step 1, select the largest bit by comparing the voltage of the detection capacitor and the reference voltage corresponding to the binary bit; Step 2, adjust the discharge path to discharge once or for a period of time with a current proportional to the actual value of the corresponding binary bit; Step 3, adjust the value of the fast storage unit corresponding to the differential multiplication and addition information in the digital circuit accordingly; Step 4, return to Step 1 and repeat until the detected capacitor reaches the threshold voltage.

[0037] This invention also proposes a constant-capacity mode, where the charge adjustment module only adjusts the voltage of the corresponding capacitor module. This mode includes multiple main capacitors and multiple detection capacitors, and a method to accelerate the operating frequency of the main capacitors: Step 1, select idle main capacitors and idle detection capacitors; Step 2, connect the idle main capacitors and idle detection capacitors, i.e., merge them; Step 3, pre-charge the merged capacitors together to a specified voltage; Step 4, perform charging and discharging of neural connections together; Step 5, disconnect the original main capacitors and original detection capacitors in the merged capacitors, and the original main capacitors enter an idle state; Step 6, perform a detection discharge operation on the original detection capacitors; Step 7, the original detection capacitors enter an idle state.

[0038] This invention also proposes a fully connected segmentation calculation method: In a layer of a neural network, the corresponding input and output neurons and their corresponding weights are divided into multiple groups; Step 1, through linear, circular, or cross-unit parallel transmission, each group synchronously exchanges input values; Step 2, in parallel, according to the program, the weight values ​​corresponding to the corresponding inputs are loaded; Step 3, each group independently calculates using the neuron calculation method of the aforementioned computing circuit system; Step 4, through digital circuits, the calculation results of other groups are added to the calculation results of this group; Step 5, return to Step 1 and repeat the operation; Step 6, obtain the final calculation result.

[0039] This invention also proposes a method for linear output: the charging and discharging modulation current in step 3 is the discharge current, and the discharge driving voltage is the ground potential; the charging and discharging process in the detection process in step 4 is the discharge, and the discharge driving voltage is the ground potential; wherein, the differential multiplication and addition information obtained in step 5 is linearly proportional to the sum of the products of the input terms and weight terms of each link, thereby realizing the linear activation function of the neuron output.

[0040] This invention also proposes a method for dynamically configuring the computing circuit system as one or more computing layers in a neural network to be computed: for a basic neuron unit in a layer of a neural network, the values ​​of input items are loaded into the input item fast access unit from a memory, external circuit, or the output module of the previous layer of neurons, and the values ​​of weight items are loaded into the weight item fast access unit; the computation of the computing layer is completed; when the computation is completed, the difference multiplication and addition information and its sign output by the basic neuron unit of the last layer are output to a memory, as the input of the subsequent computing layer, external circuit, or other processing module.

[0041] This invention also proposes a heterogeneous form for the single-symbol computation information corresponding to the input and output of a single neuron, which realizes the transformation of the single-symbol computation information form throughout the entire layer; the heterogeneous form is that one of the following is used as the input: the number of times the switched capacitor is operated, the duration of a single pulse, the pulse width modulation or the digitized value, while the output is selected in a different form than the input.

[0042] This invention also proposes a reference neuron whose output is the pulse output of the capacitor control module, used to control the charging and discharging time of the equivalent detection resistor; the reference neuron can control the single-symbol calculation information output of the capacitor control module according to parameter settings; the digital control logic selects the corresponding reference neuron's single-symbol calculation information output according to the real-time voltage range of the detection capacitor or the selected equivalent resistance value, thereby controlling the standard duration of accelerated charging and discharging in the detection process.

[0043] The present invention also proposes that the aforementioned computing circuit system be dynamically used as one or more layers of the entire neural network to be computed; the charge adjustment module loads, from memory or external circuitry, or shared or loaded from the neuron output module, the value of the input item fast access unit; the charge adjustment module loads, from memory or external circuitry, or shared or loaded from the neuron output module, the value of the corresponding weight item fast access unit; there is an equivalent resistance network that generates a specified equivalent resistance value based on the value of the weight item fast access unit, or also based on the value of the input item fast access unit, thereby adjusting the instantaneous current of the corresponding current channel; at the end of the computation, the value of the activation function fast access unit of the last layer is output to memory, its own input, external circuitry, or other modules.

[0044] The present invention also proposes an array-type memory used as a cache; the array-type memory has a row-selective write circuit to write the output of a neural network layer to the selected row; the array-type memory has a transposed block-selective read circuit to write the data of a selected block in a column to the input or weight of a neural network layer.

[0045] The present invention also proposes that the input item fast access unit is segmented bit by bit, and each segment independently generates one or more segment input item modulation signals; the weight item fast access unit is segmented bit by bit; each segment of the weight item fast access unit independently generates multiple modulation currents according to its own segment weight value and the corresponding segment input item modulation signal; the multiple modulation currents simultaneously charge and discharge the capacitor module of the corresponding segment.

[0046] This invention also proposes a layered activation function numerical source, which generates numerical information with equivalent time as the independent variable based on the setting of the activation function; the numerical information with equivalent time as the independent variable is a time value, an activation function value, a derivative value, a calculation intermediate value, or a combination thereof; the neuron output module has an activation function fast access unit; the neuron output module selects the numerical information with equivalent time as the independent variable corresponding to the symbol according to the symbol level; when the differential multiply-add information signal starts or ends the shearing, based on the numerical information with equivalent time as the independent variable of the corresponding symbol, the output value of the activation function of the corresponding neuron basic unit at the current time is obtained.

[0047] The present invention also proposes the aforementioned layer activation function numerical source, which generates activation function numerical information according to the setting of the activation function and sends it to the neuron output module of the corresponding layer;

[0048] The neuron output module generates new charge / discharge single-symbol calculation information based on the activation function numerical information; the neuron output module also generates a new activation function value based on the new charge / discharge single-symbol calculation information and another activation function numerical information.

[0049] This invention also proposes a network summarization layer; the network summarization layer has network summarization layer neurons that connect to the basic computational units of neurons in the current layer, and uses the activated values ​​of each neuron in the current layer as the input of the network summarization layer neurons; a layer computation module receives the output of the network summarization layer neurons and generates a layer parameter value; the layer computation module, based on the layer parameter value, regenerates activation function numerical information with equivalent time as the independent variable and sends it to the neuron output module of the corresponding layer.

[0050] This invention also proposes a multi-segment neuron output module that can convert the single-symbol computation signal input from the differential capacitance control module of the corresponding segment into the value of the fast access unit.

[0051] The aforementioned neuron aggregation calculation output module can perform shifting, addition, and subtraction operations on the values ​​of the fast access units of each segment neuron output module according to the corresponding bit range of its segment, and write the result value into its fast access unit.

[0052] The present invention also proposes a detection process in which a differential system has a differential comparator for detecting the voltage of the capacitor, the result of which generates an enable signal to control the operating state of other circuits.

[0053] The present invention also proposes that the capacitor control module has a comparator that uses a piecewise function to divide the reference voltage at the point of division. The comparison result divides the discharge detection process, and then generates pulses corresponding to different activation functions, and finally outputs them.

[0054] This invention also proposes that each basic neuron unit has a bias link; the input term of the bias link is set to full ratio, and the charge adjustment module adjusts the value controlled by the weight term fast access unit.

[0055] This invention also proposes that the capacitor voltage of the capacitor module is compared with a reference voltage to generate comparison detection signals P and N; the neuron output module logically synthesizes the comparison detection signals P and N, and generates an XOR gate signal based on the comparison detection signals P and N, i.e., whether the comparison detection signals P and N are at the same level; the neuron output module generates a positive or negative signal SGN when the XOR gate signal starts, i.e., when P and N become unequal, based on the comparison detection signals P and N; if the duration of the comparison detection signal P is longer, SGN is positive, otherwise it is negative.

[0056] The present invention also proposes one or more underlying reference pulse signals; the input current modulation control module has a pulse selection or shielding circuit, which selects or shields the underlying high-frequency reference pulse signal according to the input current modulation control signal to generate a short pulse input current modulation control signal, or selects one of the underlying high-frequency reference pulse signals.

[0057] This invention also proposes a system with multiple neurons, each having a shared weight term fast access unit within its charge adjustment module; each neuron's input term current modulation control module generates multiple sets of modulation signals based on the shared weight term fast access unit and its own input value; the charge adjustment module of the neuron's link adjusts the charge of the capacitor module based on the multiple sets of modulation signals; multiple neuron basic units work simultaneously or sequentially.

[0058] The present invention also proposes multiple sets of phase-staggered, non-overlapping underlying reference pulse signals; the input current modulation control module of each neuron generates multiple sets of current modulation signals according to the phase-staggered underlying reference pulse signals and its own input value; the switching of the link current channel of the neuron generates multiple sets of modulation current according to the multiple sets of current modulation signals.

[0059] The present invention also proposes that the charge adjustment module has a backup fast access unit; the output value of the neuron output module is preloaded into the backup fast access unit in blocks or batches; the charge adjustment module has a switching circuit that can switch the effective weight value to the value of the backup fast access unit, or switch the effective weight value to the weight value originally loaded in that layer of the neural network; the pre-generated output value of the neuron output module can be directly used as a real-time input value through the backup fast access unit.

[0060] This invention also proposes a method for mini-batch training: In mini-batch training, each stored data, including data from forward propagation and data from the same node, is written sequentially in adjacent rows; Step 1, during forward propagation, the array-like storage caches the output of each layer of the neural network after activation, row by row; Step 2, during backward propagation, the array-like storage caches the product of the partial derivatives of the output nodes after activation and the derivatives of the activation functions of each layer of the neural network, i.e., the multiplicative partial derivatives before activation; Step 3, the block-based reading circuit reads the multiplicative partial derivatives of each backward propagation node in the mini-batch corresponding to the block into the layer's input, and reads the multiple output values ​​of each neuron's output node from the forward propagation of the previous layer into the multiple connection weights of one neuron corresponding to the current calculation circuit; Step 4, the multiplicative sum is calculated and multiplied by the learning rate, the product being the adjustment value for the corresponding connection weight; Step 5, the corresponding connection weight is updated.

[0061] This invention also proposes that neurons share a common computational core and a common memory; the layer activation function source writes activation function information into the common memory; the computational core obtains activation function values ​​from the common memory through table lookup, copying, and calculation; and there is a method for synthesizing multiple rounds of output values: Step 1, perform forward propagation to obtain the level signal of the differential multiplication and addition information expressing the charging and discharging of the capacitor and the level signal of the symbol expressing the energy relationship of each capacitor; Step 2, update the parameters of the layer activation function value source, and continuously output the time value, the derivative change value of the activation function over time, and the current value to the neuron output module, or only output the time value and obtain other values ​​through table lookup; Step 3, if the activation function needs to be used, obtain the activation function value of each neuron through interpolation calculation of the derivative change value and the current value; Step 4, if it is necessary to realize the addition of layer output values, do not reset the output item fast access unit or the activation function fast access unit, and add the current activation function value with the previous round activation function value, repeating steps 1, 2, 3, and 4; Step 5, obtain the final neuron output value.

[0062] This invention also proposes a method for normalizing layer output values: Step 1, the network summarizing layer neurons use the values ​​of the activation function fast access units of the current main computing layer as input to the network summarizing layer neurons; Step 2, the network summarizing layer neurons calculate and generate layer parameter values, i.e., the process of powering on, detecting, and activating; Step 3, the layer parameter values ​​are given to the layer activation function value calculation chip, and the layer activation function value source regenerates normalized activation function value information with equivalent time as the independent variable based on the layer parameter values; Step 4, the activation function value information with equivalent time as the independent variable is sent to the neuron output module of the current main computing layer; Step 5, the current main computing layer regenerates the normalized activation function value based on the original activation function value and stores it in the activation function fast access unit.

[0063] This invention also proposes a backup fast access unit for the weighted current modulation control module; based on a fully connected neural network, a method for generating a self-attention QK matrix and performing multiplication is provided: Step 1, read the Q matrix parameters from an external source or storage to the fast access unit of the input current modulation control module; Step 2, in parallel, read the values ​​of the input vector to the fast access unit of the weighted current modulation control module; Step 3, perform forward propagation calculation, and the neuron output module generates a round of Q matrix result values; Step 4, copy the result values ​​from the neuron output module to the backup fast access unit of the corresponding weighted current modulation control module; Step 5, repeat steps 2, 3, and 4 until each weighted term... Step 6: The spare fast access units of the current modulation control module are filled or all Q matrix parameters are used; Step 7: Read the K matrix parameters from external storage, memory, or spare fast access units to the fast access units of the input current modulation control module; Step 8: In parallel, read the values ​​of the input vectors to the fast access units of the weight current modulation control module; Step 9: Perform forward propagation calculation, and the neuron output module generates a round of K matrix result values; Step 10: The neuron output module transfers the K matrix result values ​​to the fast access units of the input current modulation control module; Step 11: The main fast access units and spare fast access units of each weight current modulation control module are exchanged, that is, the result values ​​of the Q matrix are used; Step 12: Perform forward propagation calculation, and the neuron output module generates a round of QK matrix multiplication result values, and saves the round of QK matrix result values ​​to external storage, memory, or other idle spare fast access units; Step 13: Repeat steps 6, 7, 8, 9, 10, and 11 until all K matrices have been calculated.

[0064] Based on the calculation results from the previous steps, there are methods to generate self-attention QKV results:

[0065] Step 13: The QK matrix multiplication result is normalized and activated, then stored in another idle spare fast access unit of the weight term current modulation control module. This process is repeated multiple times until all QK activation result values ​​are obtained. Step 14: The V matrix parameters are read from external storage, memory, or a spare fast access unit and stored in the fast access unit of the weight term current modulation control module. Step 15: In parallel, the QK matrix multiplication result value for one round, i.e., one round of attention weights, is read and stored in the fast access unit of the input term current modulation control module. Step 16: Forward propagation calculation and activation are performed. The neuron output module generates a round of QKV matrix multiplication result value and saves the QKV matrix result value to external storage or memory. Step 17: Steps 14, 15, and 16 are repeated until all QKV matrices have been calculated. Step 18: If the input vector and layer input are inconsistent, the input vector is split and grouped. The final result is obtained through a combination of multiple groups and rounds of calculation.

[0066] If position encoding is required in the above steps, the data can be modified by external calculation to insert position encoding, or the vector after inserting position encoding can be generated by setting the bias term weight parameters of the weight term current modulation control module.

[0067] This invention also proposes a differential capacitor control module with a capacitance value detection circuit; the capacitance value detection circuit has multiple small capacitors with initial voltage of zero that can be selectively connected to the main capacitor of the capacitor module of the same attribute by a switch; a digital logic circuit of the same attribute is connected to the capacitance value detection circuit and the voltage comparator of the capacitor control module, and outputs a digital output of the capacitance value generated by the capacitance value change operation through the capacitance value detection circuit based on the result of comparing the main capacitor of the capacitor module with the target voltage; the neuron output module is configured to generate the differential multiplication and addition information and the sign information based on the digital output of the differential capacitor control module.

[0068] This invention also proposes that the forward-changing capacitors in the capacitance detection circuit have multiple small forward-changing capacitors with unequal capacitance values ​​and a ratio that is a multiple of 2, forming a binary capacitor network; the differential capacitor control module has a capacitance detection inverse-changing circuit; the capacitance detection inverse-changing circuit has multiple inverse-changing small capacitors whose initial voltage is the same as the initial voltage of the main capacitor of the capacitor module, and which can be selectively connected to the main capacitor of the capacitor module with the same attribute by a switch; after the capacitance detection forward-changing circuit operates the non-minimum bit of the forward-changing small capacitor, when one of the voltage comparators in the two sets of differential capacitor control modules is lower than the target reference voltage, the two sets of differential capacitance detection inverse-changing circuits synchronously operate the connected inverse-changing small capacitors, thereby increasing the voltage of the differential capacitor module.

[0069] This invention also proposes a detection output method by changing the capacitance value of the main capacitor of the capacitor module: Step 1. The differential system charge adjustment module selects the main capacitor to be connected to the positive or negative capacitor module based on the values ​​of the neuron input and weights, and the sign of their product, and adjusts the capacitance value of the main capacitor of the capacitor module. Step 2. The voltage of the main capacitor of the capacitor module and the voltage of the small capacitor in the detection capacitance inverting circuit are initialized to the same voltage, and the voltage of the small capacitor in the detection capacitance forward inverting circuit is initialized to 0. Step 3. The grounding switch of the detection capacitance forward inverting circuit is turned off, the switch of the initialization voltage source of the detection capacitance inverting circuit is turned off, and the switch of the initialization voltage source of the main capacitor of the capacitor module is turned off. Step 4. If the result of this step can be predicted, the forward converter circuit replaces the small capacitor with the corresponding capacitance value in the lower bit according to the prediction; the forward converter small capacitor with the current corresponding capacitance value is connected in parallel; the voltage comparator judges whether the main capacitor voltage is greater than the target reference voltage, and repeats this step; Step 5. If the capacitance value operation does not reach the minimum bit, the reverse converter circuit operation is executed, and the reverse converter small capacitor is connected in parallel to the main capacitors of the two differential positive and negative systems, and the reverse converter operation is repeated to raise its voltage to be greater than the target reference voltage; the forward converter circuit replaces the small capacitor with the corresponding capacitance value in the lower bit; jump back to step 4; Step 6. The differential system outputs the digital code of the corresponding capacitance value respectively; Step 7. The neuron calculates and outputs the difference value according to the output of the differential system.

[0070] This invention also proposes a computational method using forward propagation of a neural network: Step 1, selecting a charging / discharging mode based on the computational mode of the neural network layer, and selecting a specific initial voltage and a driving voltage for charging / discharging based on the charging / discharging mode; Step 2, pre-charging / discharging the capacitor module to the specific initial voltage; Step 3, generating a modulation current through a modulation control signal based on the correction values ​​of the input term fast access unit and the weight term fast access unit; Step 4, charging / discharging the capacitor module through the modulation current; Step 5, continuing to charge / discharge the capacitor module through the capacitor control module; Step 6, comparing the voltage of the capacitor module with a reference voltage or other capacitor voltages, and outputting a detection signal; Step 7, combining the positive and negative detection signals from the capacitor control module to obtain information on voltage comparison detection between the positive and negative capacitor modules. Beneficial effects

[0071] In view of the above challenges, this invention proposes a basic structure for a large-scale neural network suitable for parallel execution of digital and analog mixed circuits, as well as the structure, implementation, calculation principle and working method of digital and analog mixed circuits, and the parameter adjustment and setting method in its circuit system. More importantly, it proposes a heterogeneous construction method for supercomputer clusters.

[0072] This technology lays the technological foundation for future large-scale neural network hardware, pointing to a direction for software and hardware technologies that far surpass the human brain. It represents a major breakthrough with historical influence, achieving a long-held goal in the field. It combines existing chip technologies (such as CMOS), thin-film transistor computing, and non-volatile storage technology, solving the long-standing problem of efficiently implementing parallel ultra-large-scale neural network computing using analog circuits, far exceeding traditional analog and digital solutions. Design simplification and manufacturing friendliness: simple, practical, parallelizable, and easily standardized for manufacturing; high tolerance for defects, easy expansion and partitioning, significantly reducing chip process requirements. Huge performance leap: achieving optimizations of 10^n times across multiple dimensions (especially energy efficiency). It possesses high computing power per unit area, high energy efficiency, efficient storage read / write, and configurable high-precision calculations, making it suitable for neural network training with great potential. Profound impact: enabling large models to operate independently of servers / GPUs, embedding core intelligence into terminal devices, ushering in an era of "AI everywhere"; providing a systematic framework for overcoming conventional software and hardware technology biases and developing next-generation AI software and hardware. Attached Figure Description

[0073] Figure 1-2 is a schematic diagram of the wafer structure; Figure 3-4 is a schematic diagram of the stacked and bonded wafer layer structure; Figure 5 is a schematic diagram of wafer communication data flow; Figures 6 and 7 are circuit schematic diagrams; Figures 8-15 are circuit schematic diagrams; Figures 16 and 17 are square wave waveform diagrams; Figures 18-20 are circuit schematic diagrams; Figure 21 is a function diagram; Figures 22-24 are data relationship diagrams; Figure 25 is a circuit diagram; Figures 26 and 27 are layer connection diagrams; Figure 28 is a function diagram; Figures 29-30 are circuit schematic diagrams; Figure 31 is a data flow diagram; Figure 32 is a functional layout diagram; Figures 33-34 are circuit schematic diagrams. The best embodiment of the present invention

[0074] The following are examples of nonlinear activation functions for more complex analog outputs. Embodiments of the present invention

[0075] Unless otherwise specified, the following embodiments are basically parallel operations, that is, a large number of basic neuron units are computed in parallel, and a large number of links within a single neuron unit are also computed in parallel.

[0076] *Simplest Implementation* Figure 8 shows a hardware implementation of a neural network with two input units in the first layer and only one neuron in the last layer. The capacitor-time multiplication-accumulation calculation circuit includes a charge adjustment module, capacitor modules (7)(22), capacitor control modules (8)(16), and neuron output module (14). The charge adjustment module has two forms: constant capacitance and variable capacitance. The constant capacitance circuit realizes the input of the neuron circuit by adjusting the voltage of a fixed capacitance capacitor, while the variable capacitance circuit realizes the input by adjusting the capacitance value under a specific voltage. This embodiment is a constant capacitance circuit. The (two) charge adjustment modules include a network link input current modulation control module (21)(19) and a weight current modulation control module (5)(6)(17)(18). The number of charge adjustment modules corresponds to the number of links of the neuron to be calculated (if the links are not divided into multiple groups).

[0077] The network link input current modulation control module and the weighted current modulation control module jointly control the charging and discharging current of the capacitor module. In this embodiment, the network link input current modulation control module is a circuit module that dynamically controls the current and is used to generate the modulation current. (Known forms of modulation current include, but are not limited to: pulse width modulation (PWM), pulse density modulation (PDM), pulse position modulation (PPM), single pulse level, etc.; the modulation current is converted into charging and discharging current by the control signal through the switching circuit of the current channel, and its implementation includes direct clock generation, multi-channel signal mixing, reference signal trimming or delay insertion, etc. The current waveform also includes, but is not limited to, centrally symmetrical or edge-aligned, rectangular wave, triangular wave or sine wave, etc. The modulation current in other embodiments is the same.) The modulation current plus the switching of the internal resistance of the weighted current modulation control module (5)(6)(17)(18) realizes the control of the current. In this embodiment, the weighted current modulation control module is a small uniform resistor. In other embodiments, it can be a complex resistor network or an equivalent circuit. In this embodiment, the effect of the current flowing into the capacitor controlled by the network link input current modulation control module and the weight current modulation control module is equivalent to the link parameter w * input value x in the neural network. The network link input current modulation control module contains a fast access unit (the form of the fast access unit includes, but is not limited to: register, register file, flip-flop, latch, various SRAM / DRAM cells, array, and various new storage technology access units. Other embodiments are the same), which is set by the control information channel (1) (the control information channel and related fast access units also include functions such as enable, capacitor link port high impedance control, positive and negative setting of link parameter w and input value x, etc.), and generates current modulation (PWM, etc.) by comparing with the timing information of the timing information channel (2). Here, a part (bits) of the current modulation fast access unit represents the network input value x, and a part represents the link parameter w (which works together with the weight current modulation control module). The driving voltage channel (3) is linked to a dynamically selectable capacitor charging and discharging driving voltage source shared by the entire neural network layer. For example, the driving voltage a = 0 volts during discharge and the driving voltage a = 2 volts during charging. The specific voltage value depends on the actual needs; this is just for ease of description. The above control information channels, timing information channels, and even resistance / current channels are not limited to a single physical power-on link. The specific number of physical power-on links is determined according to the complexity of the data and actual needs.

[0078] (As is well known, a timer has multiple channels, and each channel has a fast access unit. Comparing the current with the timer's fast access unit generates multiple corresponding current modulations, which is a basic function of a microcontroller. In addition to being driven by a crystal oscillator clock, a timer can also be implemented using a delay chain (such as an inverter).)

[0079] In this embodiment, the current is basically controlled by current modulation and a standard resistor network. In addition to expressing x, current modulation is also used to express w. Sometimes, in order to save chip area, some embodiments (specific value resistor embodiments) use different specific resistor values ​​to express the link parameter w, or even use a specifically generated stable voltage source to replace current modulation.

[0080] The capacitor modules of the last layer neurons are divided into positive and negative capacitor modules. In this embodiment, the left capacitor module represents a positive value (7), and the right capacitor module represents a negative value (22). The network link input current modulation control module will select to charge or discharge the positive or negative capacitor module according to the sign bit of the fast access unit set by the internal link parameter w and the sign bit of the fast access unit set by the input value x. That is, it determines which capacitor module to charge or discharge according to the sign of (w*x).

[0081] The capacitor control modules (8) and (16) correspond to their respective capacitor modules. During the charging and discharging process of the capacitor modules, the capacitor modules are pre-charged and discharged through the initial voltage source channel (9). In the charging mode, the capacitor modules are pre-charged to the initial voltage (in this embodiment, the initial voltage b = 1 volt); in the discharging mode, the capacitor modules are pre-charged to the initial voltage (in this embodiment, the initial voltage b = 2 volts). In the current embodiment, the initial voltage source can be set to an initial voltage b = 1 volt or b = 2 volts through the layer parameters of the neural network (1 or 2 volts is just an example reference, and the specific voltage value should be set according to the actual situation). If the neural network is more complex and there are different activation functions in one layer, or other more complex requirements, then multiple initial voltage source channels (9) with different voltages are needed. After the pre-charging and discharging steps, the input current modulation control module and the weight current modulation control module are linked through the network, and the neural network is charged and discharged, which is the actual working process of the neural network. After the neural network charging and discharging steps, the capacitor discharge detection begins. The capacitor module discharges to GND through the standard resistor R inside the capacitor control module until the detection voltage c = 1 volt (the specific voltage value is set according to the actual situation; in this embodiment, it is 1 volt). The capacitor control module contains a voltage comparator. When the voltage of the discharged capacitor module is lower than the detection voltage c, the output port (15) of the capacitor control module flips, displaying a level representing the comparison result (the level is set as needed; in this embodiment, it is set to high). All operations of the capacitor control module are manipulated by its own capacitor control channel (23). (Voltage comparison is a basic function of the chip, and its circuitry is common knowledge.)

[0082] The neuron output module (14) logically merges the output signals of the capacitor control modules representing positive and negative values ​​(here, the symbol P / N represents the output of the positive and negative capacitor control modules, i.e., the positive module output is P and the negative module output is N). The layer output (13) of the neuron output module (14) represents the information of the duration difference of the discharge. In this embodiment, its information is logically equal to the level duration of P xor N. The layer output (12) of the neuron output module (14) represents the comparison of the magnitude of the positive and negative capacitor voltages or the comparison of the magnitude of the positive and negative capacitor discharge durations. In this embodiment, it is logically equal to the level of P output by the positive module before P xor N flips last. (It is known that logical operations such as xor are basic functions of digital circuits.)

[0083] Note that in this embodiment, what needs to be output is the duration difference information, not just the duration difference level itself. The duration difference information has multiple forms of expression. Although this embodiment outputs the duration difference level for processing by the microcontroller and other modules, some embodiments (such as the current modulation layer output embodiment) require converting the duration difference information into a current modulation signal for output. That is, the neuron output module (14) has a time-to-digital converter circuit for detecting the duration difference. The layer output control signal channel (11) controls the time-to-digital converter circuit of the neuron output module to capture the Pxor N and obtain its duration through the timing information input through the timing information channel (10) and store it in the internal fast access unit. Finally, based on the value of the internal fast access unit and the timing information input through the timing information channel (10), the current modulation signal is output. (It is known that detecting the square wave pulse width is a basic function of the microcontroller. At low frequencies, a counter can be used to count clock pulses, and at high frequencies, a delay chain signal / multi-phase multi-line signal is used.)

[0084] The following are the steps of the working method of the analog-digital hybrid neural network circuit in this embodiment (some steps can be performed in parallel according to actual needs, and the specific voltage values ​​are set according to actual needs):

[0085] 1. Initialization settings: Set the fast access unit values ​​of the network link input current modulation control module, including enable, high impedance control of capacitor link port, positive and negative settings of link parameter w and input value x, etc.

[0086] 2. Select appropriate resistance values ​​or resistor networks as needed.

[0087] 3. Generate an appropriate current modulation signal.

[0088] 4. Capacitor pre-charge and discharge: Use the initial voltage source channel (9) to pre-charge and discharge the capacitor module (7)(22) to a specific initial voltage (charging mode b=1 volt, discharging mode b=2 volts). Pre-charge and discharge can be achieved by directly connecting to the initial voltage source or by detection through the voltage comparator (of the capacitor control module).

[0089] 5. Perform charging and discharging. Based on the positive and negative values ​​of w and x set internally, select to perform charging and discharging operations on the positive capacitor module (7) and the negative capacitor module (22).

[0090] 6. The charging and discharging current of the capacitor module (7)(22) is jointly controlled by the network-linked input current modulation control module (21)(19) and the weighted current modulation control module (4)(20)(5)(6)(17)(18).

[0091] 7. After the charging and discharging process, during detection, the charging and discharging continues until the detection voltage is reached. (Regarding energy efficiency, the detection process can be further optimized by setting multiple reference comparison voltages and selecting different resistance values ​​for the detection discharge resistor based on the voltage range of the positive and negative capacitors during detection, thereby accelerating the detection discharge speed. Because the result is a difference and the positive and negative values ​​switch synchronously, it does not affect the final result.)

[0092] 8. Output detection result information / signal.

[0093] 9. Layer output processing: The neuron output module (14) combines the output signals (P and N) of the positive and negative capacitance control module and uses the time difference information to represent the final output. The time difference information and sign are the level and signal pulse, or the signed digital value after further time-to-digital conversion.

[0094] 10. In some embodiments, it may be necessary to convert this duration difference into a current-modulated output or other form of electrical information, which involves capturing the level duration of P xor N and outputting a current-modulated signal based on this information.

[0095] As is well known, the conversion from time pulse to digital quantity described in step 9 can be achieved through a time-to-digital converter (TDC). Specific implementations include, but are not limited to, counters, interpolators, inverter delay chains, vernier signal structures, time amplifier circuits, etc. In the case of a multi-layer neural network, if the input current modulation uses a single-pulse level, then the time pulse can be directly used as the input to the next layer of the neural network, or the entire network can use single pulses. Note that mathematically, it can be verified that if the PWM duty cycle variable in the derivation below is replaced with the pulse width of a single pulse, i.e., the pulse duration, the resulting formula, even in the counterintuitive case where the single pulses are not aligned in time, still holds true. That is, the output time difference is proportional to the product of the total pulse width and conductance of each input (the time integral of the total conductance), which still holds true due to the exponential multiplication effect. The capacitor discharge process is an exponential multiplication, with the discharge voltage V(t) = V0 * e^(-t / (RC)). e^a * e^b = e^(a+b), which is independent of the order, length, or even overlap of the single pulses corresponding to a or b. (That is, it only depends on the time integral of the total positive / negative conductance; the derivation is omitted). However, appropriate pulse repetition in PWM can average out various uncontrollable factors such as interference and inconsistencies, although it also increases power consumption. Single pulses, PWM, capacitor charge transfer, or other more complex time- and conductance-based similar forms are essentially all about the ratio of conduction time and current amplitude between various links and detection channels. The final result can be obtained by applying similar formulas to achieve the proportional relationship.

[0096] The following is a computational description of the working principle of the aforementioned hardware, and an explanation of how to generate the connection parameters w and input value x of the hardware. The calculations below are based on ideal components, neglecting leakage current, voltage and temperature variations, and other interference factors. Therefore, the calculations are approximate results. In actual neural network training, if a test statistical characteristic table of the hardware circuit is needed, the software layer will use a lookup table and interpolation method to obtain the actual input / weight / output values / derivatives of the circuit. The following calculations use the simplest PWM current modulation form as an example, but its essence is to control the current ratio of each current channel through switching; therefore, other forms of modulation are equivalent. Furthermore, in the following calculations, the pulse width, resistance, and capacitance do not require precise values; what is needed are precise and stable ratios.

[0097] The value x represents the duty cycle of the input portion in current modulation (PWM pulse width modulation; in single-pulse modulation, it's the pulse duration width, i.e., the ratio to the minimum pulse duration width). The duty cycle affects the voltage difference between the equivalent input voltage source and the capacitor voltage. Current input x = Xn = pulseXn

[0098] The value w is the duty cycle of the portion representing the equivalent resistance in current modulation (pulseRn; in single-pulse modulation, this can be omitted and set to 1; if this weighted item exists in single-pulse modulation, it represents the scaling factor of the pulse duration or current intensity), divided by the corresponding resistance Rn and capacitance cap. Current weighted current limiting w = Wn = pulseRn / Rn / cap

[0099] SSP(...), SumSelectPositive means selecting all values ​​greater than or equal to 0, summing them, and taking the absolute value.

[0100] SSN(...), SumSelectNegative means selecting all values ​​less than 0, summing them, and taking the absolute value.

[0101] V[t] represents the time function of the capacitor voltage, V'[t] is its derivative, and e is the natural constant.

[0102] 1.1. In discharge mode, the initial voltage is 'a', and the discharge drive voltage is GND voltage 0. After the neural network discharges, it enters the detection process to continue discharging, with the stop / detection voltage c=b=1. The discharge resistor used in the detection process is R.

[0103] t is a time variable, which is a standard time length that can be set and controlled.

[0104] Solve the difference equations respectively (b <= V[t] <= a).

[0105] Dsolve[{V'[t]==SSP(...,Xn*(-V[t])*Wn), V[0]==a}, {V[t]}, t]

[0106] DSolve[{V'[t]==SSN(...,Xn*(-V[t])*Wn), V[0]==a}, {V[t]}, t]

[0107] Solution results

[0108] The positive capacitance V[t] is given by gp1 = a * e^(-t*SSP(...,Xn*Wn)).

[0109] The negative capacitance V[t] is given by gn1 = a * e^(-t*SSN(...,Xn*Wn)).

[0110] During detection, the positive capacitor discharge time function tp1 = cap*R*ln[gp1 / b]=cap*R*(ln[gp1]-ln[b])

[0111] During detection, the discharge time function of the negative capacitor is tn1 = cap*R*ln[gn1 / b]=cap*R*(ln[gn1]-ln[b]).

[0112] The difference between the two, tp1-tn1, is calculated according to the logarithmic rule.

[0113] ln[a * e^(-t*SSP(...,Xn*Wn))]=ln[a]+ln[e^(-t*SSP(...,Xn*Wn)] = ln[a]-t*SSP(...,Xn*Wn)

[0114] ln[a * e^(-t*SSN(...,Xn*Wn))]=ln[a]+ln[e^(-t*SSN(...,Xn*Wn)] = ln[a]-t*SSN(...,Xn*Wn)

[0115] tp1 = cap*R*(ln[a]-t*SSP(...,Xn*Wn)-ln[b])

[0116] tn1 = cap*R*(ln[a]-t*SSN(...,Xn*Wn)-ln[b])

[0117] tp1-tn1 = -cap*R*(X1*W1+X2*W2+...+Xn*Wn)*t

[0118] As can be seen from tp1 and tn1, its activation function is a linear function. Substituting this into Wn = pulseRn / Rn / cap...

[0119] tp1-tn1 = -R*(X1*pulseR1 / R1+X2*pulseR2 / R2+...+Xn*pulseRn / Rn)*t

[0120] Here, R and R1...Rn are all proportional. The value of pulseR1*R / R1 is the value of the weighting term. If all resistors are standard resistors, that is, R1...Rn equals R, then...

[0121] tp1-tn1 = -(X1*pulseR1+X2*pulseR2+...+Xn*pulseRn)*t

[0122] If we flip the positive and negative signals of the neuron's output module, then the result of tp1-tn1 is...

[0123] R*(X1*pulseR1 / R1+X2*pulseR2 / R2+...+Xn*pulseRn / Rn)*t

[0124] (X1*pulseR1+X2*pulseR2+...+Xn*pulseRn)*t,

[0125] In simple terms, if the modulated current is generated using PWM, the total duty cycle is duty_n = pulseXn * pulseRn. The input value Xn = t * pulseXn, representing the duty cycle of the input portion. The weight value Qn = pulseRn * R / Rn, where pulseRn is the duty cycle representing the weight portion of the neuron's connection, Rn is the resistance of the weight term, and R is the standard resistor for detecting discharge. In essence, the calculation is achieved by adjusting the ratio of the standard resistor to the weight term resistance, the duty cycle of the weight term and the input term, and the total discharge time to make them equivalent to the input and weight values ​​in the software.

[0126] Additionally, there is a time t. If t is not the base time, assuming t=8, the final discharge detection time or input value should be subtracted from the time t. Since t=8 is a multiple of 2, the discharge detection time or input value can be directly shifted by 3 bits.

[0127] If there are deviations in the production process, resulting in one capacitor having a larger capacitance than the other (cap1 and cap2), then the above formula becomes...

[0128] tp1 = cap1*R*(ln[a]-t*SSP(...,Xn*Wn)-ln[b])

[0129] tn1 = cap2*R*(ln[a]-t*SSN(...,Xn*Wn)-ln[b])

[0130] tp1-tn1 =R*(cap1-cap2)*ln(a / b)-R*(X1*pulseR1 / R1+X2*pulseR2 / R2+...+Xn*pulseRn / Rn)*t

[0131] The calculation error caused by the capacitance difference: err_cap = R * (cap1 - cap2) * ln(a / b)

[0132] Because the capacitance difference is an independent term, its error value can be easily obtained through calculations with a 0% duty cycle. For digital circuits (with digital-to-analog conversion), this error can be subtracted from the final calculation result. For analog circuits (without analog-to-analog conversion), the error can be corrected by adjusting the neuron bias term connections or by setting additional bias term connections.

[0133] The standard resistor R and other resistors R1, R2...Rn may have manufacturing errors. However, we only use the resistance ratio R / Rn here, which is relatively accurate. In addition, manufacturing errors, operating temperature, voltage, line inductance, frequency, leakage current, and switching delay will all affect the calculation results. Various errors can be adjusted using appropriate fine-tuning circuits or by inserting a delay into the PWM. These adjustment circuits and how to reduce manufacturing errors are very specific design tasks, which will not be discussed in detail here. A relatively simple method for adjusting the weights is to obtain the average deviation of the weights for each link through multiple calculations based on specific parameters (such as 0), and then correct the final result by modifying the ideal weight values ​​(before using the weights).

[0134] 1.2. If the detection process begins, and the device recharges from gn1 or gp1 to a=2, the stop / detection voltage is b=a. The discharge resistor used in the detection process is R. The charging drive voltage is c=a+1, then...

[0135] The positive capacitor discharge time function during detection is tp3 = cap*R*ln[(c-gp1) / (ca)] ; tp3 = cap*R*ln[3 - 2 * e^(-t*SSP(...,Xn*Wn))]

[0136] The discharge time function of the negative capacitor during detection is tn3 = cap*R*ln[(c-gn1) / (ca)] ; tn3 = cap*R*ln[3 - 2 * e^(-t*SSN(...,Xn*Wn))]

[0137] It is evident that tp3 and tn3 are some kind of nonlinear functions.

[0138] 2.1. Now returning to the charging mode in the simplest embodiment, the initial voltage is b=1, the charging drive voltage is a=2, and after the neural network charging is complete, it enters the detection process to begin discharging to ground, with the stop / detection voltage c=b=1. The discharge resistor used in the detection process is Rt, which is a time variable.

[0139] Mathematical software solves the difference equations (b <= V[t] <= a).

[0140] DSolve[{V'[t]==SSP(...,Xn*(aV[t])*Wn), V[0]==b}, {V[t]}, t]

[0141] DSolve[{V'[t]==SSN(...,Xn*(aV[t])*Wn), V[0]==b}, {V[t]}, t]

[0142] Solution results

[0143] For a positive capacitor, V[t] is given by gp2 = a+(ba) * e^(-t*SSP(...,Xn*Wn)); for a negative capacitor, V[t] is given by gn2 = a+(ba) * e^(-t*SSN(...,Xn*Wn)).

[0144] The discharge time function for the positive capacitor during detection is tp2 = cap*R*ln[gp2 / b]; the discharge time function for the negative capacitor during detection is tn2 = cap*R*ln[gn2 / b].

[0145] The difference between the two, tp2-tn2, can be substituted into b=1 and a=2.

[0146] tp2 = cap*R*(ln[2-e^(-t*SSP(...,Xn*Wn))]); tn2 = cap*R*(ln[2-e^(-t*SSN(...,Xn*Wn))])

[0147] tp2-tn2 = cap*R*(ln[2-e^(-t*SSP(...,Xn*Wn))] - ln[2-e^(-t*SSN(...,Xn*Wn))])

[0148] As can be seen from tp2 and tn2, their activation function is the difference between two nonlinear functions tp2 and tn2. The curves of tp2 and tn2 in the first quadrant are similar in shape but have better training performance than the cap*R*ln(2)*tanh function. tanh is a mature neural network activation function with good performance. The maximum values ​​of tp2 and tn2 are cap*R*ln(2).

[0149] In other words, when nonlinear activation is required in this embodiment, the activation function used is to group the Xn*Wn groups according to their positive and negative signs, sum the sums of each group, take the absolute value (val), and then perform a nonlinear transformation of cap*R*ln(2-e^(-t*val)), and then directly take the value or take the difference between the two groups after the nonlinear transformation.

[0150] 2.2. If the detection process begins, the capacitor is recharged from gn2 or gp2 to a=2, and the stop / detection voltage a is reached. The charging resistor used during the detection charging process is R. The charging drive voltage is c=a+k, k=1. Then, the positive capacitor charging time function during detection is tp4 = cap*R*ln[(c-gp2) / (ca)] = cap*R*ln[a+k - (a+(ba) * e^(-t*SSP(...,Xn*Wn)))] = cap*R*ln[1 + (ba) / k * e^(-t*SSP(...,Xn*Wn))] = cap*R*ln[1 + e^(-t*SSP(...,Xn*Wn))]

[0151] The charging time function of the negative capacitor during detection is tn4 = cap*R*ln[(c-gn2) / (ca)] = cap*R*ln[ a+k - (a+(ba) * e^(-t*SSN(...,Xn*Wn)))] = cap*R*ln[ 1 + (ba) / k * e^(-t*SSN(...,Xn*Wn))] = cap*R*ln[ 1 + e^(-t*SSN(...,Xn*Wn))]

[0152] As can be seen from tp4 and tn4, their activation function is the difference between the same two nonlinear functions tp4 and tn4.

[0153] 3.0. Example of Positive and Negative Capacitor Detection and Comparison* Similarly, the capacitor is first pre-charged to a predetermined voltage 'a', then the network discharges for forward propagation of the neural network, and then the detection process is executed. The voltage comparator directly compares the voltages of the positive and negative capacitors. The capacitor with the larger voltage discharges to ground through a standard resistor R during the detection process until the voltages of the two capacitors are equal. The discharge time is the time difference. Its sign represents the result of comparing the magnitudes of the positive and negative capacitors at the start of detection. Its unsigned value is consistent with |tp1-tn1|, and is also a linear function. However, because the voltage comparator has an offset voltage, and various leakage currents may be inconsistent, this will affect the accuracy and efficiency of the calculation results.

[0154] 3.1. Example of Dual-Capacitor Detection and Comparison* (Each capacitor has positive and negative terminals, i.e., 4 capacitors) Here, we supplement another example of detection and discharge comparison (dual-capacitor detection and comparison), where an identical capacitor is added to the capacitor module. This capacitor does not participate in the charging and discharging of the neural network; it only plays a role during detection (the network discharges, and detection also discharges). Before detection, this capacitor is pre-charged to 'a', and then it discharges to ground (0 voltage) through 'R'. The stop / detection voltage is set to the voltages of the capacitors in the capacitor module that participate in the charging and discharging of the neural network, i.e., gp and gn, i.e., the voltages of two capacitors with the same sign are compared using analog voltage comparison.

[0155] During detection, the positive capacitor discharge time function tp0 = cap*R*ln[a / gp] = cap*R*(t*SSP(...,Xn*Wn))

[0156] During detection, the discharge time function of the negative capacitor is tn0 = cap*R*ln[a / gn] = cap*R*(t*SSN(...,Xn*Wn)).

[0157] diff0 = tp0 - tn0 = cap*R*(X1*W1+X2*W2+...+Xn*Wn)*t; From tp0 and tn0, it can be seen that its activation function is also a linear function.

[0158] 3.2 In the dual-capacitor detection process (where the network is charging and detection is also charging), the capacitor is pre-charged to b=1V before detection, and then charged through R with a driving voltage of a=2V (the specific voltage value depends on the actual setting; this is only for reference). The stop / detection voltage is set to the voltages of the capacitors participating in the neural network charging and discharging in the capacitor module, i.e., gp2 and gn2, that is, the voltages of two capacitors with the same sign are compared using analog voltage comparison.

[0159] During testing, the positive capacitor charging time function is tp5 = cap*R*ln[(a-gp2) / (ab)].

[0160] = cap*R*ln[(a-(a+(ba) * e^(-t*SSP(...,Xn*Wn)))) / (ab)] =cap*R*ln[e^(-t*SSP(...,Xn*Wn))]

[0161] The charging time function of the negative capacitor during detection is tn5 = cap*R*ln[(a-gn2) / (ab)].

[0162] = cap*R*ln[(a-(a+(ba) * e^(-t*SSN(...,Xn*Wn)))) / (ab)] = cap*R*ln[e^(-t*SSN(...,Xn*Wn))]

[0163] tp5-tn5 = -cap*R*(X1*W1+X2*W2+...+Xn*Wn)*t

[0164] As can be seen from TP5 and TN5, their activation functions are also linear functions.

[0165] The optimal method for training model parameters is to calculate them using the statistical characteristics of the circuit through table lookup and interpolation. After obtaining the final model parameters through backpropagation training, these parameters are then input into the circuit of this embodiment to achieve neural network output (for inference). For parameters exceeding the range in the software, they can be divided into multiple parameters and multiple inputs, or the hardware circuit can be improved to allow the weighted current modulation control module to select more different resistance values.

[0166] In the computational part, as summarized above, tp5-tn5, tn1-tp1, tp0-tn0, etc., are ideally equivalent to the summation of product terms commonly used in software neural networks (the constant term is simply setting the input of one of the product terms to 1), and are called differential multiplication-addition information / equivalent time difference information. tp4-tn4, tp3-tn3, tp2-tn2, etc., can also be used in neural networks, but there is no attempt or publicly available information in the software and academic communities.

[0167] Looking at the derivation process of tp5-tn5, tn1-tp1, tp0-tn0, for example, tn1, tp1, and tn1-tp1, the calculation results are independent of the capacitance value of the charging and discharging capacitors. This greatly facilitates the design and production of chips or circuits, because the capacitance value is difficult to determine precisely. Our invention's calculation results do not depend on the capacitance value, which brings a huge advantage to mass production applications. In addition, because tn1-tp1 calculates the difference, the errors in tn1 and tp1 caused by the leakage current of the positive and negative capacitor control module and the capacitor module can be mutually canceled to a certain extent (of course, the leakage current should be minimized in the design). The calculation errors caused by the resistance value deviation and the errors caused by the current modulation time deviation are relatively small, and their accuracy is high. The capacitor and resistance deviations (only related to the ratio) can be corrected by obtaining consistency parameters during testing and written into the internal memory of the neuron output module. The neuron output module automatically adjusts the XOR time difference when it is working.

[0168] As can be seen from tp5-tn5 and tp0-tn0, although the dual capacitors use twice the amount of capacitors, they can still perform the calculation of accumulating positive and negative product terms whether charging or discharging.

[0169] This embodiment implements the essential multiply-accumulate, linear activation, and nonlinear activation functions required by neural networks. Other activation functions are not strictly necessary for neural networks and can be subsequently handled by computing chips / GPUs / microcontrollers, or see more embodiments below. The multi-channel multiply-accumulate parallel circuit described above reduces the number of CMOS transistors used by at least an order of magnitude compared to a single floating-point multiplier in existing chips.

[0170] In some embodiments, grouping operations are not desired (in embodiments with single-group nonlinear activation*). Instead, it is preferable to use the traditional multiply-add form w1x1 + w2x2... followed by nonlinear transformation activation. This is consistent with the implementation of traditional neural networks. While there are hardware workarounds in this invention, they will increase system latency. Details are as follows:

[0171] The layer network uses a discharge mode. Each neuron output module has an additional equivalent conversion capacitor, pre-charged to c=1 volt. During the effective level of the P xor N output (as explained above, this duration is linear), the conversion capacitor is charged with a driving voltage a=2 through a resistor R (the resistor value here is set to the same value as the detection discharge resistor). Then, similar to the detection discharge process above, the charged conversion capacitor is discharged, also through the equivalent resistor R to ground. The discharge stops when the conversion capacitor voltage is less than or equal to the detection voltage b=c=1 volt. Similarly, the duration of the conversion capacitor detection discharge is a nonlinear transformation function with a shape similar to cap*R*ln(2)*tanh first quadrant curve.

[0172] That is, f(y) = cap * R * (ln[2 - e^(-y)]), (y = tp0 - tn0 or y = tp1 - tn1).

[0173] In effect, it's equivalent to adding a non-linear layer with one-to-one neuron connections on top of a linear activation layer.

[0174] Some embodiments may require implementing a nonlinear activation function (ADC detection type embodiment*) in the capacitor control module (8)(16) by using a high-speed ADC to read the capacitor voltage into a fast access unit, where a=2, b=1.

[0175] The positive capacitance V[t] is given by: gp2 = a + (ba) * e^(-t*SSP(...,Xn*Wn)) = 2 - e^(-t*SSP(...,Xn*Wn)).

[0176] The negative capacitance V[t] is given by: gn2 = a+(ba) * e^(-t*SSN(...,Xn*Wn)) = 2-e^(-t*SSN(...,Xn*Wn)).

[0177] Using op-amps to perform subtraction (gp2-1, gn2-1), and then importing the result into an ADC, the voltage result from the ADC is stored in a fast access unit (the value in the fast access unit). Similarly, a nonlinear function with a zero-crossing extreme of 1 and a tanh-like shape can be obtained in the first quadrant.

[0178] That is 1-e^(-y), y=t*(SSP or SSN)(...,Xn*Wn)

[0179] *An embodiment of 1:1 input-output* Description of the relationship between neuron circuit modules: When the ratio of input to output of the neural network is 1:1, as shown in Figure 11, all processes and steps, including enabling each module, writing and reading values, are controlled by a shared controller (260) of multiple neurons.

[0180] The clock source (266) sends the clock carry pulse signal to the global clock (265) and the input current modulation control module (267). The multi-bit cyclic timing signal output by the global clock (265) is compared with the value of the input register in the input current modulation control module to generate the input PWM. If necessary, the input PWM and the direct signal of the clock source can be combined (logical AND) to generate multiple refined PWM pulses. The PWM pulse is sent to the positive and negative complementary signal generator (251) to convert the PWM signal into a complementary PWM (or a complementary signal can be generated globally and then combined with the input PWM) to control the transmission gate charging and discharging switch (252) of the capacitor, which generates the actual PWM modulated charging and discharging current for the capacitor's charging and discharging channel. There are two sets of transmission gates corresponding to the positive and negative capacitors (the transmission gates can balance the parasitic charges injected when the NMOS / PMOS transistors are switched), and one is selected according to the positive or negative value of the input. The instantaneous magnitude of the PWM-modulated charging and discharging current is controlled by the resistor network (250) within the weighted current modulation control module. Its resistance value is controlled by the weight values ​​of the corresponding neural network links. The input current modulation control module and the weighted current modulation control module jointly control the charging and discharging currents of the positive and negative capacitors. This embodiment uses a discharge mode, so the discharge current ultimately flows back to the other end of the capacitor, i.e., the common terminal. The diagram shows only one link; in reality, the neural network has multiple links, meaning the capacitor has multiple discharge channels (253).

[0181] The capacitor control module (259), like the capacitor module, has positive and negative terminals. The positive capacitor control module controls the positive capacitor, and the negative one controls the negative one. Before forward propagation, the capacitor is pre-charged to a certain voltage (256), such as 0.21V, by the pre-charge module (255). Then, the capacitor is discharged through the transmission gate (257) inside the capacitor control module and the standard discharge resistor. The discharge stops when the standard pre-charge voltage, such as 0.20V, is reached. The discharge operation is achieved by the voltage comparator (254) inside the capacitor control module. This voltage comparator can select multiple reference comparison voltages through the voltage selection switch (268) to control the capacitor to achieve multiple voltages. The above is the voltage control process before forward propagation. After the capacitor module discharges during forward propagation, the next step is the discharge detection process. During the detection process, the transmission gate inside the capacitor control module is turned on, causing the capacitor module to continue discharging through the standard resistor until the capacitor voltage reaches the comparison voltage to terminate the discharge detection, such as 0.15V. The capacitor control module contains a fast logic circuit (258). Once the comparator output voltage arrives, the signal level of the entire capacitor control module output to the neuron output module (261) is flipped. The neuron output module contains an XOR circuit, which logically combines the level signals of the positive and negative capacitor control modules to generate a time difference (263) and a positive or negative sign. Finally, the time difference is converted into an activation function value (262) through a multi-bit signal (264) from the activation function value source. This value is then written to a buffer for later writing to memory, transmission to external systems, or rewriting to the input modulation control module, etc., thus completing this round of calculation.

[0182] In other embodiments, the ratio of input, weight, capacitor, and output modules can be adjusted as needed. For example, the operating frequency of the input current modulation control module and the weight current modulation control module may be much higher than the operating frequency of the capacitor module / capacitor control module / neuron output module. Therefore, multiple sets of capacitor modules / capacitor control modules / neuron output modules can be configured, allowing the lower-frequency modules to operate in rotation, thus increasing the overall operating frequency. Generally, the number of weight terms equals the number of input terms multiplied by the number of output terms, and is the largest, so optimizing its operating frequency and area is prioritized. Similarly, multiple sets of input current modulation control modules can also be configured.

[0183] *Dynamic Single-Layer Propagation Embodiment*, as shown in Figure 13, the neuron output module or intermediate module (73) has a register channel (72) to directly transmit the value of the output register to the input current modulation control module (71), sharing or jointly using the value of the fast access unit of the neuron input / output items (the input part of the fast access unit of the charge adjustment module). The value of the fast access unit of the output neuron output item is copied and written to the fast access unit of the input current modulation control module, or the two modules share a fast access unit. Compared with the dynamic alternating-layer propagation embodiment, this embodiment saves a lot of register resources, but CMOS resources are usually mainly consumed in the circuits corresponding to weight items and links, and the resources occupied by the circuits corresponding to input items and output items are very small. The working process of the two embodiments is similar. The basic computation unit of the neuron reads the weight item parameters and resistor selection parameters, as well as the output value of the previous layer network, from the storage module each time, and then uses them to calculate the output value of the current layer. In this process, the entire circuit is dynamically used to calculate each layer of the neural network, and finally outputs the result to the outside or writes it to the storage module.

[0184] The working process of a dynamic single-layer propagation implementation is as follows:

[0185] 1. Basic neuron units load data from storage modules or external sources into corresponding fast access units. (These include fast access units for various control settings, weighted items, input items, resistance selection, etc.)

[0186] 2. The capacitor begins pre-charging. (If there are fast access units related to capacitor control, they will also be reset before charging.)

[0187] 3. (Forward propagation steps 3~7) The network is powered on, that is, the input current modulation control module or intermediate module generates a modulation current according to the value of the relevant register, and charges and discharges the capacitor module through the resistor (select the charging mode / discharging mode according to the network design).

[0188] 4. Capacitor discharge is detected, and an activation function (if any) is executed. The output is stored in the fast access unit of the output item (i.e., the input of the next layer). The layer activation function value source needs to be set before executing the activation function.

[0189] 5. Copy the value of the output item's fast access unit to the input item's fast access unit of the current neuron's basic computational unit. (If shared, omit the copying process; sharing saves CMOS resources.)

[0190] 6. If there are special calculations involving layer parameters, such as SoftMax, then perform the relevant calculations for the network aggregation layer.

[0191] 7. Once all layers of the neural network have finished computing, write the computation results of the output fast access unit to the storage module or external storage.

[0192] 8. If the calculation is not finished, load other data required for the next layer of the neural network calculation (weight fast access unit, resistor selection fast access unit, related calculation settings, etc.) from the storage module or externally.

[0193] 9. Repeat steps 2 through 8 until all layers of the neural network have been computed.

[0194] Multiple dynamic layers can be arranged together symmetrically in space to balance the unidirectional flow of current in various situations.

[0195] The embodiment of multiple accumulations followed by activation*, the previous dynamic single-layer propagation embodiment, as shown in Figure 13, has the same number of input items in the input layer and output items in the output layer. Although the number of these two items is not equal in some embodiments described above, it is still fixed. At most, the number of input items in the linked input layer is reduced by closing the links between the preceding and following layers.

[0196] In this embodiment, the output module of the output layer has a counter for the activation function timing pulses. If the fundamental frequency is relatively high, the layer activation function value source can have multiple pulse lines compared to a single-line pulse counter. For example, 1x, 2x, 4x, 8x, etc. pulse lines. For instance, for an 8x line, one pulse increments the counter by 8.

[0197] The microcontroller and DMA (Memory Direct Write) corresponding to the layer activation function value source generate the programmed activation function parameters. The internal circuitry of the layer activation function value source generates the aforementioned multi-line pulses based on these parameters and the base frequency. Activation function parameters include, for example, an interval time register specifying how many base frequency pulses are needed to generate one multi-line pulse; how many base frequency pulses the interval time register increments or decrements by 1 (i.e., the interval time second derivative register); and how many base frequency pulses need to be read from memory again (i.e., the update time register). In short, the chip calculates the function parameters for pre-generating interpolation, and the internal circuitry generates the interpolation multi-line pulses based on these parameters. The layer activation function value source is positive and negative, meaning it generates two sets of multi-line pulses simultaneously, one positive and one negative. The output module selects the positive or negative multi-line pulse for counting (representing a multi-bit increment) based on its own sign.

[0198] As shown in Figure 9, a large input layer (90) is divided into multiple smaller input layers (91)(93)(95)(96). To accumulate the output values ​​of these smaller input layers, it is necessary to perform charging and discharging operations on the network capacitors corresponding to the number of neurons in the input layers, as well as detection and activation operations (forward propagation) (92). These operations use a linear activation function y=x to convert the duration difference of the xor level of the positive and negative capacitors into the value of the fast access unit of the counter. The counter adds a multi-bit increment value expressed by the multi-line pulse each time, which is the time increment value of the activation function.

[0199] The forward propagation of the input layers (91), (93), (95), and (96) all use the same circuit modules, only the input and weight parameters read from the external array storage module are different. That is, the layers divided in the figure are the parts of the multi-round operation process of the parameters of each layer, and the physical circuits are the same.

[0200] In this embodiment, the output layer (94) and the input layer (91) are powered on and perform the detection and activation processes completely, thus completing the first forward propagation of the segmented layer. This is because the activation function used is a linear function of y=x, and its value is stored in the fast access unit of the activation function timing pulse counter in the output module.

[0201] Then, the second forward propagation is performed, namely the forward propagation of the output layer (94) and the input layer (93). Of course, before the input layer (93) performs the forward propagation, the corresponding input items and weight parameters are loaded from the external memory (or the input module itself has more fast access units to temporarily store the values ​​of the input items in advance). However, it is different when the activation function is executed. In the second forward propagation, the counter of the activation function timing pulse will not be reset before activation. Because it is a linear function of y=x, the activation of the neuron is uniformly increased, and the positive and negative pulse lines are the same. Therefore, it is only necessary to continue to accumulate or decrement the pulse count during the duration of the difference in the duration of the positive and negative capacitor XOR levels when the logic level is 1. Of course, if the sign of the positive and negative capacitor logic synthesis is negative (that is, the output time of the negative capacitor module is longer), then the counter should select the pulse signal of the negative layer activation function value source, and should decrease the count value of the counter. If it is reduced to below 0, then the logic level signal related to the sign of the entire output module needs to be flipped. After the second process is completed, the result of the fast access unit of the counter is the absolute value of the arithmetic sum of all product terms in the two forward propagations.

[0202] Repeat the second forward propagation process, calculate the forward propagation of the output layer (94) and the input layer (95) (96) respectively, and then the result is the arithmetic sum of the products of input terms and weight terms of the forward propagation of each neuron according to all input layers.

[0203] In this embodiment, forward propagation of input items that are a multiple of the number of output neurons is achieved through step-by-step calculation.

[0204] Furthermore, this embodiment can also implement a residual neural network. Compared to the paired residual neural network embodiment*, this embodiment does not require pairing circuitry based on neuron coordinates or positions in hardware. Its operation is as follows:

[0205] First, during the first round of forward propagation at the first layer, the appropriate (non)linear activation function is selected.

[0206] Then, in the next layer of forward propagation after the second round, a linear activation function for y=x is chosen.

[0207] The result is that the output value of a single neuron equals the first (non)linear activation value plus subsequent linear activation values.

[0208] Self-attention matrix operation neural network example* The self-attention model calculates Q, K, and V by multiplying the input X by the parameter matrix. The calculation process is similar to the calculation process of tp0-tn0 or tp1-tn1 in the fully connected linear activation above. Although it often uses matrix mathematical operations to express and describe it, there is no difference in essence. Both involve multiplying the input terms with multiple weight terms and finally summing the multiple product terms.

[0209] Then, positional encoding information needs to be added to the Q or K matrices. This encoding information can take many forms, but the most common approach is to add a calculated positional encoding value to each value of the input vector X before generating the QKV matrices. This calculation is performed before multiplication and addition. Therefore, it can be performed by common digital circuit modules (CPU, FPGA, DMA, lookup table circuits, etc.). However, the circuit of this invention can also be used for calculation because in many subsequent embodiments, the neuron output module has a digital adder / CPU, and the general-purpose circuits for internal adders for neuron input and output functions can be shared. In other words, this part of the circuit can be directly used for positional encoding addition. Alternatively, positional encoding information can be directly generated at the interface when reading the input vector and added to each element value of the word vector, thus enhancing the computational capability of the reading interface.

[0210] In some large model designs, positional encoding information is inserted into the generated Q / K / V matrix. This insertion of positional encoding information also utilizes the bias parameters of the original neural network. As mentioned earlier, when the current modulation of the network input layer is 100% (i.e., energized at all times), the contribution of the entire resistive link to the layer output depends entirely on the resistance value and the current modulation representing the weight parameters of the resistive network. Therefore, during the generation of Q / K / V, the bias parameters can be set to add the corresponding positional encoding information to the output values ​​of the relevant neurons. In other words, when generating the Q / K / V matrix, positional encoding information (specific numbers loaded from the storage module or externally) is added to the output values ​​before activation using the bias parameters.

[0211] Then, when multiplying multiple QK matrices, the multiplication of a single QK matrix (including matrix transpose, which is essentially the sum of multiple product terms) requires writing the value of one matrix into the input term current modulation control module's register representing the input term, and the value of the other matrix into the weight term current modulation control module's fast access unit representing the weight parameters. Then, forward propagation / power-on detection activation is performed according to the process. That is, in one round of calculation, one set of input values ​​is multiplied and added with multiple sets of weight values ​​respectively, and the output value is stored in the fast access unit. The result of matrix QK multiplication is obtained through multiple rounds of calculation.

[0212] Therefore, the process of generating and multiplying the QK matrix in the self-attention mechanism, as well as the final QKV matrix operation, can be calculated using the circuitry of our already disclosed embodiments. The steps are as follows: (According to the definition of the self-attention model, Q / K / V refer to the intermediate value matrix generated by multiplying different parameter matrices with the input vector matrix).

[0213] First, during the generation of the QK matrix, the system performs multiplication and addition calculations for each row of the matrix through multiple rounds of operations. Specifically, the system reads a portion of the Q matrix parameters (one row or one column per round) from external memory or internal cache and loads it into the fast access unit of the input current modulation control module. Simultaneously, the position-encoded input vectors (i.e., word vectors, which may be multiple) are loaded in parallel into the fast access unit of the weight current modulation control module. Then, the neural network performs a forward propagation calculation, outputting the Q matrix result value for the current round (if the number of words exceeds the number of neurons, it needs to be divided into multiple batches and multiple rounds of calculation; also, the parameter matrix may have multiple rows, requiring several rounds of calculation), and temporarily stores it in the spare fast access unit of the corresponding weight module. This process is repeated until all Q matrix results are filled into the spare fast access units of the corresponding weight modules. In addition to multiple rounds of calculation, the output vector can also be calculated in blocks within a single round, simultaneously calculating QKV, with the results of each block written to the corresponding storage circuit.

[0214] Next, the system enters the K-matrix processing stage. At this point, the system loads some parameters of the K-matrix (one row or one column per round) into the fast access unit of the input module, and simultaneously loads the position-encoded input vectors (multiple word vectors) into the main fast access unit of the weight module in parallel. The neural network performs forward propagation calculations, generating the corresponding K-matrix result value, and passes this result back to the fast access unit of the input module. Then, the contents of the main and backup fast access units of the weight module are swapped (i.e., the Q-matrix result value is read), so that the Q-matrix results of each round are reused as weight values. Afterward, forward propagation is performed, and the neuron output module obtains the result of a round of QK matrix multiplication and addition operations, saving it to external memory or an idle fast access unit. This process is repeated multiple times until the entire QK matrix multiplication calculation is completed.

[0215] (If the positional encoding is inserted during the generation of Q / K / V instead of in the word vectors, then during the calculation of Q / K / V, by setting the positional encoding in the bias term parameters of the neural network, the output of the network calculation will include positional information.)

[0216] After completing the QK matrix, the system further processes it using the SoftMax normalized activation function (see related examples below), and temporarily stores the activation results externally or in memory. This process also gradually generates all QK activation results through multiple rounds of calculation.

[0217] During the V matrix generation phase, the position-encoded input vectors corresponding to multiple terms are loaded into the fast access unit of the weight term current modulation control module. One round of V matrix parameters is then loaded into the fast access unit of the input term current modulation control module. In the forward propagation computation, the neuron output module generates the V matrix result and outputs it to the spare fast access unit of the weight term current modulation control module or external / memory. This process is repeated until the spare fast access units for all weight terms are filled.

[0218] Finally, the system enters the QKV matrix generation stage. The weight term current modulation control module switches to the backup fast access unit, which contains the V matrix values ​​corresponding to each term. The result of multiplying the QK matrix (originally stored in memory or externally) by the SoftMax-normalized QK matrix (i.e., the attention weights) is read into the fast access unit of the input term current modulation control module in multiple rounds. Then, forward propagation calculation is performed to obtain the final QKV matrix calculation result. This result is then used for further calculations.

[0219] It's worth noting that, to accommodate the processing needs of large-scale input vectors, the system supports grouping input vectors for processing. That is, each batch processes only a portion of the input data, and through multiple rounds of computation, the complete output result is finally concatenated (using addition). This approach not only improves the system's flexibility but also effectively reduces hardware resource consumption. Word vectors don't necessarily have to be read into the weights; they could be placed in the inputs as well. However, word vectors are significantly more numerous than parameter matrices, and weights require more circuit resources. Therefore, they are placed in the weights, while the parameter matrices are placed in the inputs because of their smaller number.

[0220] In summary, this embodiment implements the core computational process of the self-attention mechanism through hardware circuit design, featuring high efficiency, low power consumption, and strong scalability, making it suitable for the deployment and application of Transformer-type models on low-power edge devices or dedicated AI chips.

[0221] SoftMax Implementation Example*: Calculate the activation function SoftMax, which is defined according to a well-known definition:

[0222] SoftMax = e^(Yi) / SumAll[e^(Y1), e^(Y2),...e^(Yn)]

[0223] The exponent of the natural base of the current neuron is divided by the sum of the exponents of all neurons. The above calculation process can of course be output to an external chip module for combined internal and external calculation, but our embodiment has better parallel performance. As shown in Figure 14, this embodiment also has dynamic odd layer (61) and dynamic even layer (62) alternately powered on and executed. According to the above embodiment of implementing arbitrary activation function of layer, the e^(Yi) function can be calculated, that is, the PWM output of this layer is equivalent to e^(Yi) (proportional to e^(Yi)), and then the linear sum of this layer, that is, S=SumAll[e^(Y1), e^(Y2),...e^(Yn)], is calculated in the next layer (66). This is called the S value. The function of SoftMax's whole layer normalization is mainly implemented through this part. Since SumAll[e^(Y1), e^(Y2),...e^(Yn)] only requires one neuron, then in hardware, an additional neuron unit with a special bidirectional intermediate module (67) can be designed as the network summarization layer (66). The register of the output of the corresponding neuron can be directly used by the layer calculation module (64) through the data path (65). After the layer calculation module obtains the S-value parameter, it uses its reciprocal to generate a fast access unit value (usually generated by a high-frequency microcontroller) (modulated by PWM, etc.).

[0224] Then, this value is simultaneously written to each neuron through the parallel write channel (69) connecting the entire layer to the fast access unit of the output PWM and the PWM part representing the weight value. This value is W=a / SumAll[e^(Y1), e^(Y2),...e^(Yn)], and the value of the fast access unit of the PWM representing the output value in the current layer is e^(Yi). The product of the two is proportional to SoftMax (= a*e^(Yi) / SumAll[e^(Y1), e^(Y2),...e^(Yn)] ). The proportional coefficient a needs to be adjusted according to the actual needs based on the resistance and capacitance value through the base frequency. The link to the next layer here is a one-to-one connection. Because it is a one-to-one connection, the functions of the current layer and the next layer can be merged, and the system can charge and discharge itself. That is, the current bidirectional intermediate module (63) charges and discharges the capacitor (68) of its own neuron unit through the internal resistor using PWM, and then discharges and compares the discharge time to obtain the discharge time. Then, linear activation is performed (that is, e^(Yi) is scaled proportionally and normalized according to the S value) to update the output register value. Then, the entire neural network is calculated alternately as in the dynamic alternating-layer propagation embodiment disclosed above. Because the system charges and discharges itself, the sign is known or only the absolute value is known, so the capacitor (68) is selected.

[0225] In this way, after the operation is completed, a set of multiple SoftMax activation probabilities is obtained. If multiple sets are computed in parallel, the results of multiple SoftMax activations can be obtained simultaneously.

[0226] *An example of a high-frequency, single-layer dynamic network with shared links*

[0227] Building upon the previous embodiment, the number of modules other than the weighted current modulation control module is doubled (greater than or equal to 2 times), the number of switch-related circuits used for current modulation is doubled, and the weighted current modulation control module (including the weighted fast access unit, etc.) is shared. This allows for a multiple increase in the computational output frequency of the entire network, resulting in a performance boost. Furthermore, because the number of weighted current modulation control modules in a fully connected neural network is far greater than the number of other modules, this doubling consumes relatively few circuit resources while maximizing computational power. It also balances the inconsistent delays among modules in the circuit and reduces the memory bandwidth requirements of the weighted parameters.

[0228] After doubling, multiple different networks can be computed simultaneously, meaning networks with the same weight parameters and activation functions but different input and output values. If the circuit is specifically designed for convolutional neural networks, multiple convolutional kernels can be computed simultaneously by sharing the weight current modulation control module (such as the weight fast access unit). This does not require additional weight current modulation control modules or related chip resources.

[0229] The number of fast access units for weighted terms in the current modulation control module can be multiplied according to the demand for parameters, which facilitates rapid batch-by-batch parameter replacement, thereby increasing the flexibility for high-frequency calculations.

[0230] As shown in Figure 15, this embodiment has a double design, with double the input current modulation control module (120) (or the neuron basic computing unit has two input current modulation control modules), which can generate two sets of input PWM based on the two input fast access units, the first input PWM (121) and the second input PWM (124). Then, based on the two input PWMs, new modulation current signal waves are selected or generated from the first bottom-level high-frequency reference PWM and the second bottom-level high-frequency reference PWM, that is, the first selected reference pulse signal (122) and the second selected reference pulse signal (123) are output simultaneously. Because the new modulation signal waves selected or generated by the first bottom-level high-frequency reference PWM and the second bottom-level high-frequency reference PWM are sparsely staggered in pulse phase (the input current modulation control module receives two sets of bottom-level reference pulse signals with different phases sent simultaneously by the shared pulse transmitter of the layer), the pulse signal phases of the first selected reference pulse signal and the second selected reference pulse signal are also staggered. Then, the multi-line channels (125) and (129) of each of the first neurons, which point to the corresponding weight term current modulation control modules, simultaneously send the first selected reference pulse signal (122) and the second selected reference pulse signal (123) (the symbols also have corresponding circuits). The weight term current modulation control modules insert delays and adjust the instantaneous current through the resistor network. As mentioned above, the weight values ​​are shared, so the weight term fast access unit and other parts can be shared. The internal open CMOS switches and other parts can also be partially shared through optimized design. After inserting the delay, the first weight term current modulation control module generates two sets of modulation signals, the first modulation signal and the second modulation signal. The two sets of signals control the first set of positive and negative switches (131) and the second set of positive and negative switches (138) of the first neuron, respectively (the positive and negative are distinguished according to the product of the input and the weight). In addition, the signal sent by the second input term current modulation control module is processed by the second neuron weight term current modulation control module (127) to generate two sets of modulation signals, which control the end of the second set of positive and negative switches of the first neuron and the second set of positive and negative switches (130) of the first neuron, respectively.

[0231] That is to say, the switch of the modulation current in the above current channel generates the modulation current for charging and discharging the corresponding positive and negative capacitor modules (132) and (133) according to the modulation signal. Since the transmitted signals are all pulse signals, the single-chain weight term current modulation control module (126), the first neuron weight term current modulation control module (128), and the second neuron weight term current modulation control module (127) all have channels for the driving voltage source of the capacitor module charging and discharging (discharging to ground or charging to a specified voltage). Since the pulse signals received by the weight term current modulation controller are sparsely staggered, it is possible to generate two sets of modulation currents at the same time. As shown in the figure, the neuron output module (134) and (135) and the positive and negative capacitor control module are multiples.

[0232] The number of neurons (136)(137) in the normalized network summation layer is multiplied. This allows for simultaneous forward propagation to two network layers. This means that a series of computational operations, such as charging / discharging, detecting discharge, and activation, are performed, effectively doubling the computational speed.

[0233] In this embodiment, the multiple and the number of neurons are both 2, which is only for illustrative purposes. In actual applications, the number can be expanded according to the needs.

[0234] *An Example of a Delayed Single-Layer Dynamic Network with Segmented Multi-Line PWM* In the previous example, the multi-line modulation signal for the input was used in different neurons (although the weight register was shared). In this example, the multi-line modulation signal is used in the same neuron, which can significantly improve the overall operating frequency of the neural network.

[0235] The input fast access unit is divided into multiple segments bit by bit. For example, if the input fast access unit is 16 bits, it can be divided into high 8 bits and low 8 bits (two segments), or it can be divided into high 4 bits, mid-high 4 bits, mid-low 4 bits, and low 4 bits (four segments). Each segment generates PWM independently.

[0236] As shown in Figure 16, the input fast access unit design in this embodiment is divided into two segments: a high 8-bit segment and a low 8-bit segment. The two segments generate PWM independently, namely, a high 8-bit PWM (140) and a low 8-bit PWM (144). Each of the two PWMs is combined with the underlying reference pulse signal (142) and (145) to synthesize or select a new modulation signal pulse signal, namely, a high 8-bit selected pulse signal (142) and a low 8-bit selected pulse signal (146). The two selected pulse signals are simultaneously sent to the connected weighted current modulation control module. Each connected weighted current modulation control module inserts a delay into the two selected pulse signals according to the value of its internal weighted fast access unit, thereby generating the final modulation current switching signal, namely, a high 8-bit modulation current switching signal (143) and a low 8-bit modulation current switching signal (147). The two signals independently control the switching of the modulation current for charging and discharging the positive and negative capacitor modules, namely, positive and negative high 8-bit switches and positive and negative low 8-bit switches. The capacitor module is charged or discharged based on the sign of the product term (input value and weight value), which is selected as positive or negative.

[0237] If a higher frequency or shorter pulse time is required for the underlying reference pulse signal, a common multi-line multi-phase clock signal can be used (to generate a higher resolution PWM), or a delay insertion method can be used (using delay circuits based on inverters, etc., to generate a modulated current signal such as a single pulse or PWM). It should be noted that the actual frequency and pulse width of the underlying reference pulse signal are not limited to those shown in the figure, and the current modulation is not limited to PWM.

[0238] In this embodiment, the current channels corresponding to the links of the positive and negative high 8-bit switches and the positive and negative low 8-bit switches have different resistance values. The ratio of their resistance values ​​is consistent with the maximum value of each segment. For example, the resistance value of the low 8-bit switches is 2^8 = 256 times larger than that of the high 8-bit switches, representing a smaller 1 / 256th of the current. If the weighted current modulation control module includes an optional resistor network, that is, some bits in the weighted register represent the selection of the resistor network resistance value, then the adjustment of the resistor network resistance value, for the high 8-bit current channel and the low 8-bit current channel, will still result in a ratio of 256 times. This is determined by the number of bits in the input fast access unit. During charging and discharging, the current channels corresponding to the links of the high 8-bit switches and the positive and negative low 8-bit switches converge into a capacitor module of the same symbol.

[0239] The structure of this embodiment, including multiple proportionally sized capacitor modules for charging and discharging modulated current channels, multiple resistor networks, and multiple PWM modulation signals within a single neuron circuit, effectively improves the operating frequency. For example, with 100 neurons and 100 fully connected lines, if it operates at 100M times per second (under 16-bit PWM limitations), then ideally, the number of weight parameters calculated per second is 100*100*100M=1T. If it is divided into two equal segments, then the computing power is 1T*256=256T. Dividing the input register into four segments would achieve even higher frequencies.

[0240] Not only can the input fast access unit be segmented, but the weight fast access unit can also be segmented. Assuming it's also divided into high 8 bits and low 8 bits, if the input current modulation control module sends two selected pulse signals, then the weight modulation control module independently inserts delays in the high and low 8 bits, generating four switching signals from the two selected signals: high 8 + high 8, high 8 + low 8, low 8 + high 8, and low 8 + low 8. These four switching signals control four positive and negative switches (eight switches in total, including positive and negative), and each corresponding current channel has its own corresponding resistance value, i.e., resistances in the ratios of 1, 256x1, 1x256, and 256x256. However, because the number of neuron connections is large, i.e., the number of weight current modulation control modules is large, this part of the circuit needs to be simplified as much as possible.

[0241] The resistance error at the high level is not consistent with that at the low level because it is divided into multi-wire switches and resistors. This inconsistency in error will interfere with the accuracy of the low level. In addition to production and design optimization (such as having a correction circuit inside the circuit), this situation also requires statistical detection or circuit self-testing to generate a numerical conversion table to correct the input value and weight value, so as to obtain an accurate current.

[0242] In some other embodiments, not only can the input modulation control module and the weight modulation control module be segmented, but the capacitor module, capacitor control module, and neuron output module can also be divided into multiple segments. In this case, the charging and discharging input currents of each segment's capacitor module do not need to be aggregated; instead, they are independently input into the capacitor module and aggregated after the output time difference (using digital or other addition circuits). Regarding the activation function, the layer activation function source and the multi-bit incrementing counter can also be improved with reference to this segmented pulse signal design. The corresponding output values ​​of multiple segments are summed, and the activation function value is obtained based on the summed value.

[0243] (The segmentation of the input item fast access unit may include, but is not limited to, the following forms: equal-weighted segmentation (such as every 4 bits or other numerical values), non-equal-weighted segmentation (high bit width is not equal to low bit width), dynamically adjustable segmentation (configured via registers), and the modulation signals generated independently by each segment can drive the capacitor module in parallel. Other embodiments are similar.)

[0244] *An Implementation Example of a Single-Layer Dynamic Network for Multi-Segment Multi-Line Conventional PWM* In the previous embodiment, the weighted current modulation control module directly inserted a delay into the sent short pulse signal; this embodiment uses a selection / AND gate method for multi-channel PWM synthesis, which is technically less difficult.

[0245] Assuming the total number of bits in the fast access unit is 4, divided into 2 segments, as shown in Figure 17, a basic computational unit of a neuron, its input current modulation control module generates A: high 2-bit input PWM (150) and B: low 2-bit input PWM (152) according to the input fast access unit segmentation (sign bit is calculated separately). At the same time, the base frequency of the weight current modulation control module is 1 / 4 of the base frequency of the input current modulation control module. The weight current modulation control module generates C: high 2-bit weight PWM (151) and D: low 2-bit weight PWM (153) according to the weight register (sign bit is calculated separately). The result of their respective logic AND combination is, A and C combined (154), A and D combined (155), B and C combined (156), B and D combined (157), that is, (A*4+B)*(C*4+D)=16*A*C+4*A*D+4*B*C+B*D. The numbers in the formula represent the magnitude of the current. After the switching signal is connected to the four current switches, the resistance ratio of AC in each current channel is 1 / 16, the resistance ratio of AD is 1 / 4, the resistance ratio of BC is 1 / 4, and the resistance ratio of BD is 1.

[0246] If the total number of bits in the fast access unit is 16 (excluding the sign bit), divided into two segments, then its input current modulation control module generates A: high 8-bit input PWM and B: low 8-bit input PWM based on the input fast access unit segmentation (sign bit counted separately). Simultaneously, the base frequency of the weighted current modulation control module is 1 / 256 of the base frequency of the input current modulation control module. The weighted current modulation control module generates C: high 8-bit weighted PWM and D: low 8-bit weighted PWM based on the weighted register (sign bit counted separately). (A*256+B)*(C*256+D)=(256^2)*A*C+256*A*D+256*B*C+B*D. Therefore, after connecting the switching signals to four current switches, the resistance ratios of the product terms in each current path are 1 / (256^2):1 / 256:1 / 256:1.

[0247] If the total number of bits in the fast access unit is 12 (excluding the sign), divided into 3 segments, then (a*(16^2)+b*(16^1)+c*(16^0))(e*(16^2)+f*(16^1)+g*(16^0))=a*e*(16^4)+a*f*(16^3)+a*g*(16^2)+b*e*(16^3)+b*f*(16^2)+b*g*(16^1)+c*e*(16^2)+c*f*(16^1)+c*g*(16^0); After the switching signal is connected to 9 current switches, the resistance ratio of each product term in each current path is the reciprocal of the coefficient of each product term. If the switching frequency is very high, two sets of switches can be used alternately, i.e., 9*2 current switches.

[0248] Considering factors such as switching delay, resistance error, and trace capacitance and inductance, it is necessary to convert the actual pulse width or parameter values ​​of the input and weight terms based on actual test statistics. This conversion will not affect the actual execution. In general, the connections between neurons achieve higher operating frequencies by performing parallel charging and discharging operations on the capacitor module through multiple current channels.

[0249] *An Example of Fixed-Point Calculation Using a Full-Resistance Network for Factoring Product Terms by Bit-Digit Segmentation* Using a resistor network for all weight terms allows for higher frequencies. Reducing the resolution of the input PWM ensures high accuracy with a resistor network (easier to eliminate various interference factors). Decomposing large-digit multiplications into multiple smaller multiplications provides a wider computational range, and also optimizes the speed and energy consumption of capacitor charging and discharging. If factoring involves multiple non-parallel calculations, time can be traded for chip space, because the increased frequency does not actually double the computation time with multiple calculations. Furthermore, the reduced resolution also lowers the voltage difference during capacitor charging and discharging.

[0250] Assume a neural network has n inputs and n outputs, with each neuron having n links; input values ​​X1, X2, ... Xn; weights (A11, A12... A1n), (A21, A22... A2n)... (Am1, Am2... Amn); output value Y1 = X1*A11 + X2*A12 + ... Xn*A1n, Y2... Yn; the product term of Y indicates the number of links in the neuron; assuming X1 is 16 bits, it is split into X11, X12, X13, X14 by bit; the weights are also split similarly, A11 is divided into A111, A112, A113, A114; A12... and so on; then the first data link of Y1 = (X11 + X12 + X13 + X14) * (A111 + A112 + A113 + A114); the first data link of Y1 = X11 * (A111+ A112+A113+A114)+ X12* (A111+ A112+A113+A114)+X13* (A111+ A112+A113+A114)+X14* (A111+ A112+A113+A114); If the input current modulation control module corresponding to X1 uses two channels of 4-bit PWM concurrently in two rounds, the first round is X11 and X12, and the second round is X13 and X14; each channel is associated with four segmented weight values ​​A111, A112, A113 respectively. A114 performs multiplication, yielding four product values ​​(shifting is not considered here, but will be addressed later). The sum of the shifted product terms represents X1's independent contribution to the output of the first link of the total output Y1. The contributions of the outputs of the other links of the first neuron are calculated similarly. The sum of the product terms of all input links X1 to Xn is the output value of Y1. Two operations with two 4-digit numbers can calculate a 16-digit number; similarly, four operations with two 4-digit numbers can calculate a 32-digit number, and eight operations with two 4-digit numbers can calculate a 64-digit number (integer or fixed-point number).

[0251] This embodiment is applicable to fixed-point numbers of any number of bits (if the number of bits is sufficient, the fixed-point number can be converted to a floating-point number). However, in the following design, the input item fast access unit has a value of 16 bits, divided into 4 segments, as shown in Figure 20. In the first round of forward propagation, the first input item current modulation control module (170) generates the high 4 bits (171) and the middle high 4 bits (172), the second round middle low 4 bits and the second round low 4 bits. Here, the high and low bits are determined by the binary bit representing the maximum value as the high bit. The weight item fast access unit also has a value of 16 bits, divided into 4 segments: high / middle high / middle. Low / Low are applied to the following modules respectively: the first input high 4-bit weighted term current modulation control module (173), the first input middle high 4-bit weighted term current modulation control module (174), the first input middle low 4-bit (175) weighted term current modulation control module, the first input low 4-bit weighted term current modulation control module (176), and the second input high 4-bit weighted term current modulation control module, the second input middle high 4-bit weighted term current modulation control module, the second input middle low 4-bit weighted term current modulation control module, and the second input low 4-bit weighted term current modulation control module. The input term current modulation control module corresponding to the second input PWM cannot be drawn due to image size limitations, but its existence should be easily understood. Since the input term value has two segments per round and the weight term value has four segments, each neuron in the weighted term current modulation control module should have eight modules. The dot symbol (184) in the figure indicates that there are other weighted term current modulation control modules that are not drawn but actually exist.

[0252] Since each segment of the weighted terms has only 4 bits, or 16 values, it can be easily implemented using a resistor network. The input PWM should adopt a waveform center-aligned mode, so that each pulse is centered, resulting in small interference error and achieving consistent and precise current control with fewer pulse repetitions (to accommodate changes in capacitor voltage during charging and discharging). In this round, the high 4 bits and the middle high 4 bits of the PWM are sent from the input multiple signal channel (177) to the 4*2=8 weighted term current modulation control modules mentioned above. In the neuron of this embodiment, the neuron output module, positive and negative capacitor module, and capacitor control module are also divided into 4*2=8 sets according to the number of segments, just like the weighted term current modulation control module. During forward propagation, they operate independently and synchronously according to each segment, and finally output the XOR time difference, storing the value and sign of the time difference in the fast access unit. The conversion of the time difference value is the same as in the previous embodiment, which uses the timing information of the layer activation function to convert the time difference value into a digital signal value in the fast access unit of the segment output module. However, the values ​​of the segment output modules (179) corresponding to each segment in a neuron have bit differences. Because there are bit differences between the segments of the input value and the weight value, the calculation results naturally also have bit differences. Therefore, there is a neuron summarization calculation output module (180) to obtain the output values ​​of each segment and perform addition with shift and sign. The figure also illustrates that the neuron summarization calculation output module can read or copy the segment output values ​​from the data channels (183) of the eight segment output modules. In this embodiment, the function of a digital adder is obviously added. The operating frequency of the digital adder is much faster than the addition function through the time source of the layer activation function in other embodiments. At the same time, it has shift and offset functions, which can quickly realize certain linear transformation functions. The neuron summarization calculation output module, like the neuron output module in the previous embodiment, has various activation and normalization functions.

[0253] The second round of calculation is the same as the first, except that the PWM value represented by the input item is the lower 4 bits of the second round and the lower 4 bits of the second round. The neuron summation and calculation output module needs to sum the values ​​from the two rounds in bit order to obtain the final accumulated value.

[0254] Regarding the activation function, previous embodiments used time differences to obtain the activation function. However, in this embodiment, the capacitor XOR time difference has been converted into the value of the fast access unit and summarized bit by bit, resulting in a complete digital signal. This embodiment uses a comparison-based method. It can be envisioned that the layer activation function source outputs (positive and negative paths) time values ​​and activation function values ​​(multi-line data, with the time values ​​and output module values ​​having the same number of bits). The time values ​​of the layer activation function source are compared with the values ​​calculated by the neuron output module; if they are equal, the activation function value is copied. However, this operation is not compatible with the operating frequency of other modules. Because bit segmentation reduces resolution, the PWM frequency is significantly increased, the voltage difference between capacitor charging and discharging is significantly reduced, and the speed is naturally significantly increased. Therefore, the envisioned activation function implementation method becomes a speed bottleneck. The solution is a piecewise linear interpolation method for digital circuits, matched with the number of calculation rounds, as follows:

[0255] As shown in Figure 21, the activation function curve is divided into multiple segments at equal intervals (assuming 10,000 segments), containing positive and negative activation function values ​​(192) and (190). Each segment has an initial value (191), which is the value at the intersection of the segment line and the activation function curve. The layer activation function source sends the segment time value, initial value, and derivative value (even the second derivative value) sequentially at a certain rhythm, sending positive and negative values ​​simultaneously. If the aforementioned values ​​are refreshed at a frequency of 1G, for 10,000 segments, the total time is 0.01ms, which is a relatively large delay, but acceptable. When each neuron compares the accumulated output value of the output module with the sent segment time value, it copies the initial value and derivative value. Then, linear interpolation is performed. Its serial multiplier multiplies the derivative value based on the difference between its own accumulated output value and the segment time value, and the final result is the activation function value of each neuron. This module already has the function of shifting and adding, so adding the function of the serial multiplier will not consume too many resources and time.

[0256] (Activation function value generation methods include: simple addition and subtraction counting based on data sent by the layer activation function source, direct copying, time-based accumulation / subtraction, pre-calculation lookup table (with pre-stored function-related values ​​in the table), real-time calculation (CPU dynamic calculation), piecewise linear interpolation (output slope + intercept), second-order interpolation, etc. Derivative information can be obtained through multi-line signal circuits or tables with pre-stored derivative values. Other embodiments are similar. In the bit-segmented embodiment, the linear values ​​of each segment calculation result can be obtained first based on the capacitor discharge time difference, and then the final activation function value can be obtained based on the combined segment output values ​​through table lookup and interpolation calculation or other methods.)

[0257] If necessary, this can be coordinated with the number of rounds of segmentation. In the first round of calculation, we have already obtained the high-order and mid-high-order calculated values. We extract the largest part, i.e., the highest few bits, which will not be affected by the accumulation in the second round. At this point, we can perform preliminary function segmentation, assuming there are n preliminary segments. The layer activation function source simultaneously distributes n sets of previous time values, initial values, and derivative values, i.e., subdivided segmentation or secondary segmentation. The neuron aggregation calculation output module selects the corresponding data channel of the preliminary function segment based on its own largest value part, and then copies the relevant subdivided segment values ​​when switching according to the comparison results mentioned earlier. This can speed up the communication. Finally, linear interpolation within the segment is performed, i.e., serial multiplication, to obtain the final value. If there are more rounds of calculation and more total bits, the segmentation can be finer and more convenient; if there are fewer rounds and fewer total bits, the calculation can be performed after the total accumulated value is obtained. If multiple neural networks with shared weight parameters are computed simultaneously, as in the example of a high-frequency single-layer dynamic network with shared links*, then the activation function can be computed in a pipelined manner for each layer of the neural network. That is, although the computational latency remains unchanged, the computational power is not affected by the pipelined approach.

[0258] It can be seen that the functionality of the neuron aggregation and output module is already close to that of a very simplified microcontroller CPU, and further development could achieve more complex functions. Moreover, multiple neuron aggregation and output modules can share many internal circuits, as can other modules; many internal circuits can be shared and optimized. Simultaneously, multiple neural networks or multiple neurons in a single network can also share the CPU or the forward propagation link circuitry (i.e., a large number of weighted current modulation control modules). For example, the CPU can manage the selection and copying of data in various fast access units (registers or SRAM, etc.); it can directly convert time values ​​to activation function values ​​through memory lookup tables; it can linearly transform time differences through direct multiplication; it can manage the workflow sequence of multiple neural networks, neurons, or input / output modules using the CPU and program, optimizing performance based on the delay of each process / step; and it can manage communication between larger modules and communication within and outside the chip. We will only provide an overall description of the logic and functionality here; specific details will be omitted.

[0259] The circuit of this invention uses this factorization form to perform multiplication and addition operations. It needs to control segmentation errors (when shifting and adding, the relatively low bits of the higher-order segments become the higher bits of the final number, thus amplifying the shifting and adding errors of these relatively low-order segments). If the number of neuron links is small, such as in a convolutional neural network, then accuracy is easily guaranteed; however, if it is fully connected with many neuron links, then to ensure accuracy in the middle and lower digits, there are requirements for the voltage difference between capacitor charging and discharging and the voltage comparator of the capacitor control module, which must ensure the accuracy of the calculation results.

[0260] Method 1 involves using different options for capacitor pre-charging and comparison voltage difference when performing segmented calculations in a fully connected neural network. Higher-order segments have larger voltage differences, longer PWM times, and lower calculation frequencies; lower-order segments have smaller voltage differences, shorter PWM times, and higher frequencies. This approach also improves the accuracy of the resistor network, enhances the voltage comparator, and reduces component leakage current to ensure controllable errors.

[0261] Method 2 involves changing the single-layer fully connected neural network into a multi-branch tree structure, i.e., a tree structure, and using linear activation functions in the intermediate layers (i.e., using linear activation functions at the multi-branch points).

[0262] Method 3 involves calculating the multiplication and accumulation of a finite number of links each time (multiple sets of finite numbers can be calculated simultaneously), and then summing the results at the end. No network design modifications are needed. Although it's fully connected, the output is summarized in the form of an addition tree. This is a multi-branch tree of addition (addition can be calculated using multi-layer networks in our invented circuits, or digital circuits, with the values ​​of each output module used for data transfer and shifting between modules in a tree structure for addition and subtraction). This also brings the advantage of network flexibility; it can function as both convolutional neural networks and fully connected neural networks, with high accuracy under controlled conditions, and the total number of bits can be flexibly defined. Furthermore, this software design does not affect the circuit function of the basic computational units of neurons; it only depends on the number of neuron links—meaning it's an adjustment to the computational program. However, when using Method 2, it's important to understand that in fully connected networks, the input terms are shared, but the computation is performed with a finite number of links; in convolutional computation, the input terms are independent, but the weights are shared. For a fully connected system, as shown in Figures 22 / 23 / 24, the segmented neuron input current modulation control module group (200) with a finite number of members has neuron links, that is, through the weighted current modulation control module corresponding to each neuron, it is linked to other modules (202) and (205) of the output of a single neuron, including capacitor module, capacitor control module, segment neuron output module and neuron summary calculation output module, etc. The multiple inputs and their multiple input terms PWM of the first weighted current modulation control module group are sent to the corresponding weighted current modulation control module of the neuron. The weights of each link in each group are different. For example, the segment weights (201) of the first link group are different from the segment weights of the second link group. After two rounds of calculation, it is the turn of the second weighted current modulation control module group (204) to perform calculation, and then the third, fourth and so on to perform calculation. With 4 input groups, each group containing n inputs, and each input undergoing 2 rounds of calculation in 2 segments, the current modulation control module is divided into 4*2=8 weighted terms across 4 segments. Through 16 capacitor modules (positive and negative 8x2=16) and a capacitor control module, the final result is obtained through a summation calculation output module (one neuron per group), totaling 4 neuron summation calculation output modules. This completes the calculation for that neuron block. Afterwards, the input values ​​of the first neuron block (210), the second neuron block (207), and subsequent neuron blocks (208) are cyclically transferred in their corresponding fast access units (206). This calculation is repeated, with the final result based on the previous calculation plus the accumulated value from subsequent calculations, until the loop ends. The final summation value of each neuron summation calculation output module, after activation, is the final output value of the fully connected layer.

[0263] Because the input is also 4-bit segmented, the generated PWM only has 16 values, which has a very low resolution. Therefore, it is convenient to use the method of inserting delays by circuit elements to generate PWM instead of according to the base frequency, so that a higher frequency PWM can be obtained.

[0264] This embodiment can simultaneously compute convolutional and fully connected networks through more complex transfer mechanisms. The ratio of neuron inputs to outputs and the number of groups / segments can be set according to actual needs. Directly computing activation functions using a high-performance general-purpose CPU is also a feasible solution, suitable for mixed-precision computation and network training.

[0265] To illustrate the characteristics of arbitrary bit length, as shown in Figure 18, the fast access unit in this section has 8 bits, divided into two segments: the high 4 bits and the low 4 bits. The first input current modulation control module (162) generates two PWM signals: the first calculated high 4-bit input PWM (160) and the low 4-bit input PWM (161). The two PWM signals are sent separately: the high 4-bit input PWM is sent to the first neuron first chain weight current modulation control module (164) and the second neuron first chain weight current modulation control module; the low 4-bit input PWM is sent to the first neuron second chain weight current modulation control module (165) and the second neuron second chain weight current modulation control module.

[0266] The value of the weight term fast access unit of the weight term current modulation control module is an 8-bit value. The resistance value of the corresponding resistor network is selected and generated, but it needs to be multiplied with the input term PWM. Therefore, the first neuron contains two sets of weight term current control modulation modules, positive and negative capacitor modules and capacitor controllers, segment neuron output modules, etc. The weight values ​​of the two sets of weight term current control modulation modules are the same (multiplied with the two 4-bit input terms PWM respectively). The second neuron operates in the same way. This embodiment implements the discharge mode. The weight term current modulation control module itself has a grounding channel (163). Therefore, the channel from the input term current modulation control module to the weight term current modulation control module is the channel for PWM and other information and control. The channel from the weight term current modulation control module to the capacitor module is the actual discharge current channel. Then, the segment neuron output module generates the time difference value through the charging and discharging time difference and sign of the positive and negative capacitors, as well as the timing signal of the layer activation function source (168). The neuron summary calculation output module obtains the cumulative sum value through shift addition. If normalization is necessary, it is also performed through the network aggregation layer (169), and then normalization and secondary activation are implemented in the same way as other embodiments that require normalization. Here, the first neuron and the second neuron are the same, and the calculation of the 8-bit value can be completed in one round of calculation.

[0267] Example of training a backpropagation network using hybrid CPU computation* Backpropagation and forward propagation both involve the same process of accumulating product terms. The inverse activation function is a function related to the original activation function and its derivative. It's necessary to simultaneously perform the accumulation calculation of product terms during backpropagation, calculate the change in weight terms, and adjust the values. Backpropagation from the current neuron output value to the weight terms typically requires calculating the partial derivative of the weight terms in conjunction with the learning rate. Other particularly complex backpropagation methods can be found in the following calculation process, using the CPU and the circuitry of this invention for collaborative computation. Basic calculation of backpropagation: [Partial derivative of weight term (single link) DW] = [Partial derivative of neuron output (error gradient) DO] * [Derivative value of activation function DF] * [Input value of neuron link DX]. The learning rate η can be adjusted either during the total loss / error calculation or during parameter updates, depending on the chosen parameter adjustment scheme. (The calculation of second-order backpropagation and second-order partial derivatives for learning rate adjustment during parameter updates, or other adaptive learning rate optimization algorithms, are similar. If the calculation process is too complex, it can be handled by the main control CPU; details are omitted here.) The propagation process for the neuron output partial derivative DO is similar to that of forward propagation, only reversed. The neuron link input value DX is the output value of the previous layer's neuron cached during forward propagation. The product of the neuron output partial derivative DO, activation function derivative DF, and neuron link input value DX should be calculated during backpropagation at this layer. This saves data caching because during forward propagation, only the neuron output partial derivative DO, activation function derivative DF, and neuron output value need to be cached. The previous example of bit-by-bit segmentation has already been described, as it incorporates a shift adder. Following digital circuit principles, multi-cycle multipliers are also based on shift adders; single-cycle tree multipliers also utilize shift adders. In this example, non-multiplicative-accumulator multiplication calculations and function lookup interpolation calculations use digital multipliers, while bit-by-bit segmentation summarization in multiplicative-accumulator calculations also utilizes a shift adder—that is, part of the CPU's computational function (shared by multiple neurons).

[0268] The parameter update method used in this embodiment is mini-batch training. This involves backpropagating multiple sets of independent data to calculate the adjusted weight parameters, caching them first, and then updating the weights after summing and averaging. This embodiment utilizes multiple circuit blocks working collaboratively, with each local circuit operating only dynamically at a single layer. Taking a 64x64 dynamic layer as an example (the number of neurons in the hardware can be expressed in more layers through software control and composite calculations), it has 64 neurons, each neuron has 64 connections, the weight parameters are 64*64 units, and the neuron output value is 64 units. If the number of mini-batch training batches is 64, then the required cache size for the neuron output values ​​is 64x64 units.

[0269] Figure 31 shows the backpropagation data flow of a neuron in the backpropagation calculation process. A certain neural network has multiple sets of data, each of which performs first-order backpropagation independently. The partial derivative of the neuron output to a certain neuron in the previous layer, DO, is calculated by multiplication and accumulation. For example, the backpropagation error data of the first set (320) -> the backpropagation error data of the first set of the next layer (323), the backpropagation error data of the second set (321) -> the backpropagation error data of the second set of the next layer (324)... the backpropagation error data of the fifth set (322) -> the backpropagation error data of the fifth set of the next layer (325)... up to the nth set of backpropagation error data -> the backpropagation error data of the nth set of the next layer. The vertical dashed line represents the n independent backpropagation neuron output values ​​o1, o2... o5... on of a neuron in the next layer; the n independent activation function derivative values ​​j1, j2... j5... jn of this neuron (derived from the independent output values ​​of the forward propagation neurons); and the cached output values ​​x1, x2... x5... xn of the second layer below corresponding to the input value of a link of this neuron. First, we use a digital multiplier to calculate o1*j1, o2*j2...o5*j5...on*jn, which equals k1, k2...k5...kn. Then, we use the circuit of this invention to perform a multiplication-accumulation calculation, i.e., S(328) = k1*x1(326)+k2*x2...+k5*x5(327)+...kn*xn. In this way, the adjustment value S of the weight parameter of a certain neuron and a certain link in the backpropagation calculation of a small batch of n sets of data is calculated.

New weight value Wnew

Old weight value Wold

Learning rate η

Old weight value Wold

Learning rate η

Old weight value Wold

[0270] Figure 33 is a schematic diagram of a simple dynamic single-layer calculation circuit. The multiple multiplication and accumulation calculations mentioned above can be performed using this circuit. More complex functions can be combined with other disclosed embodiments. The input current modulation control module (306) has a right-hand signal path (309), a first PWM signal (307) segmented by bit, and a second PWM signal (308), which are sent in two rounds in parallel to multiple rows to the weighted current modulation control modules (305) of the corresponding rows of the array. Input data is read from a fast-access memory. The positive and negative capacitor modules (303) each have current channels (304) that are directed column-wise to multiple weighted current modulation control modules. In discharge mode, the charge in the capacitor module is released back to the common ground of the capacitor module through the multiple weighted current modulation control modules of the corresponding columns, as well as the PWM switches and resistor networks set therein. The resistor network can be set with multiple segmented paths; the circuit cannot be clearly contained in the figure, so details are omitted. The weight parameters of the resistor network are pre-read from the memory module. There is a positive / negative capacitor control module (302) and a neuron output module (301). Then, as in other embodiments, the neural network data propagation process includes network discharge, detection discharge, and activation, and finally the data (300) is output to a fast-access memory module. The training of the neural network is backward propagation, so compared with forward propagation, the layers of the neural network corresponding to the input and output are reversed.

[0271] As shown in Figure 32, taking the design of a fast access array memory (SRAM, etc.) as an example, in order to achieve parallel read and write, if the number of data output by the neurons in the physical layer of the circuit is different from the number of data (bits or bytes, etc.) in a row of the array memory, for example, the amount of data in a row of the array memory is several times the amount of data in the neurons, or, the output of a neuron in one round of forward and backward propagation is 8 bits, while the array memory is 32 bits, then the known processing is multiple read and write operations, bit masking / bit selection, selectively writing / reading specific data addresses in the selected row (316), and in more complex cases, a value may be distributed bit by bit in multiple rows, but the details are omitted here. The key point is that there is a certain number of transpose read or block read circuits in the array memory (in the process of multi-chip collaborative computing, the single-chip array memory only needs to meet the computing needs of a small number of layers corresponding to the dynamic layer). In mini-batch training, it is assumed that the data of each layer in the forward propagation has been filled. For example, several rows of data (313) and (314) that are close to each other are the output data of the same neuron in the same neural network in multiple batches of forward propagation. The number of rows in this data block, without considering masking and selective read / write, is the number of training batches. The block and row selection module (317) has the function of enabling row or block selection; the column read / write and enable (311) has the functions of enabling, reading / writing by row, and reading / writing by data bit masking. The column read / write enable (312) and block selection (315) together enable the transpose / block read / write function, which can (once or multiple times) read and transpose each batch of data in parallel and apply it to the corresponding neuron input current modulation control module (column to row within the block), that is, perform multiplication and addition calculations according to the data flow process and order described above, that is, read the data into the input of the dynamic layer in sequence according to the requirements or program settings, and save or output the layer output results to the outside. Finally, the entire backpropagation process is completed.

[0272] *An Example of Detecting Accelerated Discharge*: The charging and discharging of neural network links has great potential to increase the frequency, but the comparator of the capacitor control module has a limited operating frequency. At high frequencies, the smallest unit of delay chain time conversion cannot match the resolution of capacitor discharge time detection (the time resolution requirement of the link pool node is much higher than that of the input and weights). As a result, the capacitor, capacitor control module and neuron output module need to be matched with a higher ratio to balance the operating frequency. The neuron output module is already complex, which takes up a lot of chip area.

[0273] The discharge detection process can be accelerated by adjusting the resistance value of the discharge detection resistor. The principle is as follows (V0 > Vb > Vc): In the first case, there is no acceleration. A capacitor C, with an initial voltage V0, discharges to ground through a resistor R. Discharge stops when the voltage drops below Vc. The detection time is T. In the second case, there are two discharge resistors, R and R / k2, which start discharging in parallel simultaneously. When the detection voltage Vb is reached, resistor R / k2 stops discharging. When the detection voltage Vc is reached, all resistors stop discharging. Therefore, two discharge times, Tb and Tc, are obtained. T can be expressed using Tb and Tc: T = k2 * Tb + Tc. Derivation: Tc = t0 + t1, Tb = t0. The capacitor discharge voltage V(t) = V0 * e^(-t / (RC)), and the two-stage discharge V(t1) = V(t0) * e^(-t1 / (RC)) = (V0 * e^(-t0 * (k2+1) / (RC))) * e^(-t1 / (RC)) = V0e^(-(k2 * t0 + t0 + t1) / (RC)). Therefore, T = k2 * t0 + t0 + t1, T = k2 * Tb + Tc. Based on the property of exponential multiplication, it can be deduced that the relative delay or time shift of the discharge of the two resistors and their respective lengths do not affect the formula. We construct one (or more) reference neurons (352), control their inputs, let the single-pulse time standard output of their single-symbol capacitor control module be Tu, and let the number of parallel detection discharge resistors be j, then the actual pulse width time is Tu / j. This time is used to control the discharge time of the preceding resistor R / k2, i.e., Tb = Tu / j, T = k2 * Tu / j + Tc. Assume the target (binary number) T = 0B1001. If Tu = 0B0010, then control the switching of the two resistor networks, setting k2 / j = 4, so Tc should be 1. Tc is the actual time value obtained from the accelerated discharge detection in case two. T is the value we actually need to calculate; what we actually need to do is calculate T using Tc and the preset time Tu. Vb is a concept we set for ease of understanding; in reality, there should be a Va greater than Vb. We check if T is greater than 0B1000 + margin; if it is large enough, we execute accelerated discharge detection. Now, extending to the case of accelerated discharge in n stages, the actual discharge resistance of each stage is R / (1+k2+k3+ … +kn), … ,R / (1+k2+k3),R / (1+k2),R; similarly, it is easy to obtain T=kn*T1 / j1+k[n-1]*T2 / j1 … k2*T[n-1] / j[n-1]+Tn / jn.If the comparator determines that the current detection capacitor voltage is sufficiently large, then by using multiple standard neurons to generate standard output pulse widths (e.g., 0B000001, 0B000100, 0B10000...) through standard inputs, and controlling the number of parallel connections of kn, discharge can be accelerated in multiple stages with different limit safety resistance values. After obtaining the positive and negative time differences Tn, the logic circuit then adds back the predetermined standard value according to the sign based on the accelerated discharge operation of the existing positive and negative capacitors. This allows for obtaining the numerical conversion value of the time difference with high accuracy and frequency.

[0274] This embodiment uses a single-pulse pulse width time input, as shown in Figure 19. The main capacitor (365) in the capacitor module can be turned on by one of the multiple detection capacitors (366) and (367) through switches (363) and (364). In some embodiments, the detection discharge can be operated in parallel with the pre-charging (361) or the charging and discharging process of the network node (362), increasing the overall operating frequency. In this embodiment, multiple detection capacitors are used independently for voltage comparison without interfering with each other, improving the accuracy and performance of the comparator, and setting the priority comparison order. The capacitor control module with the same sign has multiple sets of multiple voltage comparators (358) (different sets can be comparators of different types and performance), and each comparator can be connected to different reference voltages (359) through switches. During the design, the reference voltage can be calculated back from the binary discharge time to determine whether the initial voltage is greater than a certain voltage corresponding to a binary output pulse width value. The first set of comparators and the first spare detection capacitor can obtain the approximate voltage and sign before the second set of operations, and submit the results to the control logic (360) and the control module (357). The second set of comparators and the second spare detection capacitor can be compared with multiple reference voltages in real time during the actual discharge detection process. This design allows the control logic (360) (357) to determine the minimum resistance or maximum discharge current of the resistor network (356) in each accelerated discharge stage of the discharge detection process, as the voltage of the detection capacitor changes before or during the discharge detection. The comparison results can also be used to assist the operation of other circuits. Parallel resistors (355) (354) with discharge switches (356) are usually conveniently used for the resistor network. The operation time for detecting each resistance value of the accelerated discharge comes from the single-symbol pulse width time output (353) of the reference neuron. The reference neuron can reduce some components as needed compared to the normal neuron. The reference neuron (352) has a similar structure to other neurons, with a capacitor module (351) and the same input link node (350), but it does not need positive and negative capacitors (positive and negative can also be set to reduce errors), and it outputs a single pulse directly from the single-symbol capacitor control module. Its neurons use independent standard inputs and standard weights. There are multiple reference neurons, corresponding to different common standard pulse width time outputs. The output of the single-symbol capacitor control module of the neuron, after being adjusted by the control logic based on the pulse width and time of the reference neuron's output (i.e., after accelerated discharge), is then, as in the previous embodiment, calculated to obtain the positive and negative time difference, i.e., the digital value of the XOR pulse width. Then, according to the accelerated discharge strategy, the known preset time value that was reduced during the acceleration process is added back. If necessary, activation transition is calculated, and finally, the output value is generated.

[0275] For analog layers with single-pulse input and output, a multi-channel output accelerated discharge design can be used. This involves a capacitor controller with multiple independent resistor networks (equivalent to shift operations, corresponding to different bits) of varying resistance values ​​for detection and discharge. Similarly, the neuron output module has multiple independent operation modules. The output pulses of the capacitor controller (with different shift operations) undergo positive and negative XOR logic synthesis, and finally, the pulses (with different shift operations) are output independently, meaning multiple lines (with different bits) synchronously output single pulses. Correspondingly, each input item and neuron connection in the next layer also has multiple independent current channels (and independent equivalent resistances) to cooperate with the operation of the current layer. The resistor networks in the next layer corresponding to different shift positions for input / weight / detection processes can have their resistance values ​​or time lengths modified as needed to perform reverse shift corrections, or pulses of different bits can be operated in separate rounds. The so-called round operation means that the output pulses of the upper layer (equivalent to different shifts) are output in batches, and the weighted resistor network of the link current channel of the lower layer is also shifted according to the batch or read from the fast access memory and reloaded after shifting.

[0276] *An Example of a Multi-Layer Dynamic Network with Analog Intermediate Layers* Based on the above example of a single-layer dynamic network, instead of calculating the weight parameters every time they are read from memory, multiple layers are used, as shown in Figures 25 and 26. Data between layers is processed like an assembly line. There are two digital layers at the top and bottom, and one or more digital / analog layers in between. Although the calculation frequency of each layer remains the same, multiple layers are processed at once, thus doubling the computational performance. Moreover, the intermediate layer between the output of the current layer's neurons and the input of the next layer's neurons can be optimized. If no training is required and the accuracy requirement is not high (or the network is trained to be insensitive to accuracy), then the first and last two layers (385), i.e., the input and output layers of the entire computational circuit, are digital, while the output of the intermediate layer (384) is a single XOR pulse (408) and a sign level generated directly by capacitor discharge, and the input is a single pulse. Because the intermediate layer omits digital conversion, its operating frequency, area, and energy consumption can be further optimized. In addition to using linear pulse output, the activation function of the intermediate layer can also use other simulated nonlinear outputs constructed based on the principles of tp4-tn4, tp3-tn3, tp2-tn2, etc., which can limit the maximum output value range and ensure the stability of the single pulse width time range (see Figure 28 for subsequent related embodiments). During external training, the above nonlinear functions can be used for training, and fluctuation errors are added to the neuron output during training to adapt to the production deviation of hardware parameters. The network has a residual calculation circuit (383) that can be matched with a residual network. A single intermediate simulation layer has a self-looping information channel from output to input (382), which can also be used as a dynamic layer for single pulse simulation, that is, it can be used as a different layer by changing the parameters. In addition, the connection method of the intermediate simulation layer can be further shown in Figure 27 with a more complex network structure (386). The thick lines in the figure represent multi-line signals, that is, single pulse signals or digital signals of multiple neuron inputs and outputs (380)(381). As mentioned earlier regarding the capacitance exponential multiplication-addition effect, using single-pulse signals, even if out of order, does not affect the calculation results. The XOR single pulses and signs of multiple neurons output from multiple intermediate analog layers are sequentially and time-divisionally input to the next layer at the node, thus aggregating the outputs of multiple layers to the input of the same layer. Of course, the transmission between layers requires a matching number of neurons. In this embodiment's complex network structure, single pulses are used as input and output between upper and lower analog intermediate layers. The working mode is as follows:

[0277] As shown in Figure 25, the capacitor module contains two main capacitors (400) and three detection capacitors (401). Switches (402), (403), (404), and (407) control the connection relationships between each capacitor and the neuron's charging / discharging link node (405), the pre-charging node, and the capacitors themselves. This embodiment has three parallel connection states: 1. (Pre-charging) Main capacitor + detection capacitor; 2. (Network discharging) Main capacitor + detection capacitor; 3. (Detection / discharging) Detection capacitor. During the parallel processes of pre-charging and network discharging, a certain main capacitor and a certain detection capacitor need to be connected together. Pre-charging requires connection to the pre-charging port (408), and network discharging requires connection to the neuron's charging / discharging link node. There is also a parallel state where the detection capacitor works in conjunction with the capacitor control module and the neuron output module (411). Because multiple sets of capacitors work in parallel, the single-pulse output representing the analog signal can be given to the input of the same layer. Meanwhile, more detection capacitors and connection switches can be set, presenting detection capacitance values ​​in binary multiples of different proportions. By selecting the capacitance value, the binary multiple amplification rate of the detection discharge time can be controlled. That is, the shift operation of the XOR single pulse pulse width time output. The comparator of the capacitor control module can be set with multiple reference voltages (406), suitable for neurons with different numbers of links (the maximum value of the main capacitor voltage change is different). In addition, the voltages of the positive and negative capacitors can be compared first, so that the control logic (413) obtains the positive and negative signs (410) before the XOR pulse output (409), rather than afterward. Similarly, the setting and selection of the resistance value of the detection resistor (414) can also control the binary multiple amplification rate of the detection discharge time. The circuit for generating the negative pulse (412) in the figure is similar to that of the positive pulse and is omitted. A challenge in chip manufacturing is the complexity of using bit-by-bit segmentation for analog intermediate layers. Therefore, in the binary weighted resistor network of the current modulation control module, if the number of bits is large, the resistance value increases exponentially bit by bit. This necessitates the use of resistors with different resistivities and corresponding manufacturing processes, resulting in significant errors in the resistor values ​​within the chip due to manufacturing deviations in resistivity. Through repeated testing with specific weight parameters and standard inputs, we can determine the error resistance and error resistivity of each bit by detecting and statistically analyzing the pulse width (internal or external) of the linear output of this layer's forward propagation. This allows us to correct the weight parameter values ​​based on the detection results using the main CPU or external calculations before applying the weight parameters. Note: Considering chip / circuit area limitations, the functions of the digital and analog layers can also be combined into a single dynamic layer that can be used for both analog and digital applications.

[0278] *Example of Non-Linear Activation Function with More Complex Analog Output* , when a single pulse width time is used as the analog output, if it is a linear activation function with an output range like ReLU6, the present invention only needs to simply shield the negative time difference pulse output part. More complex function output can be realized through secondary charging and discharging, that is, using the single pulse generated by the time difference level of positive and negative capacitance detection, the (additional) detection capacitor is charged and discharged again through a standard equivalent resistance (only a single sign, one of the positive and negative capacitors is required), and then a new detection charging and discharging process is re-entered, and a new single pulse corresponding to the activation function output value conforming to software habits is generated. tp4-tn4, tp3-tn3, tp2-tn2 and the like have been proved through formula derivation that, with the arrangement of charging and discharging steps and the adjustment of voltage parameters, various forms of non-linear analog single-pulse activation functions can be generated through simple capacitor charging and discharging. Furthermore, segmented activation functions can be generated through the combination of voltage comparison logic and charging and discharging of multiple detection capacitors. Based on the exponential multiplication characteristic of capacitor charging and discharging, (e^a)*(e^b)=e^(a+b), even if there is truncation or out-of-order between the segments of the output segmented function single pulse, it will not affect the subsequent calculation results.

[0279] As shown in Figure 28, taking tp2-tn2 as an example, tp2-tn2 = cap*R*(ln[2-e^(-t*SSP(...,Xn*Wn))] - ln[2-e^(-t*SSN(...,Xn*Wn))]), if there is no input to SSN, then tp2-tn2 = cap*R*(ln[2-e^(-t*SSP(...,Xn*Wn))]. This is a function similar to tanh, which passes through the origin and has an origin derivative of cap*R. Let x=-t*SSP(...,Xn*Wn) and a=cap*R; construct f(x)=a*x, x<=b;h(x)=a*(ln[2-e^(-(x-b)))]+a*b,x>b; these two functions can form a first-quadrant piecewise function g(x) = IF[x<b, f(x), h(x)] , and through the selection and setting of signs, there can be origin-symmetric first-quadrant and third-quadrant curves, or coordinate axis-symmetric curves, or a single quadrant curve can be shielded / selected through sign setting (450).

[0280] As shown in Figure 29, the piecewise function hardware circuit includes a more powerful neuron output module (430). It incorporates the function of generating time pulses by charging and discharging the original integrated positive / negative capacitor control module, i.e., outputting XOR time difference pulses (434) to the signal synthesis selection logic circuit (439). It also includes a secondary charging and discharging conversion circuit. The working process is as follows: when XOR is logically 1, one of the positive and negative capacitors has stopped discharging, while the remaining positive / negative capacitor continues to discharge (discharging mode). The positive / negative capacitor control module has multiple voltage comparators for comparing multiple standard reference voltages (435), one of which is the standard reference voltage at the piecewise function dividing point (greater than the cessation comparison voltage), dividing the discharge process of the remaining positive / negative capacitor into two parts: linear segment discharge and nonlinear segment discharge. Of course, if the voltage of the remaining discharging positive / negative capacitor is less than the standard reference voltage at the piecewise function dividing point when XOR is switched to 1, then it only discharges in the linear segment. That is, the comparison result (438) with the standard reference voltage is also sent to the signal synthesis selection logic circuit in real time. The signal synthesis and selection logic circuit, based on the comparison result, divides the XOR time difference pulse into pulse time outputs (442). The linear segment outputs directly, while the nonlinear segment controls the connection of the charging and discharging switch (assuming pre-charging has been completed) to the activation capacitor (432) and the charging and discharging voltage source (431), and performs nonlinear conversion through the activation capacitor control module (445). Similar to the standard positive and negative capacitor process, this stage also has a reference capacitor (433) and a reference capacitor control module to counteract various effects such as leakage current and charge loss injection. The time difference (444) (443) between the two in the nonlinear stage is also input to the signal synthesis and selection logic circuit, which then outputs the pulse time (442). As mentioned in other previous embodiments, due to the calculation process of the exponential multiplication of capacitor charging and discharging, the division and disorder of single-pulse charging and discharging do not affect the calculation of the time difference result of dual-capacitor discharge detection. Therefore, the pulse time outputs of the linear segment and the nonlinear segment are equivalent outputs of the piecewise function pulse time. Similar to other embodiments, this neuron output module also outputs a positive or negative sign level (441) based on the comparison between the positive and negative capacitor voltages. Mathematically, the mathematical relationship between the values ​​before and after activation can be selected in a form similar to Figure 28, tp2-tn2, or a symmetrical variant based on multiple quadrants with a sign setting. Note: In large-scale model applications, the single-clock analog output can also use additive attention methods to replace the attention calculation process based on digitization and matrix dot product in the previous embodiments.

[0281] Furthermore, because the key circuit is an analog signal circuit based on single-pulse pulse width time calculation, it cannot adjust the positive and negative deviations (originating from deviations such as production, voltage, and temperature) through the results of digital time conversion and numerical detection deviations (stored in registers, etc.) as in other embodiments. However, it can adjust the positive and negative deviations using positive / negative input adjustment nodes (436). The input adjustment node and the positive / negative current channel node (437) of the neuron link input are the same circuit (similar to the neuron bias term, multiple links can be set according to adjustment needs).

[0282] Similar to a biological spiking neural network*, biological neural networks are constantly operating. Charging current and leakage current simultaneously act on the neuron capacitors. When the neuron capacitor voltage exceeds a threshold, it is activated, and the output is determined by the frequency of the activated neuron's output and the number of excited neurons. Our circuit, after modification, can achieve the same function. It compares the voltages of the positive and negative capacitors, and the comparison result is output to the signal synthesis and selection logic. Simultaneously, the positive / negative capacitor control module no longer outputs discharge pulses but instead uses a predetermined voltage source connected to a large resistor as active leakage current. The neuron links charge and discharge the positive and negative capacitors respectively (in the opposite direction to the leakage current). When the positive voltage minus the negative voltage exceeds a specific voltage threshold, the circuit emits a fixed-width positive pulse signal (conversely, when the negative voltage exceeds a certain threshold of the positive voltage, a fixed-width negative pulse signal is emitted). If training makes the input range of the corresponding layer greater than or equal to 0, and only outputs a fixed-width positive pulse signal, then it is closer to a biological neural network. This signal can be connected to the single-pulse analog input embodiment disclosed herein (in the form of a ratio based on the total pulse length or the ratio of the number of capacitor handling operations, etc.), meaning that the input and output of two circuits with different operating principles can be connected (it can be used for a brain-computer interface).

[0283] The aforementioned circuit, which expresses output based on frequency modulation and the number of excited neurons, resembling a biological neuron, is inefficient compared to other embodiments using single-pulse pulse width-time simulation. It suffers from large errors, low operating frequency, high computational unit / resource consumption, weak activation function, and difficulty in training. Its only redeeming feature is sparse network computation; however, this function can also be achieved using gating circuits. The enable-type gating circuit described below is different from embodiments such as the gated activation function SwiGLU*, and does not generate a gating signal in subsequent calculations of detection and activation. As shown in Figure 29, the positive and negative voltage comparator (446) compares the positive and negative capacitor voltages after the two devices of the neuron network are charged and discharged and before the detection phase begins. At high frequencies, this is actually performed synchronously with the detection, that is, the sign of the neuron input multiplication and accumulation value is quickly obtained and used for enable and gating. On the one hand, it can quickly enable the subsequent circuits such as (445), (444), and (431) inside, reducing energy consumption; on the other hand, it can be used for enable and gating (440) of external circuits. For example, the gating signal is output to other neurons to shield the output of external neurons or turn off their working state. Furthermore, the gating and enable signals can be stored in the fast access unit, and the working state of subsequent neurons can be controlled according to the read gating and enable signals.

[0284] *Fully Connected Segmentation Calculation Example* For fully connected systems, as shown in Figures 22 / 23 / 24, the segmented neuron input current modulation control module group (200) with a finite number of members has neuron links, that is, through the weighted current modulation control module corresponding to each neuron, it is linked to other modules (202) and (205) of the output of a single neuron, including capacitor module, capacitor control module, segment neuron output module and neuron summary calculation output module, etc. Multiple inputs and their multiple input terms PWM of the first weighted current modulation control module group are sent to the corresponding weighted current modulation control module of the neuron. The weights of each link in each group are different. For example, the segment weights (201) of the first link group are different from the segment weights of the second link group. After two rounds of calculation, it is the turn of the second weighted current modulation control module group (204) to perform calculation, and then the third, fourth and so on to perform calculation. With 4 input groups, each group containing n inputs, and each input undergoing 2 rounds of calculation in 2 segments, the current modulation control module is divided into 4*2=8 weighted terms across 4 segments. Through 16 capacitor modules (positive and negative 8x2=16) and a capacitor control module, the final result is obtained through a summation calculation output module (one neuron per group), totaling 4 neuron summation calculation output modules. This completes the calculation for that neuron block. Afterwards, the input values ​​of the first neuron block (210), the second neuron block (207), and subsequent neuron blocks (208) are cyclically transferred in their corresponding fast access units (206). This calculation is repeated, with the final result based on the previous calculation plus the accumulated value from subsequent calculations, until the loop ends. The final summation value of each neuron summation calculation output module, after activation, is the final output value of the fully connected layer.

[0285] *An Example of a Computational Circuit System Based on Wafer-Level Stacked Bonding* The main multiplication and addition calculation of the circuit shown above relies on the theorem based on capacitance and the relationship between equivalent time length and capacitance value that we discovered. Moreover, it is closely integrated with SRAM, RRAM, etc., so the power consumption is extremely low. Therefore, it is very suitable for high-density stacking.

[0286] Stacked wafers perform even better when combined with microfluidic liquid cooling, as shown in Figures 1 and 2. The planar liquid flow main channel 902 connects to the vertical liquid flow main channel 903. There are multiple vertical liquid flow main channels because they need to connect multiple wafers, so the flow rate is much greater than that of the planar liquid flow main channel. Furthermore, each vertical liquid flow main channel flows in the opposite direction to its adjacent vertical liquid flow channels; that is, the liquid flow direction of the vertical liquid flow channels is two-dimensionally intersecting at 906 and 907. After wafer-level stacking and bonding, the wafer module only needs to collect the channel coolant according to the direction of the vertical liquid flow main channel and generate positive and negative pressures. The section 908 between the intersecting vertical liquid flow main channels with opposite flow directions in the planar liquid flow main channel naturally has positive and negative pressures, generating corresponding liquid flow and carrying away the operating heat. The planar liquid flow main channel in the figure coincides with the die partition line because this computing circuit generates relatively little heat. If it were partitioned, it would not require liquid cooling and could dissipate heat naturally. However, this does not mean that there cannot be other cooling liquid flow channels besides the planar liquid flow main channel. To prevent external liquids or liquid channels from affecting the wafer working area, the planar liquid flow main channel should be placed on the back side of the wafer, in addition to being located in the die segmentation area.

[0287] As shown in Figure 1, wafer 900 contains numerous undivided die arrays 901. Each die may contain, as needed, complete neural network computing circuitry, including power supply, communication, storage, general-purpose computing circuitry, neural network digital computing, and analog computing circuitry. The back-side bonding pad 904 and the front-side bonding pad 905 are among multiple communication bonding pads on both sides of the die. For ease of explanation, other bonding pads are not shown in the figure. Communication bonding pads, through line selection circuitry, can selectively connect to computing circuits of adjacent dies or to other upper / lower bonding pads on the same die. That is, bonding pads can establish independent or shared data channels with neighboring dies (and of course, the computing cores within neighboring dies can also communicate independently). Therefore, if a damaged die is encountered in a stacked network of wafers, that die can be shielded, and the normally functioning dies and data flow cells of the stacked wafers can be moved as a whole to surrounding normal data flow cells, like sand. Specifically:

[0288] 1. After the wafer is tested or self-tested, the chip damage information is obtained from the bus or other communication links.

[0289] 2. Select wafers with a similar number of damaged dies, i.e., a consistent ratio of normal working dies to yield. We use these well-matched wafers with consistent yield for stacking and bonding, layer by layer. Wafers that do not meet the wafer consistency requirements or the requirements for the number of neuron layers in the large model can be separated into dies for other uses. As shown in Figure 3, the upper, middle, and lower wafers 910, 911, and 912 prepared for stacking and bonding are all selected wafers with consistent yield and a consistent number of normally functioning dies. The crossed symbols in the figure represent damaged dies, located at different positions on the upper, middle, and lower wafers.

[0290] 3. Use a program to calculate the data flow migration path for each unit, determine the feasibility of migrating data flows between upper and lower layers, and ensure proper matching. A batch of wafers with consistent yield yields contains many wafers; we select a few for use in one of the stacked bonding wafer modules. Specifically, the selection method involves using a computer to virtually test and sequentially select wafers, generating a search tree for the wafer stacking order. For each possible stacking order, connections are made between the upper and lower defect points, and the shortest, non-intersecting path is chosen from among the various paths. As shown in Figure 4, for the lower layer defect 913 and the middle layer defect 914, based on the shortest connection path, the communication link of the first normal die in the path is connected to the communication port bonding pad 916 of the defect, and the second normal die uses the communication port bonding pad 915 of the first normal die. In essence, the communication port of the wafer above the die in the path is linked to the bonding pads of the adjacent dies in the path, according to the path and direction. This allows data to be transmitted (bidirectionally) to the adjacent die in the next layer corresponding to the path. This shifts the overall communication of the dies involved in the lower-layer connection path, resulting in locally misaligned communication between the lower-layer and upper-layer dies. Of course, if the loss of defective pixels is negligible, since only a few dies are damaged, especially when the number of stacked bonding layers is small, it is feasible to directly shield the entire stack based on the coordinates of the defective pixels in each layer. For example, in the diagram, there are three defective pixels, one on each wafer. Shielding would then cover 3 coordinate positions, shielding 3x3 dies. Alternatively, if there are good dies at the shielded coordinate positions, they can be individually separated using laser cutting. Furthermore, if the production volume is high enough and the batch size is large enough, it is possible to select wafers with consistent defect positions from the stack. In short, it is feasible to ensure that stacked wafers can be matched.

[0291] Of course, the bonding pads and communication circuitry of the communication ports may also be damaged, but because of their small area, the yield rate is relatively high. Also because of their small area, multiple sets can be easily installed within the die for switching purposes. In other words, the number of inter-layer communication ports should be multiples of the required number to ensure high reliability.

[0292] 4. Based on the wafer number and stacking order calculated by the program, stack and bond the wafers.

[0293] In some implementations of neural networks, a single die may contain multiple computational units. Each computational unit is equivalent to only one layer of the same depth in the neural network, or a slice of one layer (i.e., some neurons and some connections; digital models are subsequently summed a second time through digital communication and the CPU, while analog models are summed a second time through multiple clock communications and multiple capacitor charging and discharging operations). Even if a die contains multiple core dual-capacitor multiply-accumulate computational matrix circuits, the number of neurons and connections in a large model's layers may far exceed what multiple dual-capacitor multiply-accumulate computational matrix circuits can handle in one cycle or one operation. Furthermore, the next layer may also require multiple layers of the current depth. Therefore, horizontal communication and data transfer across multiple dies on the same wafer are essential. Neighboring dies have independent, controllable communication lines (internal die communication naturally outperforms inter-die communication, which will be omitted here), enabling cross-die computational summation. Independent communication is supplemented by a bus and shared communication lines to complete the data transfer between dies.

[0294] As shown in Figure 5, the die array 920 on the stacked wafers is divided into multiple independent working areas 921 as needed. The neural networks within the independent working areas work independently, and their data is vertically transmitted between the die arrays of the upper and lower wafers in the stack. The stacked wafers may be classified as digital or analog.

[0295] The digital capacitor multiply-accumulate circuit uses digital signals for input and output. The output of the lower layer is vertically transmitted to the current layer, becoming the input of the current layer. The input of the current layer is transmitted horizontally in a ring (adjacent dies have independent communication channels). Each time the input is updated, the output value is calculated and accumulated according to the network structure. Then, activation / gating / residual operations are performed before the data is vertically transmitted to the next layer for subsequent calculations. Because horizontally adjacent dies have multiple channels, single-row or single-column dies can also be used, cyclically transmitting data in independent rings (923).

[0296] Analog capacitor multiply-accumulate circuits use analog signals for input and output, with a single pulse width and duration. Storage and addition are based on capacitance, and the temporary storage time is affected by leakage current. Therefore, communication between upper and lower layers of the working area requires channel link selection and time-division switching. Horizontal communication needs to be independent and parallel, as well as time-division traversal. That is, a die within the working area needs to be able to select and establish / connect to data channels with any die within the working area, i.e., matrix-type connections / star-type selectable connections. Each operation independently connects to different dies without repetition. After multiple operations (after full connection), it's equivalent to transmitting and calculating all inputs. After multiple operations, full communication connection is completed. Of course, the weights are pre-written in the die's internal RRAM / SRAM storage circuits and are loaded when needed.

[0297] In addition, communication between stacked dies can also be one-to-many, that is, one transmitting stacked die synchronously sends data to multiple receiving dies in the vertical direction; or many-to-one, multiple stacked dies send output data to a certain die in the vertical direction in a time-division manner; or even many-to-many, multiple transmitting stacked dies send data in a time-division manner and multiple receiving stacked dies receive data in parallel.

[0298] *An Example of a Multi-Channel, Bit-Segmented Time Pulse Output Calculation Circuit System Based on Neurons* In an analog capacitor multiply-accelerate calculation circuit, combined with the multi-channel output accelerated discharge design in the *Detection Accelerated Discharge Example*, a long single time pulse can be decomposed into multiple short time pulses with independent channels, each corresponding to a different binary bit. That is, the pulse width of each short time pulse represents the value of a portion of a bit segment in a binary number. For example, in the 16-bit number 0001 1100 1010 0101, the value / information of the 1100 bit segment can be transmitted using the pulse width of a single time pulse from one of the channels.

[0299] To address the consistency issue of time pulse resolution subdivision unit standards, i.e., to ensure that the input and output time pulse widths are consistent at the same value, based on the analog numerical conduction relationship between the input time of the single pulse RC discharge and the detection discharge time difference (refer to the simplest embodiment), the unit pulse width between the various computing core circuits located on each wafer can be achieved using cross-wafer pulse input. The output of the internal RC detection discharge time difference of the reference circuit serves as the reference source for the unit calculation of the analog capacitor multiply-accumulate computing circuit. The output of this time difference can be directly used as the input signal or as the standard for internal signal correction of the computing core.

[0300] *An Example of an Analog Computing Circuit System with All Bit Outputs* In the previous example, the analog capacitor multiply-accumulate computing circuit core faced significant challenges in communication transmission. This was because clock signals were difficult to store (capacitor storage was unsustainable due to leakage), resulting in high costs for its time information storage circuit. Furthermore, the need for a time pulse reference per unit time hindered communication. In other words, it relied on real-time communication and struggled with delayed communication. This necessitated more parallel numerical channels and inter-core channels for parallel communication, requiring more horizontal communication lines on the same wafer level – a star or regional star topology. Unlike digital capacitor multiply-accumulate computing circuits, it couldn't perform serial numerical communication, serial inter-core link communication, or two-line / ring inter-core link communication. Star parallel communication was far more complex and costly to design than two-line / ring parallel communication, requiring more independent communication channels.

[0301] The problem can be solved by further modifying the bit-by-bit segmentation into a full-bit output transmission signal. For example, a 16-bit number can be decomposed into a 16-channel multi-pulse signal, or an n-bit number can be decomposed into an n-channel multi-pulse communication. Alternatively, through software and process improvements, the output of a 16-bit neuron can be divided into two rounds of 8-channel multi-pulse signals, or n bits = m*x, i.e., m rounds of x-channel multi-pulse signals. The circuit is similar to that in Figure 19. Through the combination of digital circuits in control logic 360 and analog circuits (multiple voltage references 359 and comparators 358, etc.), the operation of detecting discharge is converted into a specific value in the register, that is, the pulse width value of the time pulse is converted into digital information. Compared to the more complex shared general-purpose digital computing core of digital microcontrollers in many embodiments, this analog circuit is simpler and can achieve higher frequencies.

[0302] More importantly, this design in this embodiment achieves decoupling of the timing of collaboration between computing cores. In addition to the essential power, timing, and shared / bus communication between multiple parallel neural network layers on a stacked wafer, the computation between the analog computing cores corresponding to the neural network subnets, especially at high frequencies, presents significant design challenges if wafer-level collaborative manipulation via timing is required. This embodiment greatly optimizes the design of such collaborative circuits.

[0303] *An Example of Using Switched Capacitors Instead of Resistor Networks* Switched capacitor networks, acting as resistors, can replace conventional network resistors, also known as capacitor-resistors. In our circuit, as shown in Figure 12, a current channel of a single neuron's link is connected to a switched capacitor (423) of different capacitance values ​​via a left parallel switch (424) and a right parallel switch (425). The left and right parallel switches select one path and then alternately conduct, thereby transferring the charge of the capacitor module (422) to the common terminal of the capacitor module. Its effect is equivalent to a conventional resistor single-pulse fixed-width operation, meaning the ratio of charge release is not affected by the main capacitor voltage. In terms of control, the number of times the switched capacitors are transferred is equivalent to the pulse width of a conventional resistor single pulse. Switched capacitor networks can be used for network weights, for detecting discharge resistance, or both. A preferred design is that the neural network input values ​​(421) and (422) control the number of times the switched capacitors are transferred, and the neural network weight values ​​(426) control the selection of the capacitance value of the capacitor network. Together, they control the discharge current of the corresponding symbol neuron's capacitor module. If the discharge detection also uses a switched capacitor to control the current, the pulse width conversion of the time pulse needs to be changed to a switching capacitor transport counter. Note: The switch used in the figure is a high-frequency switch that can reduce charge injection (parasitic capacitance).

[0304] *An embodiment of an interleaved parallel switched capacitor calculation circuit system* The equivalent time can be approximated using the number of operations of the switched capacitors. In chip design, capacitor matching can often achieve higher precision than matching resistors using conventional materials and processes, and is less susceptible to temperature and voltage variations. Furthermore, high-quality, specialized resistors often involve different production processes and management / chip-out cost issues. Using capacitors uniformly can solve these problems, although resistive circuits may offer higher energy efficiency and performance. Switched capacitors operate one-time, making timing optimization easier, and the larger the equivalent resistance of the switched capacitor, the smaller its area, meaning lower cost.

[0305] However, switched capacitors and resistors cannot be simply substituted because neural networks have many connections and different weights. This leads to variations in the charge distribution and release ratio due to inconsistent parallel connections of different numbers of switched capacitors. While the discharge of a resistor is differentiable in time, the operation of a switched capacitor is discrete, unless the switched capacitor is infinitesimally small. It's easy to see that if n equivalent switched capacitors discharge sequentially, the result of discharging the main capacitor in parallel with a real resistor is equivalent; however, if the n equivalent switched capacitors discharge synchronously, the result of discharging in parallel with a real resistor will have errors. To reduce or eliminate these errors, the capacitance value of the switched capacitors must be much smaller than the total capacitance. If we make the ratio of the capacitance value of each switched capacitor to the total capacitance very small, this results in an excessively large main capacitor, occupying too much area and causing the voltage drop across the main capacitor to change too little, ultimately amplifying the threshold voltage detection error. In the circuit of this invention, the switched capacitors are used as equivalent resistors, requiring specific processes and circuit design to achieve more accurate calculations and compensate for the errors caused by the parallel operation of the switched capacitors, because the synchronous discharge of multiple switched capacitors involves complex mathematical relationships. Without changing the capacitance value of the switched capacitor and the ratio of the total capacitance, the following solutions are available:

[0306] First, suppose a neuron has n positive or negative connections, and each connection may have multiple capacitors, for a total of m capacitors. We can group the capacitors corresponding to the connections, and operate on each group sequentially. For example, each group could have m / 8 capacitors. This significantly reduces the error. If we utilize the natural time delay of the control circuit signal to stagger the switching operations of capacitors in different connections, we can also reduce the computational error caused by parallel capacitor operations, for the same reason as grouping capacitors sequentially.

[0307] Second, a design that does not change the capacitance value of the parallel capacitor is used, so that the switching capacitors are connected in parallel without changing their ratio with the main capacitor, thereby compensating for parallel connection errors.

[0308] Assuming the input is 1 bit and the weights are both 2 bits, meaning the capacitor count w corresponding to a single-link weight is 0-3, subsequent calculations using multiple rounds will achieve higher bit values. Each calculation undergoes standardization correction via digital circuitry to avoid accumulated analog calculation errors. The calculation flow and circuit design are as follows:

[0309] A neuron has n links and n input terms, one of two sets of circuits (positive or negative). The input section has one main capacitor C0 and weighted switching capacitors C11 / C12 / C13...C1w, C21 / C22 / C23...C2w, C31 / C32 / C33...C3w, ..., Cn1 / Cn2 / Cn3...Cnw. The detection section has one main detection capacitor Ct and m detection switching capacitors Ct1 / Ct2 / Ct3...Ctm. In integrated circuits, capacitors are assembled from standard capacitor blocks Cs for proper matching. Here, for ease of calculation and matching, the capacitance values ​​of the switching capacitors are all Cs, and the capacitance values ​​of other capacitors are also multiples of Cs. The capacitance values ​​of the input and detection sections are proportional, can be combined and connected, and multiple sets of input and detection sections can be used in succession to increase the operating frequency.

[0310] As shown in Figure 6: The weighted switching capacitors are divided into two types / groups: pre-connection capacitors and post-connection capacitors. During the pre-charging process, the pre-connection capacitors and the main capacitor are charged together to the operating voltage 'a'. The total capacitance of the connection section, CInput, is calculated as C0 + Sum[pre-connection capacitor]. CInput serves as the reference capacitance for operation. The total capacitance of the connection, CInput, is calculated as C0 + Sum[C11 / C12 / C13...C1w] + Sum[C31 / C32 / C33...C3w]..., where C0 includes Ct + Sum[detected pre-connection capacitor]. When each connection discharges, the pre-connection capacitors are disconnected from CInput (after discharge, they are reconnected to the main capacitor in a specific order), while the post-connection capacitors are connected to the main capacitor. This ensures that the total capacitance CInput does not decrease due to the disconnection of the pre-connection capacitors, nor does it increase due to the connection of the post-connection capacitors. The total capacitance remains constant during the switching process. Assuming the quantities of the two types of capacitors are s1 and s2, the capacitance value of the CInput capacitor is (s1-s2)*Cs. Since s1 and s2 partially cancel each other out, the capacitance change is smaller, thus significantly reducing the error. Ideally, the two types of capacitors should be evenly distributed, such as in an alternating arrangement. In this embodiment, one column of capacitors corresponds to a 2-digit weight value, meaning 0-3 operating capacitors are selected based on the corresponding bit range value (0-3). Capacitors C11, C12, and C13 are pre-linked capacitors, which are on at the start of operation of switch 930 in group 1. Therefore, at the start of operation, capacitors C11, C12, and C13 are connected in parallel with the main capacitor C0. Capacitors C21, C22, and C23 are post-linked capacitors, which are off at the start of operation of switch 931 in group 2. The pre-linked and post-linked capacitors are arranged in an alternating column arrangement; for example, groups 1, 3, 5... are pre-linked, and groups 2, 4, 6... are post-linked. Capacitor C0 is functionally connected to other capacitor multiplication and addition circuits (932). The focus here is on parallel capacitor calculations, so this will be omitted and not repeated. The working process of the equivalent resistance of the switched capacitors is as follows: Switch A is used to connect the pre-connected capacitor to the main capacitor, followed by switch B. When A is disconnected, B is turned on. As is known, the switches connecting each capacitor to ground either have corresponding operations before or after (to avoid short circuits). That is, the capacitor corresponding to A immediately discharges to ground, and the capacitor corresponding to B charges from C0. Then, waiting for the capacitor voltage to balance, A turns on again, and B turns off again. Immediately afterwards, the capacitor corresponding to B discharges to ground, and the capacitor corresponding to A charges from C0. Finally, waiting for the capacitor voltage to balance again.

[0311] The operation steps for pre-connecting and post-connecting capacitors may differ slightly in sequence, but their positive and negative voltage difference ΔV = a*k*(n1-n2) / (1+n1*k / 2) / (1+n2*k / 2). k=Cs / Cinput, where a is the total target capacitor pre-charge voltage, n1 is the number of switched capacitors in the positive circuit, and n2 is the number of switched capacitors in the negative circuit in one round of input operation. Both n1 and n2 are greater than or equal to 0. Compared to a scheme using a mixture of pre-connecting and post-connecting capacitors, a scheme using only post-connecting capacitors or only pre-connecting capacitors has a maximum voltage difference error that is 1-2 orders of magnitude smaller.

[0312] If we group the links, connecting multiple neurons—for example, two links per group (with a maximum of six capacitors in a single round with two weights, we can divide them into three pre-linked capacitors and three post-linked capacitors)—a control circuit manages this grouping to ensure that the number of pre-linked and post-linked capacitors is roughly equal. This results in a more even distribution of the two types of capacitors and a smaller error. In other words, if the weights and input values ​​are uneven, it might lead to unequal numbers of pre-linked and post-linked capacitors. We rebalance this through the control circuit. The previously fixed-line switched capacitors become dynamically switched capacitors. As long as the number of switched capacitors is the same in each round, the effect is the same. Grouping the links to achieve a comparable number of capacitors means the control circuit is modular, making the design more convenient.

[0313] As shown in Figure 7: If space allows, it would be better if each link had a multiple of switched capacitors. For example, in the case of a single-round operation with a weight of 2 bits, which would normally have 3 capacitors, it could be set to a multiple, i.e., 6 capacitors: 3 pre-linked capacitors and 3 post-linked capacitors. Thus, each unit of weight corresponds to two capacitors: one pre-linked and one post-linked. The operation that was originally a single switched capacitor becomes that the weight of the first link of a neuron corresponds to two sets of capacitors, a and b. A is a pre-linked capacitor, and the switch 935 connected to the main capacitor is on before the operation; b is a post-linked capacitor, and the switch 936 connected to the main capacitor is off before the operation. If the weight value is 1, then the corresponding capacitors are C11a and C11b; if the weight value is 2, then the corresponding capacitors are C11a and C11b, and C12a and C12b. Similarly, the connections 937 between the main capacitor and other circuits, as well as other necessary circuits, are not repeated here.

[0314] Similarly, the switching capacitors of the detection units for both positive and negative systems should also adopt a pre-linked and post-linked design. The detection unit capacitor Ccheck = Ct + Sum[Ct1a / Ct2a / Ct3a...Ctma]; Ct1b / Ct2b / Ct3b...Ctmb are post-linked switching capacitors, and 'a' is a pre-linked switching capacitor. Ccheck can be designed as part of Cinput, similar to the previous embodiment with a multi-layered dynamic network having an analog intermediate layer. Multiple sets of Cchecks can be connected or disconnected from Cinput via switches. These multiple sets of Cchecks and Cinputs form a pipeline-like, cyclical combination operation. That is, Cchecks, as part of Cinput, are pre-charged together and then receive charging and discharging from the neuron connections together. They are separated only during detection, and idle Cchecks are re-integrated into the Cinput to be completed.

[0315] In summary, the positive and negative circuits operate independently: 1. One end of each capacitor in the input section and the detection section is connected, and the other end is connected to a common terminal, i.e., connected in parallel as Cinput; 2. Cinput is charged to voltage a; 3. Each link is selected according to the weighted segment value, i.e., 0-3 switching capacitors in this embodiment; 4. According to the segment value of the input value, i.e., 0-1 in this embodiment, multiple rounds of discharge operations are performed on the corresponding links, i.e., each round of operation makes the selected pre-link and post-link switching capacitors operate synchronously, so that Cinput keeps its capacitance value as constant as possible during operation; 5. Prepare for detection, Cinput and Ccheck are disconnected; 6. Ccheck discharges at a specific discharge rate according to the voltage, i.e., when the voltage is high, multiple detection switching capacitors discharge simultaneously; 7. Ccheck stops when it reaches voltage b; 8. The digital circuit combines the discharge operations of Ccheck to obtain the equivalent number of discharges; 9. The number of discharges in the positive and negative circuits is subtracted to obtain the detection equivalent discharge time difference, i.e., the difference in the number of detection discharges; 10. Activate the circuit and output.

[0316] In general, the goal is to ensure that the parallel operation of the switched capacitors does not deviate from the true resistance, keeping the error within the required accuracy range of the circuit. The results of each round of operations are then combined through subsequent digital circuit operations to obtain a higher-bit / precision overall calculation result.

[0317] The above calculation error has been reduced to a practical range. If further improvement is needed, additional pre-linked or post-linked switched capacitors can be added to the above parallel operation scheme. In this case, the total number of pre-linked and post-linked capacitors will not be the same. However, the number of switched capacitors that can eliminate the error and the total number of positive and negative systems (n+ / n-) are mathematically related, but this relationship does not depend on the input of neurons within the group, but only on the weights of neurons within the group. That is, in addition to calculating the number of these additional switched capacitors in real time, we can also pre-calculate them and load the relevant settings parameters when loading the weights. These parameters are used to further correct the type and number of additional switched capacitors for error correction. Alternatively, without changing the number of pre-linked and post-linked capacitors within the group, we can simply adjust the ratio of the capacitance values ​​of the pre-linked and post-linked capacitors to reduce and equalize the error of the simultaneously parallel-operated switched capacitors within a certain range.

[0318] *An Example of a Wafer-Level Heterogeneous Bonding Computing System for Neural Networks Based on Hybrid Information Communication* This example provides an innovative path for building small-volume, high-performance, and energy-efficient ultra-large-scale AI computing systems. The vertical communication between wafer-level dies / computing cores (avoiding bad pixels) and the horizontal bilinear / ring-type parallel communication within the wafer (two independent communications) in the previous examples are based on digital or analog equivalent time information or capacitance information. Building upon this, wafers with different process nodes and advantages are vertically integrated; computing wafers based on traditional complementary metal-oxide-semiconductor (CMOS) technology are three-dimensionally heterogeneously integrated with analog / memory wafers based on thin-film transistor (TFT) technology, and then used in conjunction with optical waveguide roll-up films, thereby achieving an optimal balance between performance and cost at the system level. The global clock, power network, pulse width or voltage reference source, and high-speed shared communication bus (such as a parallel bus based on through-silicon vias, TSVs) required for auxiliary neural network computing and general-purpose computing cores are well-known and will be omitted here.

[0319] In this system, CMOS wafers and TFT wafers play complementary roles. Traditional CMOS processes excel at manufacturing high-density, high-frequency logic transistors and multi-layered, high-precision metal circuitry, making them ideal for high-speed vector computation and complex control. However, integrating high-resistance resistors, large-capacity capacitors, low-leakage transistors, and analog components such as optical transmitters and receivers, even if feasible, often requires additional mask layers or the introduction of special materials, leading to increased process complexity and cost.

[0320] In contrast, while TFT technology lags behind CMOS in transistor density and switching speed, its ability to fabricate large-area, highly uniform thin-film devices on glass or flexible substrates allows for the convenient large-scale integration of high-precision resistors, capacitors, and even non-volatile memory elements and optoelectronic devices at a lower cost and with simpler processes. This heterogeneous integration approach enables high-speed digital computing, high-density analog / storage, and high-performance communication to be realized on their respective most suitable process platforms and tightly interconnected.

[0321] At the ends of the stack, other functional chips can be further integrated using heterogeneous bonding technology, such as heterogeneous wafers specifically for sensing, or thin-film circuits integrating optoelectronic devices. Crucially, flexible optical waveguides or flexible circuit boards (circuit boards or various chips soldered together) can be bonded and integrated to achieve high-speed, low-power optical interconnects between layers, modules, and even with external systems.

[0322] *Variant Capacitor Time Calculation Circuit Example* As shown in Figure 30, this example uses a variable capacitance circuit. Its charge adjustment module can change the number of small capacitors connected to the differential capacitor module according to the neuron input value and weight value, thereby changing the capacitance value of the main capacitor of the differential capacitor module before detection. In this example, the capacitor module 950 and the charge adjustment module 962 are drawn together. The small capacitors and switches of the capacitor network belong to the charge adjustment module, which is divided into positive 962 and negative 961. The input 963 of the neuron unit circuit in this example is a 1-bit digital signal. Each round is a 1-bit input of the corresponding bit of each input value of the neuron, which controls the conduction switches 951 of the (positive and negative) differential capacitor modules 962 and 961 and the main capacitor 950; while the weight 960 is a 2-bit numerical signal, which controls the weight setting switch 952 corresponding to the three product terms input capacitor 953. Through multiple rounds of operations, the multiplication of each segment in response to factorization is completed, and finally the results of the multiplication of all segments are merged to complete the multiplication-accumulation operation.

[0323] During the weight loading phase, the weight setting switch sets the input capacitor 953 of the product term to be incorporated into the main capacitor according to the corresponding bit of the weight value. The operation of the neuron circuit input information causes the corresponding product term input capacitor to be incorporated into the main capacitor via the corresponding conduction switch 951, meaning the capacitance value of the main capacitor changes; this is called the variable capacitance input phase. The main capacitor C0 prevents the capacitance from being too small, balances the change in total capacitance value caused by the unequal number of product term input capacitors in each round of operation, and ensures that the discharge detection time is within a detectable preset range, while also reducing circuit noise. Assume the number of product term input capacitors whose connection is changed is n, and the capacitance value of the product term input capacitor is Cs. There are three variable capacitance forms: PSum[Cs] represents selecting and summing the product terms with positive signs (weights and inputs), i.e., the number of positive product term input capacitors * Cs; NSum[Cs] represents selecting and summing the product terms with negative signs (weights and inputs), i.e., the number of negative product term input capacitors * Cs. n = |PSum[Cs]| + |NSum[Cs]|.

[0324] Capacitance increase form: The change in capacitance connected to the main capacitor is increased, that is, the product term input capacitor, which was originally disconnected from the main capacitor, changes the capacitance value of the main capacitor from C0 to C0positive+|PSum[Cs]| and C0negative+|NSum[Cs]| respectively.

[0325] Capacitance reduction form: The change in capacitance connected to the main capacitor is reduced. That is to say, the input capacitor of the product term was originally connected / pre-connected to the main capacitor. The capacitance values ​​of the main capacitor in the positive and negative systems change from C0 to C0positive-|PSum[Cs]| and C0negative-|NSum[Cs]|, respectively.

[0326] Mixed Capacitance Configuration: The positive and negative product term input capacitors are not separated into positive and negative systems, but are all in the mixed system / main system. The negative product term input capacitor is originally connected to the main capacitor, while the positive product term input capacitor is originally disconnected from the main capacitor. The reference system / auxiliary system has no product term input capacitors, only the main capacitor C0. The changes in the positive and negative product term input capacitors connected to the main capacitor can be both increases and decreases. If the product term is negative, the corresponding switch is closed from the pre-connected negative product term input capacitors that were originally connected to the main capacitor, separating them from the main capacitor. If the product term is positive, the corresponding switch is opened from the positive product term input capacitors that were originally disconnected from the main capacitor, connecting them to the main capacitor. The main system's main capacitor value changes from C0 to C0-|PSum[n*Cs]|+|NSum[n*Cs]|, while the auxiliary system's main capacitor value remains unchanged at C0. Clearly, the number of adjustable product term input capacitors in the main system is twice that of the positive and negative systems. (A constant-capacity calculation circuit can also be divided into a main system and an auxiliary system. The main system has twice the number of charging and discharging current channels as either the positive or negative system. The corresponding negative product term current channel is subtracted from the default charging and discharging channels based on the weight of the negative sign and the input bit segment value, while the corresponding positive product term current channel is increased in the number of charging and discharging channels based on the weight of the positive sign and the input bit segment value.) Of course, because the number of switches and quantity channels differs between the main and auxiliary differential systems, the error is greater.

[0327] As mentioned in the previous examples, tp1-tn1 = R*(cap1-cap2)*ln(a / b)-R*(X1*pulseR1 / R1+X2*pulseR2 / R2+...+Xn*pulseRn / Rn)*t. The difference in time difference caused by the capacitance difference is calculated as: err_cap=R*(cap1-cap2)*ln(a / b). That is, cap1-cap2 is also proportional to tp1-tn1. When t=0, tp1-tn1 = err_cap. R is the equivalent resistance of the detection circuit, cap1-cap2 is the capacitance difference of the differential capacitors, and tp1-tn1 is the charging and discharging time difference of the differential capacitors.

[0328] After the weighting and variable capacitance input phases, the pre-charging phase begins. The main capacitor of the main / auxiliary or positive / negative system, along with the input capacitor of the connected product term, is charged to the reference voltage Va through the pre-charging circuit 956 of the capacitor control module, and then discharged to the reference voltage Vb through the sensing resistor network 957. That is, the current main capacitor voltage is compared with the voltage of the reference voltage module 955 by one or more voltage comparators 954, and the operation of the digital circuit is triggered based on the comparison result. Assuming the sensing resistor value is R, the calculation formula can be obtained:

[0329] Capacity expansion method: The discharge time difference between the positive and negative systems is err_cap = (|PSum[Cs]|-|NSum[Cs]|)*R*ln[Va / Vb].

[0330] Capacity reduction form: The discharge time difference between the positive and negative systems is err_cap =-(|PSum[Cs]|-|NSum[Cs]|)*R*ln[Va / Vb].

[0331] Mixed capacity configuration: The discharge time difference between the main and auxiliary systems is err_cap = (|PSum[Cs]|-|NSum[Cs]|)*R*ln[Va / Vb].

[0332] By adjusting the sign and the value of the R*ln[a / b] part, we can make the final result equal to |PSum[Cs]|-|NSum[Cs]|, which is equal to the sum of the specific segments of each input value and the specific segments of each weight value, or in a specific proportion to it.

[0333] If necessary, we can convert the discharge time difference into a digital signal. Conventional conversion methods are based on digital circuits using the system clock or the reference time signal of the inverter delay chain. In this invention, the fixed-capacitance embodiment mentioned earlier also achieves digitization through multiple voltage comparators and multiple sets of values ​​to accelerate discharge. However, this embodiment is a variable-capacitance type, where the number of input capacitors in the product term is unpredictable (without complex circuitry). Although the size of the capacitor or the remaining detection discharge time of the positive and negative systems can be inferred from the voltage change rate per unit time and the time length of a specific voltage difference, the circuitry is relatively complex.

[0334] In this embodiment, accelerated operation and digitization can be achieved relatively simply by detecting voltages Vb, Vc, Vd... in multiple stages and attempting accelerated discharge. In this embodiment, the digitization of the positive and negative / master / auxiliary systems is independent. Finally, the digital circuit of the neuron output module is used to perform addition and subtraction of the digital values ​​of the positive and negative / master / auxiliary systems to obtain the cumulative sum of the product terms for the current bit segment. Accelerated discharge is detected by adjusting the resistance value of the detection resistor network 957 (ideally in a binary multiple ratio). Initially, a smaller detection resistor value and a larger detection current are used to accelerate discharge. Assuming a resistance of 1 / 8R, the discharge is performed for m1 standard units of time (using RC discharge, inverter delay chain, or crystal clock as time references). If the voltage of a single system (positive / negative / main / auxiliary) is lower than the first-stage detection voltage Vb, the detection resistor value is doubled to 1 / 4R, and the operation continues for m2 standard units of time. Similarly, if the voltage is lower than Vc, the detection resistor value is doubled to 1 / 2R, and the operation continues for m3 standard units of time. Likewise, if the voltage is lower than Vd, the detection resistor value is doubled to 1R, and the operation continues for m4 standard units of time. This cycle continues until the detected capacitor voltage is lower than the stage target voltage and the resistance is at its maximum. The digital control logic 958 can directly obtain digitized discharge time information by accumulating the products of time lengths m1, m2, m3, m4... and their operating conductances (the reciprocals of the resistance values).

[0335] The detection resistor can also be replaced by the equivalent resistance of a switched capacitor network, in which case the output is the number of discharge operations of the switched capacitor network. As mentioned in the previous embodiment, the parallel connection of switched capacitors is a discrete operation that cannot be infinitely differentiated, resulting in errors compared to a real parallel resistor connection. In the circuit of the variable capacitance charge adjustment module, the capacitance value of the main detection capacitor is not fixed, and the equivalent resistance value of the switched capacitor varies and has errors when the switched capacitor is not small enough. However, the single-symbol calculation information (equivalent time pulse or equivalent time difference pulse) output by the variable capacitance input capacitor time calculation circuit can be input into the fixed capacitance capacitor time calculation circuit. Using a switched capacitor network in the fixed capacitance capacitor time calculation circuit, since the capacitance values ​​of both the main detection capacitor and the switched capacitor network are stable and preset, a precise equivalent resistance value can be obtained. That is, a precise equivalent switched capacitor capacitance value can be generated, achieving an effect completely consistent with the resistor network. The operation of detecting different resistance values ​​per unit time of accelerated discharge, mentioned earlier, is achieved through a single operation of a switched capacitor with different capacitance values. The same accelerated discharge attempt begins by using a detection switching capacitor with a larger capacitance to accelerate the discharge (positive or negative, primary or secondary). If the system voltage is lower than the target voltage for the current stage, then a detection switching capacitor with double the equivalent resistance is used. This process is repeated through multiple stages to achieve accelerated discharge, digitization, and the output of differential multiplication / addition information / equivalent time difference information 959. If the voltage of the detection capacitors in both the positive and negative or primary / secondary systems is lower than the target voltage in the above stages, and the equivalent resistance of the detection switching capacitor has reached its maximum, then the detection ends. Note: When the capacitance of the switching capacitor is very small, the charge transferred each time will be affected by the thermal noise of the switching resistor. The charge carried away by the switching capacitor in a single operation will still be affected by noise, requiring a noise reduction design.

[0336] *An Example of Equivalent Time Pulse Form Conversion* As shown in Figure 10, the circuit is similar to other embodiments. The equivalent time pulse or equivalent time difference information is used as the input 980 of the conversion circuit. The charge adjustment module adjusts the charge of the (voltage-initialized) capacitor module 981. During the discharge detection process, the capacitor control module's control logic 982 implements a new equivalent time pulse 983. Based on the equivalent time pulses output by two sets of such differential circuits, the time difference between the two is calculated, thus realizing the conversion between the equivalent time input and output.

[0337] *Varicapacitive Detection Example* Principle: A capacitor C0 and C1 are connected in parallel and pre-charged to voltage A. Then, a zero-voltage capacitor C2 is connected in parallel, causing voltage A to drop to the target voltage B. According to the law of charge conservation, (C0+C1)*A=(C0+C1+C2)*B, therefore C2=(C0+C1)*(AB) / B. In the previous example of the variable-capacitive input capacitor time calculation circuit, the difference in capacitance value changed by the positive and negative charge adjustment modules was proportional to the difference in discharge time of the equivalent resistance of the detection circuit. In this example, the difference in capacitance value changed by the positive and negative charge adjustment modules is proportional to the difference in capacitance value added by the detection circuit to reduce the detection capacitor voltage to the target voltage, i.e., a differential system, Cpositive2 - Cnegative2 = (Cpositive1 - Cnegative1)*(AB) / B. In summary, the calculation result, like in the previous example, is proportional to the cumulative value of the neural network input and weights. Circuits for various types of charge adjustment modules and capacitor control modules can be used interchangeably and converted to each other, and their proportional coefficients can be set through the circuit.

[0338] As shown in Figure 34, the variable capacitance detection embodiment is similar to the previous embodiments. Its differential system is divided into positive and negative modules, and its charge adjustment module is divided into a positive module 990 and a negative module 991. The main capacitor C0 of the capacitor module is placed within the charge adjustment module box. Similar to the previous embodiments, the variable capacitance charge adjustment module segments its input and weight terms bit-wise. The input term has one bit segment, and the weight term has two bit segments (i.e., four values). The product of the input and weight is also two bits with four values. Based on the sign of the product, each product term is controlled and incorporated into the corresponding capacitor of the positive and negative capacitor modules, thus changing the capacitance value of the main capacitor of the positive and negative system capacitor modules. After the main capacitor value is adjusted, it is pre-charged to the initial voltage, assumed to be A. The difference between this embodiment and the variable capacitance input capacitor time calculation circuit is that the detection circuit within its capacitor control module also uses a variable capacitance circuit. In a single differential system, there is a positive capacitance value detection circuit 992 and a negative capacitance value detection circuit 996. The positive converter circuit for capacitance detection contains multiple capacitors 993 with switches. Their initial voltage is 0, and the capacitors are initially grounded via a grounding switch 994. After detection begins, the grounding switch is closed, and the capacitor module is connected to the main capacitor switch 995, which is then opened. The positive converter capacitors are then connected to the main detection capacitor one by one. Assuming the capacitance of the positive converter capacitor is Cs, as these small capacitors are connected to the main capacitor, the main capacitor voltage gradually decreases. When the voltage comparator determines that the main capacitor voltage is less than or equal to the target reference voltage B, then according to charge conservation, C2 = (C0 + C1) * (AB) / B. Because these small capacitors are controlled by the digital logic circuit 997, the digital logic circuit can determine the number of positive converter capacitors connected to the main capacitor through the operation record at this time. Then, the number of small capacitors in the differential positive and negative systems is subtracted, resulting in (Cpositive1 - Cnegative1) * (AB) / B, which is the product of the sum of the input and weighted corresponding bit segments and (AB) / B. However, a large number of such small capacitors are needed, resulting in a long operation time. For example, adding 64 2-bit 4-value product terms requires a resolution of at least 256. If the capacitance value of the main capacitor C0 is equal to the maximum value of the adjustable capacitance C1, the resolution needs to be 512, requiring 512 switching operations, making the operation cumbersome. (In the diagram, the small capacitors used for expanding the main capacitor's capacity in the detection process are connected in parallel to the main capacitor, and their switches are directly connected to the small capacitors and the main capacitor. However, connecting the main capacitor to the small capacitors via a tree-like, series, or other form of switching circuit can achieve a similar effect.)

[0339] The problem of lengthy operation can be solved by using a capacitor network with unequal values ​​matching the binary representation. For example, a 512 resolution requires 9 bits, so the capacitor network for detecting discharge can be set to 9, each with a capacitance value twice that of the next. Considering error and adjustment, we can set the capacitance at corresponding bits, with more small-value capacitors. In other words, using a binary-matched capacitor network with unequal capacitance values ​​accelerates the detection process and also achieves digitization. Compared to the previous form of switching capacitor equivalent resistance, where the capacitor network performs multiple identical discharge operations, this embodiment, while also involving multiple operations of the capacitor network, operates on different small capacitors each time, thus increasing the capacitance value of the main capacitor module each time. The advantages are: no error from switching capacitor equivalent resistance; and the switching of small capacitors in the detection capacitor network of this embodiment occurs between the main capacitor and the small capacitors, without repeated grounding, making it less susceptible to switching thermal noise and external charge interference.

[0340] For example, at the start of detection, we first connect the small, positively changing capacitor (with a larger capacitance value) in the capacitor network. If the voltage is already lower than the target voltage B, we can connect the small, inverting capacitor, i.e., the small capacitor in the capacitance inverting circuit 996. The capacitance inverting circuit has a switch 998 connected to the pre-charge voltage source. The initial voltage of the small capacitor 999 inside is the same as that of the main capacitor, both being A. Connecting the inverting capacitor means that the total charge of the system increases, and the voltage of the main capacitor is also raised. Note the difference between the parallel capacitor in the inverting circuit and the positive circuit. The positive circuit operates independently in the differential system, while the inverting circuit operates simultaneously in the differential system. That is, in both the positive and negative differential systems, the main capacitors of both have the same capacitance value (Cp) connected in parallel with the inverting capacitor, and the same amount of charge is added. In other words, because the charge and capacitance expansion brought to the main capacitor by the inverting circuit in the differential system are equal, the effect of the inverting circuit will be canceled out in the neuron output circuit, i.e., when the differential system finally calculates the difference. If the voltage is exactly B at this point, then the detection result C2 = (C0 + C1 + Cp) * (AB) / B. In this embodiment, the voltage comparator only determines whether the voltage of the main capacitor is less than B. Through the expansion operation of the inverter circuit, the voltage of the main capacitor is already greater than B. If not, the inverter circuit will continue to connect more small inverter capacitors with a voltage of A. Because the previous operations corresponded to capacitors with higher capacitance values, we will next operate on capacitors in the forward inverter system corresponding to capacitors with lower capacitance values. That is, we gradually operate the small forward inverter capacitors with corresponding capacitance values ​​from high to low. If the voltage of the main capacitor in the capacitor module is lower than the stage voltage target B, we simultaneously raise the voltages of both differential systems until the voltages of the main capacitors in both systems are greater than B, and then replace the small forward inverter capacitor with the lower capacitance value. If the voltage of the main capacitor does not fall below B due to the operation, then we continue the operation of the current position, that is, continue using the small capacitor with the capacitance value from the previous operation. This process continues until the operation reaches the minimum bit and the voltage is less than B. Only then does the operation of the differential system end, and its differential output 1000 is handed over to the neuron output module to calculate the difference or activation value between the two.

[0341] Additionally, we need to reduce the number of operations in the inverter circuit, as this consumes energy and time. Therefore, we can use multiple comparators. By comparing multiple comparators with multiple reference voltages, we can roughly determine the voltage difference between the current detection main capacitor and the target voltage. Based on this approximate voltage difference, we can predict the allowable value of the currently connected detection capacitor, avoiding operations where the voltage is less than B, and reducing the number of such operations. In other words, by simulating the hardware circuit, we can predict whether the current operation will be lower than B. If the prediction is yes, we directly reduce the number of bits in the operation and the corresponding capacitance value.

[0342] Specific steps:

[0343] 1. The differential system charge adjustment module selects the main capacitor to be connected to the positive or negative capacitor module based on the values ​​of the neuron input and weights, as well as the sign of their product, and adjusts the capacitance value of the main capacitor of the capacitor module.

[0344] 2. The main capacitor voltage of the capacitor module is initialized to A, the small capacitor voltage of the forward converter circuit of the capacitor control module is initialized to 0, and the small capacitor of the reverse converter circuit of the capacitor control module is initialized to A.

[0345] 3. The forward converter circuit of the capacitor control module turns off the grounding switch, the reverse converter circuit of the capacitor control module turns off the switch of voltage source A, and the main capacitor of the capacitor module turns off the switch of voltage source A.

[0346] 4. If the operation result can be predicted, then if there is a prediction that the main capacitor voltage is less than B, the detection circuit replaces the small capacitor with the corresponding capacitance value of the lower bit, and this step is repeated; the small capacitor with the current corresponding capacitance value of the positive converter is added in parallel; the voltage comparator determines that if the main capacitor voltage is greater than the target reference voltage B, this step is repeated.

[0347] 5. If the capacitance operation does not reach the minimum bit, execute the inverting circuit operation. The main capacitors of the positive and negative differential systems are simultaneously connected to the inverting small capacitor. Repeat this inverting operation until both voltages are raised to greater than B. Detect the positive conversion circuit and replace the small capacitor with the corresponding capacitance value of the lower bit, and jump back to step 4.

[0348] 6. Differential systems output digital codes corresponding to the capacitance values, or information on the changing capacitance values.

[0349] 7. The neuron calculates and outputs the difference (C+1 - C-1)*(AB) / B based on the output of the difference system, or calculates the activation value later.

[0350] The tolerance values ​​and operating steps in this embodiment are for reference only. Specific applications need to be set according to actual conditions.

[0351] *An Example of Heterogeneous Input / Output Conversion* The form of the equivalent time pulse used for detection can be different from the form of the input. For example, the output may be the number of times the capacitor is switched, while the input is a single pulse. Or the output may be a single pulse, while the input may be the number of times the capacitor is switched. Alternatively, one of the inputs or outputs may be PWM pulse width modulation; or it may be digital information representing the equivalent time pulse; or it may be the capacitance value information incorporated into the detection of a varactor network. These can also have linear or nonlinear relationships. This achieves the conversion of different forms of input and output. In addition to using the conversion form to coordinate the calculations of different computing cores, the most important purpose of this is to improve performance. For example, the single-pulse form has lower energy consumption and higher performance than multiple operations of the capacitor, but the circuit for converting it to a digital output is relatively complex. Therefore, in a multi-layer analog neural network within a single chip, the middle layer can use the single-pulse form, while the first layer input or the last layer output can use the capacitor switching operation or digital form.

[0352] *In summary*, the above embodiments have sufficiently illustrated the implementation of the present invention. The present invention has various potential or hybrid implementations, and the specific numerical examples can be set according to actual needs, and are not limited to the content of the partial textual description. Industrial applicability

[0353] This invention relates to an artificial intelligence computing circuit, principle, and method, which can be widely used in printed circuits and integrated circuits. The resulting products exhibit excellent performance in various aspects and can meet the application requirements of relevant industries. Therefore, this invention possesses industrial applicability.

Claims

1. A parallel computing system for neural networks suitable for wafer-level stacking, wherein the basic neuron unit includes: Charge adjustment module, capacitor control module, capacitor module, neuron output module; The capacitor control module and the capacitor module are differentially configured, each with positive and negative settings, or main and auxiliary settings. The capacitor control module of the corresponding attribute has a circuit for charging and discharging the capacitor module of the corresponding attribute, and can independently charge and discharge the capacitor module of the corresponding symbol to a specific voltage to prepare for simulated neural network detection. The charge adjustment module changes the voltage or capacitance of the capacitor module according to the weight and input of the neuron, thereby changing the amount of charge before detection. The capacitor control module has a current channel for detecting charging and discharging, which can charge and discharge the corresponding capacitor module. The capacitor control module contains a voltage comparator. During the process of detecting changes in capacitor voltage, the voltage of the capacitor in the capacitor module is compared with the target reference voltage, or the voltage of the capacitor in the capacitor module is compared with the voltage of other capacitors with the same sign attribute, generating a voltage comparison signal, and terminating the detection based on this signal. Based on the duration or operation of the charging and discharging current channel from conduction to the generation of the termination detection voltage comparison signal, and the corresponding bit or ratio, each of the differential capacitor control modules outputs single-symbol calculation information. The single-symbol calculation information corresponds to the duty cycle in pulse width current modulation, the pulse width time in single-pulse pulse width modulation, the number of switching capacitor operations in switched capacitor current modulation, and the capacitance value of the capacitor added in the discharge operation of the varactor capacitor network. The neuron output module generates difference multiplication and addition information based on the single-symbol calculation information of the difference; The neuron output module generates a symbol level or its digitized signal, i.e., symbol information, based on the comparison of differential single-symbol computation information or the voltage comparison of differential capacitor modules. The output of the basic neuron unit, including the differential multiplication-addition information and the symbolic information, is transmitted or stored in its own input or in external circuits.

2. The computing circuit system according to claim 1, characterized in that: The charge adjustment module includes an input current modulation control module and a weighted current modulation control module, which can adjust the voltage of the capacitor module by modulating the charging and discharging current, thereby adjusting the amount of charge. There is a reference time source signal with one or more bits; The input current modulation control module generates an input current modulation control signal based on an external signal, or a signal from a neuron output module, or a reference time source signal and the value of the input fast access unit loaded thereon. The weighted current modulation control module, based on the value of its own weighted fast access unit, and based on the corresponding input current modulation control signal, further selects, trims, or inserts a delayed control signal to generate the modulation current, or adjusts the equivalent resistance of the current channel. Multiple current channels with corresponding neural network links can charge and discharge the capacitor modules with corresponding attributes according to the settings. The current channel has a switch for generating a modulated current, which is controlled by the control signal for the modulated current to generate a modulated current calculated by an analog neural network for charging and discharging the capacitor module. The signs of the input item fast access unit and the weight item fast access unit together control the selection of positive / negative or main / auxiliary capacitor modules for charging and discharging. After the modulation current calculated by the simulated neural network is used to charge and discharge the capacitor module, the equivalent resistance is detected to continue charging and discharging the capacitor module.

3. The computing circuit system according to claim 1, characterized in that: The charge adjustment module includes a circuit for adjusting the capacitance value of the main capacitor of the capacitor module, including a parallel switch and a parallel capacitor; the signs of the input item fast access unit and the weight item fast access unit jointly control the selection of positive / negative or main / auxiliary capacitor modules for capacitance adjustment; after the capacitance value of the main capacitor of the capacitor module is adjusted, its voltage is simultaneously adjusted to a specific voltage, and then the detection stage begins.

4. The computing circuit system according to claim 1, characterized in that: The capacitor module contains a combination of capacitors and switches; the capacitor module includes a main capacitor and a detection capacitor; in the constant capacity mode, that is, the charge adjustment module only adjusts the voltage of the corresponding capacitor module, wherein a certain main capacitor and a certain detection capacitor are charged and discharged together through a switch combination; the remaining disconnected or uncombined detection capacitors are charged and discharged independently or their voltages are compared.

5. The computing circuit system according to claim 1, characterized in that: There is a circuit with multiple layers for detecting charge and discharge. The first layer detects charge and discharge to generate differential multiplication and addition information to control the charge and discharge of the detection capacitor in the next layer. The simulation layer generates a linearly or non-linearly activated output by selecting the charging and discharging mode and charging / discharging level depth of the detection capacitor according to the settings; the output is subsequently used as input to its own layer or other layers.

6. The computing circuit system according to claim 5, characterized in that: There is a multi-layer neural network, with analog layers in the middle and digital layers at the beginning and end; the output of the analog layer uses equivalent single-pulse pulse width modulation; the equivalent single-pulse pulse width modulation means that the output of a neuron is a pulse signal of one or more channels, and the output of a single channel is one or more independent single-pulse signals.

7. The computing circuit system according to claim 5, characterized in that: There is a multi-layer neural network, and each layer of the neural network contains multiple basic neuron units. The differential multiplication and addition information output by the basic neuron units in the last layer of the multi-layer neural network is expressed in binary, i.e., digitized. The output data of the multi-layer neural network is transmitted in the form of the digitized differential multiplication and addition information and the symbolic information.

8. The computing circuit system according to claim 7, characterized in that: The dies within the wafer contain arrays of the basic neuron units described above, corresponding to the multilayer neural network described above; adjacent dies have serial or parallel communication lines for transmitting their input or output data; the input or output data of the multilayer neural network includes the digitized differential multiply-add information and the symbol information described above; the die array or division within the wafer is used to communicate independently and in parallel with adjacent dies within the region using a ring or bilinear configuration.

9. The computing circuit system according to claim 8, characterized in that: The wafer contains dies with bonding pads; the wafers are stacked together, and the stacks are connected by the bonding pads; the bonding pads include independent communication channels between the dies in the stack; the dies in the wafer have independent communication channels connecting adjacent dies; the independent communication channels between the dies and the independent communication channels between adjacent dies can be combined into a transfer channel, so that data between the dies in the stack and data between adjacent dies can be communicated.

10. The computing circuit system according to claim 9, characterized in that: The stacked wafer array contains damaged dies; the damaged dies have independent communication channels above and below and independent communication channels for the normal operation of adjacent dies; the working dies connected between the array coordinates corresponding to the damaged dies of two wafers in the stack pass through the transfer channel, so that data transmission avoids the damaged dies.

11. The computing circuit system according to claim 2, characterized in that: Current modulation is achieved through the operation of switched capacitors; the capacitor module has a main capacitor; the switched capacitors are divided into a pre-connection group and a post-connection group; the pre-connection group is connected to the main capacitor of the capacitor module before each current modulation of the switched capacitors occurs; the post-connection group is connected to the main capacitor of the capacitor module after each current modulation of the switched capacitors occurs; the capacitors of the pre-connection group and the post-connection group operate synchronously to correct the calculation deviation caused by the parallel discharge of the switched capacitors; while the capacitors of the pre-connection group are disconnected from the main capacitor and grounded or connected to other voltage sources, the capacitors of the post-connection group are disconnected from ground or other voltage sources and connected to the main capacitor, and the original connection state is synchronously restored after the voltage stabilizes.

12. The computing circuit system according to claim 11, characterized in that: The switched capacitors are divided into groups with equal pre-link and post-link numbers. The smallest unit of a group is a pre-link capacitor and a post-link capacitor. The switching capacitors in each group are operated in a staggered sequence.

13. The computing circuit system according to claim 1, characterized in that: The detection equivalent resistance is a switched capacitor network or a resistor network; the detection circuit has accelerated discharge during the charging and discharging process, and the size of the detection equivalent resistance is determined according to the voltage of the main capacitor to be detected and the target reference voltage of the detection stage, that is, the size of the modulation current of the charging and discharging. By using digital control logic, combining the operational results with the equivalent resistance value of the operation, detection acceleration and digitization are achieved.

14. A wafer-level stacked parallel computing system for neural networks, wherein the basic unit of a neuron has two differential capacitors, characterized in that... Computational methods with neurons: Step 1: Obtain the values ​​of the input items and the weight items; Based on the sign of the product of the input item's value and the weight item's value, if it is a capacitor voltage adjustment mode, open the current channel of the corresponding capacitor; or if it is a capacitor capacitance adjustment mode, open the variable capacitance channel of the corresponding capacitor. Step 2, in parallel, the two differential capacitors are connected to a specific voltage source and pre-charged and discharged to a specific voltage; Step 3, or if it is the mode for adjusting the capacitance value, turn on the adjustment switch of the two differential capacitors; If it is the mode of adjusting capacitor voltage, in the corresponding current channel of the neuron link, the value of the input item and the value of the weight item are converted into the charging and discharging modulation current of the capacitor, that is, the product of the equivalent conductance and the current conduction time generates the modulation current to charge and discharge the capacitor. Step 4: After the capacitor charge adjustment is completed, the detection process begins, and charging and discharging continues until the threshold voltage is reached, generating two differential single-symbol calculation information. The single-symbol calculation information corresponds to the duty cycle in pulse width current modulation, the pulse width time in single-pulse pulse width modulation, the number of switching capacitor operations in switched capacitor current modulation, and the capacitance value of the capacitor added in the discharge operation of the varactor capacitor network. Step 5: Based on the two single-symbol calculation information of the difference, generate the difference multiplication and addition information and symbol, or its digitized value, as the output of the basic neuron unit.

15. The method for calculating neurons in a computing circuit system according to claim 14, characterized in that, A method for generating an activation function output through a multi-layered charge / discharge detection circuit: Step 1: The first-layer charging and discharging circuit is set to a linear mode and generates a linear output of the differential multiplication and addition information; Step 2: The digital logic circuit uses the maximum value of the linear segment of the segmented activation function. If the value is exceeded, the remaining time of the first-level linear output is used as the input of the second-level charging and discharging circuit. Step 3: The second-layer charging and discharging circuit is set to a non-linear mode and generates a non-linear output of the differential multiplication and addition information. Step 4: The digital logic circuit combines the outputs of the differential multiplication and addition information from the first and second layers; Step 5: The combined differential multiplication and addition information is used as input to its own or external circuits.

16. The method for calculating neurons in a computing circuit system according to claim 14, characterized in that, When using switched capacitors to modulate current, there are methods to correct for the discrete error in the calculation of parallel switched capacitors: During parallel operation, the two capacitors operate synchronously, that is, they complete the operation in approximately the same time. Step 1: Before the parallel operation of the switched capacitors, some of the pre-connected capacitors are connected to the main capacitor and disconnected from the charging and discharging voltage source, while some of the later-connected capacitors are disconnected from the main capacitor and connected to the charging and discharging voltage source. Step 2: When the switched capacitors are connected in parallel, the pre-connected capacitor is disconnected from the main capacitor and connected to the charging and discharging voltage source, and the subsequent connected capacitor is connected to the main capacitor and disconnected from the charging and discharging voltage source. Step 3: After the voltage of the switched capacitor stabilizes, restore the original on and off states; Step 4: After the main capacitor voltage stabilizes, the parallel charging and discharging operation of the switched capacitors is completed.

17. The method for calculating neurons in a computing circuit system according to claim 14, characterized in that, The method of calculating by segmentation: Step 1: The input values ​​and weight values ​​of the neurons are each divided into segments with fewer bits according to binary digits; Step 2: Select the input segment value and weight segment value of the neuron according to the factorization law; Step 3: The selected input segment value and weight segment value are used to charge and discharge the capacitor module through current modulation; Step 4: After the capacitor is charged and discharged, the subsequent detection operation is performed, and the segmented differential multiplication and addition information is obtained. Step 5: Return to step 2 and repeat the operation until all terms of the factorization are calculated. Step 6: Obtain the final equivalent difference multiplication and addition information. The digital circuit completes the summation of each term in the factorization by shifting and adding.

18. The method for calculating neurons in a computing circuit system according to claim 14, characterized in that, There are methods to accelerate discharge detection, obtain the aforementioned differential multiplication-addition information, and digitize it: Step 1: Based on the comparison between the voltage of the detection capacitor and the reference voltage corresponding to the binary bit, select the largest bit; Step 2: Adjust the discharge path to discharge once or for a period of time with a current proportional to the actual value of the corresponding binary bit. Step 3: The fast storage unit corresponding to the differential multiplication and addition information in the digital circuit adjusts its value accordingly; Step 4: Return to step 1 and repeat until the detected capacitance reaches the threshold voltage.

19. The method for calculating neurons in a computing circuit system according to claim 14, characterized in that, In the constant-capacity mode, where the charge adjustment module only adjusts the voltage of the corresponding capacitor module, there are multiple main capacitors and multiple detection capacitors. There are methods to accelerate the operating frequency of the main capacitors. Step 1: Select the main capacitor and the detection capacitor for the idle state; Step 2: The main capacitor in the idle state and the detection capacitor in the idle state are turned on, that is, the capacitors are combined. Step 3: The merged capacitors are pre-charged together to a specified voltage; Step 4: The merged capacitors work together to charge and discharge the neuronal links. Step 5: In the merged capacitors, the original main capacitor and the original detection capacitor are disconnected, and the original main capacitor enters an idle state. Step 6: The original detection capacitor performs a detection discharge operation; Step 7: The original detection capacitor enters an idle state.

20. The method for calculating neurons in a computing circuit system according to claim 14, characterized in that, Methods for calculating full link segmentation: In a single layer of a neural network, the corresponding input and output neurons and their weights are divided into multiple groups. Step 1: Through parallel transmission in linear, circular, or cross-unit manner, each group synchronously exchanges input values; Step 2: In parallel, according to the program, load the corresponding weight values ​​for the respective inputs. Step 3: Each group independently calculates using the calculation method of the neurons in the aforementioned computing circuit system; Step 4: Use digital circuitry to add the calculation results of this group to the calculation results of other groups. Step 5, return to step 1, and repeat the operation; Step 6: Obtain the final calculation result.

21. The method for calculating neurons in a computing circuit system according to claim 14, characterized in that, One method for achieving linear output: The charge-discharge modulation current in step 3 is the discharge current, and the discharge drive voltage is the ground potential. The charging and discharging process in step 4 is called discharge, and the discharge driving voltage is ground potential. In step 5, the differential multiplication and addition information obtained is linearly proportional to the sum of the products of the input terms and weight terms of each link, thereby realizing the linear activation function of the neuron output.

22. The method for calculating neurons in a computational circuit system according to claim 14, characterized in that, A method for dynamically configuring the computing circuit system as one or more computing layers in a neural network to be computed: As a basic neuron unit in a layer of neural network, it loads the values ​​of input items into the input item fast access unit from memory, external circuits, or the output module of the previous layer of neurons, and loads the values ​​of weight items into the weight item fast access unit; completes the calculation of the computation layer; when the calculation is finished, it outputs the difference multiplication and addition information and its sign output by the basic neuron unit of the last layer to memory, as input to the subsequent computation layer, external circuits, or other processing modules.

23. The computing circuit system according to claim 1, characterized in that: The single-symbol computation information corresponding to the input and output of a layer of neurons is in a heterogeneous form, and the entire layer realizes the transformation of the single-symbol computation information form; the heterogeneous form is that one of the following is used as the input: the number of times the switched capacitor is operated, the duration of a single pulse, the pulse width modulation or the digitized value, while the output is selected in a different form than the input.

24. The computing circuit system according to claim 13, characterized in that: There is a reference neuron whose output is the pulse output of the capacitor control module, which is used to control the charging and discharging time of the equivalent detection resistor; the reference neuron can control the single-symbol calculation information output of the capacitor control module according to the parameter settings. The digital control logic controls the standard duration of accelerated charging and discharging during the detection process based on the real-time voltage range of the detection capacitor or the selected equivalent resistance value, and simultaneously selects the single-symbol computation information output of the corresponding reference neuron.

25. The computing circuit system according to claim 1, characterized in that: The computational circuit system is dynamically used as one or more layers of the entire neural network to be computed; the charge adjustment module loads the values ​​of the input fast access units from memory or external circuitry, or shared or loaded from the neuron output module; the charge adjustment module loads the values ​​of the corresponding weight fast access units from memory or external circuitry, or shared or loaded from the neuron output module; there is an equivalent resistance network that generates a specified equivalent resistance value based on the values ​​of the weight fast access units, or also based on the values ​​of the input fast access units, thereby adjusting the instantaneous current of the corresponding current channels; at the end of the computation, the value of the activation function fast access unit of the last layer is output to memory, its own input, external circuitry, or other modules.

26. The computing circuit system according to claim 25, characterized in that: An array-type memory is used as a cache; the array-type memory has a row-selective write circuit to write the output of a neural network layer to the selected row; the array-type memory has a transposed block-selective read circuit to write the data of a selected block in a column to the input or weight of a neural network layer.

27. The computing circuit system according to claim 25, characterized in that: The input item fast access unit is segmented bit by bit, and each segment independently generates one or more segment input item modulation signals; the weight item fast access unit is segmented bit by bit; each segment of the weight item fast access unit independently generates multiple modulation currents according to its own segment weight value and the corresponding segment input item modulation signal; the multiple modulation currents simultaneously charge and discharge the capacitor module of the corresponding segment.

28. The computing circuit system according to claim 25, characterized in that: A layered activation function numerical source generates numerical information with equivalent time as the independent variable based on the activation function setting. The numerical information with equivalent time as the independent variable is a time value, an activation function value, a derivative value, a calculation intermediate value, or a combination thereof. The neuron output module has an activation function fast access unit. The neuron output module selects the numerical information with equivalent time as the independent variable corresponding to the symbol based on the symbol level. When the differential multiply-add information signal starts or ends its shearing, the output value of the activation function of the corresponding neuron basic unit at the current time is obtained based on the numerical information with equivalent time as the independent variable of the corresponding symbol.

29. The computing circuit system according to claim 28, characterized in that: The layer activation function value source generates activation function value information according to the setting of the activation function and sends it to the neuron output module of the corresponding layer; the neuron output module generates new charge-discharge single-symbol calculation information based on the activation function value information; the neuron output module generates a new activation function value based on the new charge-discharge single-symbol calculation information and another activation function value information.

30. The computing circuit system according to claim 29, characterized in that: There is a network summarization layer; the network summarization layer has network summarization layer neurons that connect to the basic computational units of neurons in the current layer, and uses the activated values ​​of each neuron in the current layer as the input of the network summarization layer neurons; there is a layer computation module that receives the output of the network summarization layer neurons and generates a layer parameter value; the layer computation module, based on the layer parameter value, regenerates the activation function value information with equivalent time as the independent variable and sends it to the neuron output module of the corresponding layer.

31. The computing circuit system according to claim 27, characterized in that: It has multiple segment neuron output modules that can convert the single-symbol computation signal input from the differential capacitance control module of the corresponding segment into the value of the fast access unit; The aforementioned neuron aggregation calculation output module can perform shifting, addition, and subtraction operations on the values ​​of the fast access units of each segment neuron output module according to the corresponding bit range of its segment, and write the result value into its fast access unit.

32. The computing circuit system according to claim 1, characterized in that, During the detection process, the differential system contains a comparator for the differential detection capacitor voltage, and the result generates an enable signal to control the operating state of other circuits.

33. The computing circuit system according to claim 1, characterized in that, The capacitor control module has a comparator that uses a piecewise function to divide the reference voltage. The comparison result divides the discharge detection process, and then generates pulses corresponding to different activation functions, and finally outputs them.

34. The computing circuit system according to claim 1, characterized in that: Each neuron's basic unit has a bias link; the bias link input is set to full scale, and the charge adjustment module adjusts the value controlled by the weight term fast access unit.

35. The computing circuit system according to claim 1, characterized in that: The capacitor module's capacitor voltage is compared with a reference voltage to generate comparison detection signals P and N. The neuron output module logically synthesizes the comparison detection signals P and N, and generates an XOR gate signal based on P and N, indicating whether the comparison detection signals P and N are at the same level. The neuron output module generates a positive and negative signal SGN when the XOR gate signal starts, i.e., when P and N become unequal. If the comparison detection signal P lasts longer, SGN is positive; otherwise, it is negative.

36. The computing circuit system according to claim 2, characterized in that: There is one or more underlying reference pulse signals; the input current modulation control module has a pulse selection or shielding circuit, which selects or shields the underlying high-frequency reference pulse signal according to the input current modulation control signal to generate a short pulse input current modulation control signal, or selects one of the underlying high-frequency reference pulse signals.

37. The computing circuit system according to claim 25, characterized in that: There are multiple neurons, and their charge adjustment modules have a shared weight term fast access unit inside; each neuron's input term current modulation control module generates multiple sets of modulation signals according to the shared weight term fast access unit and its own input value; the charge adjustment module of the neuron's link adjusts the charge of the capacitor module according to the multiple sets of modulation signals; multiple neuron basic units work simultaneously or sequentially.

38. The computing circuit system according to claim 37, characterized in that: There are multiple sets of phase-staggered, non-overlapping underlying reference pulse signals; each neuron's input current modulation control module generates multiple sets of current modulation signals based on the phase-staggered underlying reference pulse signals and its own input values; the switching of the neuron's link current channel generates multiple sets of modulated currents based on the multiple sets of current modulation signals.

39. The computing circuit system according to claim 25, characterized in that, The charge adjustment module has a backup fast access unit; the output value of the neuron output module is preloaded into the backup fast access unit in blocks or batches; the charge adjustment module has a switching circuit that can switch the effective weight value to the value of the backup fast access unit, or switch the effective weight value to the weight value originally loaded in that layer of the neural network. The pre-generated output value of the neuron output module can be directly used as a real-time input value through the backup fast access unit.

40. The computing circuit system according to claim 26, characterized in that, There is a method for mini-batch training: in mini-batch training, data is stored each time, propagated in the same direction, and the stored data of the same node is written to adjacent rows in sequence; Step 1, during forward propagation, the array-type memory caches the output of the activated layer in each computation of the neural network, row by row; Step 2, in backpropagation, the array-type memory caches the product of the partial derivative of the activated output node of each layer of the neural network in each calculation and the derivative of the activation function, that is, the multiplicative partial derivative before activation. Step 3, the block-based reading circuit reads the multiplicative partial derivative values ​​of each backpropagation node in the mini-batch corresponding to the block into the layer input, and reads the multiple output values ​​of each neuron output node of the forward propagation of the previous layer into the multiple link weights of one neuron corresponding to the current calculation circuit. Step 4: Calculate the cumulative value and multiply it by the learning rate. The product is the adjustment value for the current corresponding link weight. Step 5: Update the corresponding link weights.

41. The computing circuit system according to claim 37, characterized in that, Neurons share a common computational core and a common memory; the layer activation function source writes activation function information into the common memory; the computational core obtains activation function values ​​from the common memory through table lookup, copying, and computation; there is a method for synthesizing multiple rounds of output values: Step 1: Perform forward propagation to obtain the level signal of the differential multiplication and addition information expressing the charging and discharging of the capacitor and the level signal of the sign expressing the energy relationship of each capacitor. Step 2: Update the parameters of the layer activation function numerical source, and continuously output the time value, the time derivative of the activation function, and the current value to the neuron output module, or output only the time value and obtain other values ​​by looking up a table; Step 3: If an activation function is needed, the activation function value of each neuron is obtained by interpolating the derivative change value and the current value. Step 4: If it is necessary to add the layer output values, do not reset the fast access unit of the output item or the fast access unit of the activation function, and add the value of the current activation function to the value of the previous activation function. Repeat steps 1, 2, 3, and 4. Step 5: Obtain the final neuron output value.

42. The computing circuit system according to claim 30, characterized in that, There are methods to normalize the output values ​​of the implementation layer: Step 1: The network summarizing layer neurons use the values ​​of the fast access units of the activation functions of the current main computing layer as input to the network summarizing layer neurons. Step 2: The neurons in the network summarization layer calculate the parameter values ​​of the generation layer, which is the process of powering on, detecting, and activating. Step 3: The layer parameter values ​​are given to the layer activation function value calculation chip. The layer activation function value source calculates and regenerates the normalized activation function value information with equivalent time as the independent variable based on the layer parameter values. Step 4: The activation function numerical information with equivalent time as the independent variable is sent to the neuron output module of the current main computation layer; Step 5: The current main computing layer regenerates the normalized activation function value based on the original activation function value and stores it in the activation function fast access unit.

43. The computing circuit system according to claim 39, characterized in that, The weighted current modulation control module has a backup fast access unit; based on a fully connected neural network, there is a method to generate the self-attention QK matrix and perform multiplication. Step 1: Read the Q matrix parameters from external storage or memory into the fast access unit of the input current modulation control module; Step 2, in parallel, read the value of the input vector into the fast access unit of the weighted term current modulation control module; Step 3: Perform forward propagation calculations, and the neuron output module generates a round of Q-matrix result values; Step 4: Copy the result value from the neuron output module to the spare fast access unit of the corresponding weight term current modulation control module. Step 5: Repeat steps 2, 3, and 4 until all spare fast access units of the current modulation control module for each weighted item are filled or all Q matrix parameters are used. Step 6: Read the K matrix parameters from an external storage unit or a backup fast access unit to the fast access unit of the input current modulation control module. Step 7, in parallel, read the value of the input vector into the fast access unit of the weighted term current modulation control module; Step 8: Perform forward propagation calculation; the neuron output module generates a round of K matrix result values. Step 9: The neuron output module transfers the K matrix result value to the fast access unit of the input current modulation control module. Step 10: The main fast access unit and the spare fast access unit of each weighted current modulation control module are exchanged, that is, the result value of the Q matrix is ​​used. Step 11: Perform forward propagation calculation. The neuron output module generates a round of QK matrix multiplication result value and saves the round of QK matrix result value to external storage or memory, or other idle and readily accessible units. Step 12: Repeat steps 6, 7, 8, 9, 10, and 11 until all K matrices have been calculated. Then, based on the calculation results from the previous steps, there are methods to generate self-attention QKV results: Step 13: The QK matrix multiplication result value is normalized and activated, and then stored in other idle standby fast access units of the weight term current modulation control module. This process is repeated multiple times until all QK activation result values ​​are obtained. Step 14: Read the V matrix parameters from an external or storage or backup fast access unit to the fast access unit of the weighted term current modulation control module. Step 15: In parallel, read the result value of the QK matrix multiplication in one round, i.e., the attention weight in one round, and send it to the fast access unit of the input current modulation control module. Step 16: Perform forward propagation calculation and activation. The neuron output module generates a round of QKV matrix multiplication result value and saves the round of QKV matrix result value to external storage or memory. Step 17: Repeat steps 14, 15, and 16 until all QKV matrices have been calculated. Step 18: If the input vector and the layer input are inconsistent, the input vector is split and grouped, and the final result is obtained through a combination of multiple groups and multiple rounds of calculation. If position encoding is required in the above steps, the data can be modified externally to insert the position encoding, or the vector after inserting the position encoding can be generated by setting the bias term weight parameters of the weight term current modulation control module.

44. The computing circuit system according to claim 3, characterized in that: The differential capacitor control module has a capacitance value detection circuit; the capacitance value detection circuit is equipped with multiple small capacitors with an initial voltage of zero, which can be selectively connected to the main capacitor of the capacitor module of the same attribute by a switch. The digital logic circuit, which shares the same attribute, is connected to the detection capacitance forward converter circuit and the voltage comparator of the capacitor control module. Based on the comparison between the main capacitor of the capacitor module and the target voltage, it outputs a digital output of the variable capacitance value generated by the variable capacitance operation of the detection capacitance forward converter circuit. The neuron output module is configured to generate the differential multiplication and addition information and the sign information based on the digital output of the differential capacitor control module.

45. The computing circuit system according to claim 44, characterized in that: The forward-changing circuit for detecting capacitance value contains multiple small forward-changing capacitors with unequal capacitance values ​​and a ratio that is a multiple of 2, forming a binary capacitor network. The differential capacitor control module includes a reverse-changing circuit for detecting capacitance value. This reverse-changing circuit contains multiple small reverse-changing capacitors whose initial voltage is the same as the initial voltage of the main capacitor of the capacitor module and which can be selectively connected to the main capacitor of the capacitor module with the same attribute by a switch. After the forward-changing circuit for detecting capacitance value operates on the small forward-changing capacitor that is not the minimum bit, when one of the voltage comparators in the two sets of differential capacitor control modules is lower than the target reference voltage, the two sets of reverse-changing circuits for detecting capacitance value synchronously operate on the connected small reverse-changing capacitors, thereby increasing the voltage of the differential capacitor module.

46. ​​The computing circuit system according to claim 3, characterized in that, There is a detection output method that changes the capacitance value of the main capacitor of the capacitor module: Step 1. The differential system charge adjustment module selects the main capacitor to be connected to the positive capacitor module or the negative capacitor module based on the values ​​of the neuron input and weights, as well as the sign of their product, and adjusts the capacitance value of the main capacitor of the capacitor module. Step 2. Initialize the main capacitor voltage of the capacitor module and the small capacitor voltage of the capacitance detection inverter circuit to the same voltage, and initialize the small capacitor voltage of the capacitance detection forward converter circuit to 0; Step 3. Detect the forward converter circuit and turn off the grounding switch; detect the reverse converter circuit and turn off the initialization voltage source switch; and turn off the initialization voltage source switch for the main capacitor of the capacitor module. Step 4. If the result of this step can be predicted, the positive converter circuit is tested and the corresponding small capacitor of the lower position is replaced according to the prediction; the positive converter small capacitor of the current corresponding position is connected in parallel; the voltage comparator judges that if the main capacitor voltage is greater than the target reference voltage, this step is repeated. Step 5. If the capacitance value operation does not reach the minimum bit, execute the inverting circuit operation, simultaneously connecting the main capacitors of the differential positive and negative systems to the inverting small capacitor, and repeat this inverting operation to raise its voltage to be greater than the target reference voltage; check the forward conversion circuit and replace the small capacitor with the corresponding capacitance value of the lower bit. Jump back to step 4; Step 6. The differential system outputs the corresponding digital code of the capacitance value; Step 7. The neuron calculates and outputs the difference based on the output of the difference system.

47. The computing circuit system according to claim 2, characterized in that, The working method of computation with forward propagation of neural networks: Step 1: Select the charging and discharging mode according to the calculation mode of the neural network layer, and select a specific initial voltage and charging and discharging driving voltage according to the charging and discharging mode. Step 2: The capacitor module is pre-charged and discharged to a specific initial voltage; Step 3: Generate a modulation current using a modulation control signal based on the correction values ​​of the input item fast access unit and the weight item fast access unit. Step 4: The capacitor module is charged and discharged by the modulation current. Step 5: The capacitor module continues to be charged and discharged through the capacitor control module. Step 6: The voltage of the capacitor module is compared with the reference voltage or other capacitor voltages, and a detection signal is output. Step 7: Combine the detection signals from the positive and negative capacitor control modules to obtain information on the voltage comparison detection between the positive and negative capacitor modules.