Parallel polynomial gradient computation system and method based on cmos synapse transistor

By combining a parallel array architecture and voltage pulse programming in the floating substrate operation mode of CMOS neural synaptic transistors, high-precision polynomial gradient calculations were achieved, solving the process compatibility and reliability issues of memristor solutions, reducing production costs, and improving the manufacturability and computational efficiency of the system.

CN121981183BActive Publication Date: 2026-06-26SHENZHEN HOTCHIP TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN HOTCHIP TECH
Filing Date
2026-04-03
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In existing technologies, memristor solutions suffer from problems such as incompatibility between manufacturing processes and standard CMOS processes, low device yield, large differences between devices, state drift, and high power consumption, making it difficult to achieve high-precision parallel polynomial gradient calculations.

Method used

By employing CMOS synaptic transistors in floating substrate operation mode and combining them with a parallel array architecture, the conductivity state of the CMOS transistors is adjusted through a voltage pulse programming circuit to achieve dual-mode characteristics of neuronal and synaptic behavior. A neuromorphic computing system based on standard CMOS process is constructed. Utilizing mature CMOS process manufacturing technology, new materials such as memristors and back-end integration processes are avoided, resulting in a highly reliable, low-power, and easily integrated parallel polynomial gradient computing system. This solves the problems of poor process compatibility, low yield, and large device differences faced by memristor solutions, and achieves high-precision polynomial gradient calculation.

Benefits of technology

It achieves high-precision polynomial gradient calculation based on standard CMOS process, reduces production cost, improves system manufacturability and yield, overcomes device differences and state drift problems in memristor schemes, supports short-term and long-term flexibility, and shortens calculation time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121981183B_ABST
    Figure CN121981183B_ABST
Patent Text Reader

Abstract

The application discloses a parallel polynomial gradient calculation system and method based on a CMOS nerve synapse transistor, and relates to the technical field of calculation systems.The parallel polynomial gradient calculation system based on the CMOS nerve synapse transistor adopts a basic calculation unit containing at least one CMOS transistor, and configures the CMOS transistor to exhibit the dual-mode characteristics of neuron behavior and synapse behavior in a floating substrate operation mode, and meanwhile, combines a parallel array architecture, so that the application realizes a neuromorphic calculation hardware based on a standard CMOS process, utilizes mature CMOS manufacturing technology, avoids special materials and post-integration processes required by emerging devices such as memristors, thereby improving the manufacturability and yield of the system, and reducing production costs.Meanwhile, the inherent device consistency and long-term stability of the CMOS transistor effectively overcome the problems of large differences between devices and state drift in the memristor scheme, and realize high-precision polynomial gradient calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computing system technology, specifically to a parallel polynomial gradient calculation system and method based on CMOS neural synaptic transistors. Background Technology

[0002] Neuromorphic computing is a novel computing paradigm that draws inspiration from the information processing mechanisms of biological nervous systems. Its core lies in directly implementing the neuronal and synaptic functions of artificial neural networks through hardware circuits. In recent years, with the rapid development of big data and artificial intelligence technologies, traditional architectures have faced bottlenecks such as the memory wall and power consumption wall. Neuromorphic computing hardware has gained traction due to its in-memory computing and parallel processing capabilities. In existing technologies, researchers have proposed various hardware acceleration schemes to achieve efficient polynomial gradient calculations to accelerate optimization algorithms. For example, IBM's NorthPole chip integrates storage and computing units, enabling large-scale matrix operations on-chip and improving the efficiency of neural network inference. Furthermore, memristor-based cross-array structures have been widely studied for simulating vector-matrix multiplication. By storing weights in the conductance of the memristor, it can perform multiplication and accumulation operations in the analog domain, thus supporting hardware acceleration of optimization algorithms such as gradient descent. These schemes have demonstrated excellent energy efficiency and computational throughput in specific application scenarios.

[0003] However, the aforementioned existing technologies still have the following problems in practical applications. For memristor-based cross-array schemes, although memristors have non-volatility, programmability, and high-density integration potential, memristors typically use special materials such as metal oxides, and their manufacturing processes are incompatible with standard CMOS processes. They need to be additionally integrated in the back-end processes of chip manufacturing, which increases manufacturing costs and complexity. Moreover, the yield of memristor devices is generally low, making it difficult to ensure that all cells can work properly in large-scale arrays. Furthermore, differences in the conductivity characteristics between devices lead to decreased computational accuracy and difficulty in algorithm convergence. Memristors are prone to state drift under continuous operation, and their resistance values... The long-term stability of weights is affected by changes in time or the number of pulses. At the same time, leakage current in the cross array further aggravates power consumption and signal interference, limiting the expansion of array size. For digital accelerator chips such as NorthPole, although they use mature CMOS technology, their computing units are still based on traditional digital logic, making it difficult to fully utilize the parallelism and energy efficiency advantages of analog computing. Moreover, complex digital circuit resources are required when implementing high-order polynomial gradient calculations. Therefore, how to achieve parallel computing hardware that combines neuron simulation and synaptic plasticity while maintaining the maturity and high reliability of CMOS technology has become a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a parallel polynomial gradient calculation system and method based on CMOS synaptic transistors. By employing a basic computing unit containing at least one CMOS transistor and configuring the CMOS transistor to exhibit dual-mode characteristics of neuronal and synaptic behavior in a floating substrate operation mode, combined with a parallel array architecture, this invention realizes neuromorphic computing hardware based on standard CMOS technology. Utilizing mature CMOS manufacturing technology, it avoids the special materials and back-end integration processes required by emerging devices such as memristors, thereby improving the system's manufacturability and yield, and reducing production costs. At the same time, the inherent device consistency and long-term stability of CMOS transistors effectively overcome the problems of large inter-device differences and state drift in memristor solutions, achieving high-precision polynomial gradient calculation.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a parallel polynomial gradient calculation system based on CMOS neural synaptic transistors, comprising:

[0006] Multiple basic computing units, each containing at least one CMOS transistor, the CMOS transistor being configured to exhibit dual-mode characteristics of neuronal and synaptic behavior in a floating substrate operation mode;

[0007] A voltage pulse programming circuit, connected to each basic computing unit, is used to apply voltage pulses to the CMOS transistor to adjust its conductivity state, thereby programming the basic computing unit to simulate synaptic weights.

[0008] The parallel array architecture organizes the multiple basic computing units into an array for parallel execution of polynomial gradient calculations.

[0009] Preferably, the floating substrate operation mode refers to floating the substrate terminals of the CMOS transistor and controlling its threshold behavior by adjusting the gate voltage and substrate bias to realize the excitation and recovery process of the neuron.

[0010] Preferably, in the floating substrate operation mode, the CMOS transistor can enhance and suppress synaptic weights by applying voltage pulses of different amplitudes, widths or frequencies, and support short-term and long-term plasticity.

[0011] Preferably, each basic computing unit is a 2-transistor unit, including a first CMOS transistor and a second CMOS transistor, wherein the first CMOS transistor is configured to simulate a neuron and the second CMOS transistor is configured to simulate a synapse, and the substrates of the first CMOS transistor and the second CMOS transistor are both floating.

[0012] Preferably, the gate of the first CMOS transistor is connected to the input signal line, and the drain and source are connected to the power supply and ground, respectively. The neuron threshold can be adjusted through a substrate resistance adjustment circuit. The gate of the second CMOS transistor is connected to the pulse programming line, and the drain and source are connected to the input and output lines, respectively. Its channel conductance is changed by the gate voltage pulse to store synaptic weights.

[0013] Preferably, the parallel array architecture is a cross array structure, wherein each basic computing unit is located at the intersection of word lines and bit lines, word lines are used for input signals or control signals, and bit lines are used for output signals, so that multiple units can perform multiplication and accumulation operations simultaneously.

[0014] Preferably, the system further includes a row driver and a column readout circuit, wherein the row driver is used to provide a voltage pulse sequence to the word line, and the column readout circuit is used to acquire the bit line current or voltage and convert it into a digital signal to complete gradient calculation.

[0015] Preferably, the system further includes a control logic unit, which generates the update weights required for the gradient descent algorithm according to a preset polynomial function, and controls the row driver to generate corresponding voltage pulses to program the synaptic weights of each basic computing unit.

[0016] Preferably, the control logic unit is further configured to read the array output in each iteration, calculate the polynomial gradient value, and adjust the voltage pulse parameters based on the gradient value to update the weights until convergence.

[0017] This invention also discloses a parallel polynomial gradient calculation method based on CMOS synaptic transistors, applied to the aforementioned parallel polynomial gradient calculation system based on CMOS synaptic transistors, comprising the following steps:

[0018] Step S1: Configure multiple basic computing units as CMOS transistors operating in floating substrate mode and organize them into a parallel array;

[0019] Step S2: Apply initialization voltage pulses to each basic computing unit through a voltage pulse programming circuit to set the initial synaptic weights;

[0020] Step S3: Apply the polynomial variable values ​​as input signals to the rows of the array, and obtain the polynomial values ​​and their gradient-related intermediate results by parallel computing of the array for each unit.

[0021] Step S4: Read the column output of the array, convert it into a digital signal through the readout circuit, and calculate the gradient according to the preset gradient descent algorithm;

[0022] Step S5: Based on the calculated gradient, the control logic unit generates an update pulse, and the synaptic weights of the corresponding basic calculation units are adjusted through the voltage pulse programming circuit.

[0023] Step S6: Repeat steps S3 to S5 until the polynomial gradient satisfies the convergence condition or reaches the predetermined number of iterations.

[0024] The technical effects and advantages of this invention are as follows:

[0025] 1. This parallel polynomial gradient calculation system based on CMOS synaptic transistors, by employing a basic computing unit containing at least one CMOS transistor and configuring the CMOS transistor to exhibit dual-mode characteristics of neuronal and synaptic behavior in a floating substrate operation mode, combined with a parallel array architecture, realizes neuromorphic computing hardware based on standard CMOS technology. Utilizing mature CMOS manufacturing technology, it avoids the special materials and back-end integration processes required by emerging devices such as memristors, thereby improving the system's manufacturability and yield, and reducing production costs. At the same time, the inherent device consistency and long-term stability of CMOS transistors effectively overcome the problems of large inter-device differences and state drift in memristor solutions, achieving high-precision polynomial gradient calculation.

[0026] 2. This parallel polynomial gradient calculation system based on CMOS synaptic transistors applies voltage pulses of different amplitudes, widths, or frequencies to CMOS transistors through a voltage pulse programming circuit to adjust their conductivity state and program basic calculation units to simulate synaptic weights. This achieves fine-tuning and multi-level storage of weights, enabling the system to support both short-term and long-term plasticity. Specifically, it can flexibly adjust the polynomial coefficients corresponding to each basic calculation unit according to the needs of the gradient descent algorithm. Compared to traditional memristor arrays, which can only achieve a finite number of weight updates, this invention expands the dynamic range and adjustment accuracy of weights through diverse combinations of pulse parameters, thereby improving the convergence speed of the optimization algorithm and the accuracy of the final solution.

[0027] 3. This parallel polynomial gradient calculation system based on CMOS neural synaptic transistors organizes multiple basic computing units into a cross array architecture and combines row drivers, column readout circuits, and control logic units to construct a complete parallel polynomial gradient calculation pipeline. This enables multiple basic computing units to perform multiplication and accumulation operations simultaneously. The control logic unit generates the updated weights required by the gradient descent algorithm according to a preset polynomial function and coordinates the voltage pulse programming circuit to complete the iterative update of the weights. This closed-loop hardware-algorithm co-design realizes parallel processing of the entire process from input variables to gradient output, which greatly shortens the calculation time of a single iteration. It is especially suitable for real-time optimization scenarios of high-order or multivariate polynomials. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 This is a diagram of the overall system architecture of the present invention;

[0030] Figure 2 This is a diagram of the parallel array architecture of the present invention;

[0031] Figure 3 This is a flowchart of the method of the present invention;

[0032] Figure 4 This is the internal logic diagram of the control logic unit of the present invention;

[0033] Figure 5 This is a schematic diagram of neuronal behavior simulation according to the present invention. Detailed Implementation

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] This embodiment discloses a parallel polynomial gradient calculation system based on CMOS neural synaptic transistors, such as Figures 1 to 5 As shown, this system utilizes the dual-mode characteristics of standard CMOS transistors in floating substrate operation mode to construct programmable neural-synaptic basic computing units, and achieves efficient polynomial gradient calculation through a parallel array architecture, thereby overcoming the bottlenecks in process compatibility, reliability and integration of emerging devices such as memristors.

[0036] System overall architecture reference Figure 1The parallel polynomial gradient calculation system provided by this invention includes: a control logic unit, a row driver, a voltage pulse programming circuit, a parallel array architecture, and a column readout circuit. The control logic unit, as the core controller of the system, is used to generate the update weights required for the gradient descent algorithm according to a preset polynomial function and to coordinate the work of each module. The row driver is connected to the word line of the parallel array architecture and is used to provide a voltage pulse sequence to the word line. These pulses can be used as input signals or as programming signals. The voltage pulse programming circuit is connected to each basic computing unit and is used to apply voltage pulses with specific parameters to the CMOS transistor to adjust its conduction state, thereby programming the basic computing unit to simulate synaptic weights. The parallel array architecture is organized by multiple basic computing units and is used to perform multiply-accumulate operations in polynomial gradient calculation in parallel. The column readout circuit is connected to the bit line of the parallel array architecture and is used to collect the current or voltage on the bit line and convert it into a digital signal for the control logic unit to perform gradient calculation and decision-making.

[0037] Parallel array architecture reference Figure 2 The parallel array architecture preferably adopts a cross-array structure, in which multiple word lines (row lines) intersect multiple bit lines (column lines) perpendicularly. A basic computing unit is set at each intersection point. Word lines are used for input signals or control signals, and bit lines are used for output signals. This layout enables multiple basic computing units to perform multiplication and accumulation operations simultaneously, thereby realizing large-scale parallel computing. The row driver is located on one side of the word line and provides voltage pulses to each word line; the column readout circuit is located on one side of the bit line and collects and converts the output of each bit line.

[0038] The basic computing unit is the core component of the system. Each basic computing unit contains at least one CMOS transistor, which is configured to operate in a floating substrate operation mode. The floating substrate operation mode means that the substrate terminal of the CMOS transistor is floating and not connected to a fixed potential, so that its threshold behavior is sensitive to external bias. In this mode, the threshold voltage of the transistor can be controlled by adjusting the gate voltage and the substrate bias, simulating the activation and recovery process of neurons. The activation of a neuron is the generation of a pulse when the input exceeds the threshold, and the recovery of a neuron is the return to the resting state after activation. At the same time, by applying voltage pulses of different amplitudes, widths or frequencies, the channel conductance of the transistor can be changed to simulate the enhancement (long-term enhancement, LTP) and inhibition (long-term inhibition, LTD) of synaptic weights, and support the modulation of short-term plasticity (STP) and long-term plasticity (LTP).

[0039] In a preferred embodiment, each basic computing unit is a 2-transistor unit, i.e., an NS-RAM unit, including a first CMOS transistor and a second CMOS transistor. The first CMOS transistor is configured to simulate a neuron, and the second CMOS transistor is configured to simulate a synapse. The substrates of both the first and second CMOS transistors are floating. The specific connection method is as follows: the gate of the first CMOS transistor is connected to the input signal line, and the drain and source are connected to the power supply and ground, respectively. The neuron threshold is adjustable through a substrate resistance adjustment circuit. The gate of the second CMOS transistor is connected to the pulse programming line, and the drain and source are connected to the input line and the output line, respectively. Its channel conductance is changed by the gate voltage pulse to store synaptic weights. The substrate resistance adjustment circuit may include a variable resistor or another transistor for finely adjusting the substrate bias of the first transistor, thereby changing its threshold.

[0040] Control logic unit reference Figure 4 The control logic unit internally implements the iterative process of the gradient descent algorithm. First, it determines the coefficients for gradient calculation based on a preset polynomial function, such as the user-specified polynomial to be optimized. In one iteration, the control logic unit sets the initial synaptic weights of the basic computational units through row drivers and voltage pulse programming circuits. These initial synaptic weights correspond to the polynomial coefficients. Then, the polynomial variable values ​​are applied as input signals to the word lines of the array. Through the multiplication and accumulation operations of the parallel array, the column readout circuit obtains the outputs of each bit line. These outputs represent the polynomial value under the current input and gradient-related intermediate results, such as the values ​​of each monomial. Based on these outputs, the control logic unit calculates the gradient of the current coefficients using the gradient descent algorithm and generates updated weights. Subsequently, it controls the row drivers and voltage pulse programming circuits to generate corresponding voltage pulses, adjusting the synaptic weights of the corresponding basic computational units to complete one weight update. This process is repeated until the gradient converges or the predetermined number of iterations is reached.

[0041] Voltage pulse programming circuits can generate pulses with various parameters, including pulses of different amplitudes, widths, and frequencies, such as... Figure 4 As shown, high-amplitude, wide-width, or high-frequency pulses are typically used to achieve long-term plasticity (LTP), resulting in enhanced synaptic weights; while low-amplitude, narrow-width, or low-frequency pulses are used for short-term plasticity (STP) or weight suppression. The programming circuit selects appropriate pulse parameters according to the instructions of the control logic unit and applies them to the pulse programming line of the target basic computing unit, thereby precisely controlling the synaptic weights of each unit.

[0042] Method and process reference Figure 3The present invention also provides a parallel polynomial gradient calculation method based on CMOS synaptic transistors, which is applied to a parallel polynomial gradient calculation system based on CMOS synaptic transistors. The method includes the following steps:

[0043] Step S1 involves configuring multiple basic computing units as CMOS transistors operating in a floating substrate mode and organizing them into a parallel array. Specifically, this includes ensuring that the substrates of all transistors are floating and setting an initial threshold range through a substrate resistance adjustment circuit.

[0044] Step S2: Apply initialization voltage pulses to each basic calculation unit through the voltage pulse programming circuit to set the initial synaptic weights. These initial weights can be set randomly or based on prior knowledge, corresponding to the initial coefficients of the polynomial.

[0045] Step S3: Apply the polynomial variable values ​​as input signals to the rows of the array. Calculate the output of each unit in parallel using the array to obtain the polynomial values ​​and their gradient-related intermediate results. For example, for the polynomial:

[0046] ;

[0047] Each basic computing unit can store one coefficient. The input x is broadcast via the row lines, and the unit output is... After the bit lines are converged, f(x) is obtained. The gradient calculation requires... These intermediate results can be obtained directly from the unit's output or through additional readouts.

[0048] Step S4: Read the column output of the array, convert it into a digital signal through the readout circuit, and calculate the gradient according to the preset gradient descent algorithm. The gradient descent algorithm can be the standard batch gradient descent or stochastic gradient descent. The gradient value is obtained by weighted summation of each intermediate result.

[0049] Step S5: Based on the calculated gradient, the control logic unit generates an update pulse, and the synaptic weights of the corresponding basic calculation units are adjusted through the voltage pulse programming circuit. The update rule is: new weight = old weight - learning rate × gradient. The amplitude, width or frequency of the update pulse is determined according to the required amount of change.

[0050] Step S6: Repeat steps S3 to S5 until the polynomial gradient satisfies the convergence condition (e.g., the gradient norm is less than a threshold) or the predetermined number of iterations is reached.

[0051] Example 1: This example uses the univariate quadratic polynomial f(x)=ax²+bx+c as an example to illustrate the system's workflow. The system is pre-configured with three basic calculation units to store coefficients a, b, and c respectively, and the control logic unit is set with a learning rate η and the number of iterations N.

[0052] Initialization: Initialization pulses are applied to the three units through the voltage pulse programming circuit to set a, b, and c to initial values, such as all zero or random small values. The synaptic weights of each unit are the corresponding coefficients.

[0053] Iteration begins: Take a set of training samples (x, y), where y is the target value. Apply x to the row lines of the array. For each cell, the outputs are ax², bx, and c, respectively. After the bit lines are collected, the following is obtained:

[0054] f(x) = ax² + bx + c;

[0055] Meanwhile, x², x, and 1 can be directly provided as intermediate results within the unit (through an appropriate readout mechanism), and the column readout circuit converts these analog values ​​into digital signals and sends them to the control logic unit.

[0056] Gradient calculation: The control logic unit calculates the partial derivative of the loss function L=(f(x)-y)² with respect to each coefficient:

[0057] ;

[0058] The gradient values ​​are these partial derivatives (or corrected by the algorithm).

[0059] Update weights: The control logic unit generates update values.

[0060] ;

[0061] Then, the control voltage pulse programming circuit generates corresponding pulses, which are applied to the pulse programming lines of the cells storing a, b, and c respectively, adjusting their channel conductance to achieve weight updates. For example, when a needs to be increased, an enhancement pulse, such as a high-amplitude wide pulse, is applied; when a needs to be decreased, a suppression pulse, such as a low-amplitude narrow pulse, is applied.

[0062] Repeat: Repeat steps 2-4 for the next sample until all samples have been processed or the convergence condition has been met.

[0063] Through iteration with multiple samples, the system eventually converges to the optimal coefficients, completing the polynomial fitting.

[0064] Example 2: This example uses a tenth-degree polynomial For example, to demonstrate the system's massively parallel capabilities, the system contains 11 basic computing units arranged in an array of one row (or one column). Since each unit independently stores a coefficient and can compute simultaneously... Therefore, a single forward propagation can yield the complete value of f(x), which is much faster than traditional serial computation. During the gradient calculation phase, each unit outputs simultaneously. The control logic units read and calculate gradients in parallel, and then update all units in parallel, realizing true parallel gradient descent. This architecture is particularly suitable for high-dimensional polynomial fitting and function approximation problems. The computation time is independent of the polynomial degree and is only limited by the array size and readout speed.

[0065] Example 3: This example considers a bivariate quadratic polynomial f(x,y)=ax²+by²+cxy+dx+ey+f. The system needs to store six coefficients, therefore requiring at least six basic computational units. To input two variables x and y simultaneously, the array can be designed as a two-dimensional structure: each basic computational unit simultaneously receives a certain combination of x and y. One implementation is to group the units by rows and columns. For example, the first row of units processes x-related terms, the second row processes y-related terms, the third row processes xy terms, and so on. During input, x is applied to the word line of the first row, and y is applied to the word line of the second row. The unit processing xy terms may need to receive multiple combinations of x and y simultaneously. The inputs x and y can be obtained through two input lines or by using a multiplier. Alternatively, the programmability of the units can be utilized to map the powers of x and y to different units. For example, unit 1 stores a with the input x²; unit 2 stores b with the input y²; unit 3 stores c with the input xy; unit 4 stores d with the input x; unit 5 stores e with the input y; and unit 6 stores f with the input 1. After inputting x and y, all units compute in parallel, and the bit lines are summed to obtain f(x,y). The gradient calculation is similar to that in Example 1, where partial derivatives are calculated with respect to each variable, and the corresponding coefficients are updated. This example demonstrates the scalability of the system in multivariate function optimization.

[0066] A key technical detail to emphasize is the physical mechanism of the floating substrate operation mode: When the substrate of a CMOS transistor is floating, the substrate potential is determined by the charge state of the gate, source / drain, and the substrate itself. Changes in the gate voltage affect the substrate potential through capacitive coupling, thereby changing the threshold voltage. This effect makes the transistor exhibit characteristics similar to neuron integration-firing. At the same time, by applying programming pulses, charges can be captured in the gate oxide layer or at the substrate interface to achieve non-volatile or semi-volatile weight storage, simulating synaptic plasticity.

[0067] The amplitude of the programming pulse determines the change in channel conductance; the pulse width affects the charge injection time, which in turn affects the amount of change; the pulse frequency is related to short-term enhancement / suppression. High-frequency pulses tend to induce long-term effects, while low-frequency pulses produce short-term effects. By carefully designing the pulse waveform, multi-level weight storage can be achieved, thereby improving computational accuracy.

[0068] The function of the substrate resistance adjustment circuit is as follows: By changing the equivalent resistance from the substrate to ground, the neuron threshold can be dynamically adjusted so that the same unit has different response characteristics in different modes, thereby adapting to different computational needs. For example, a low threshold is used in the early stage of computation to achieve fast convergence, and the threshold is increased in the later stage to stabilize the weights.

[0069] The readout methods for parallel arrays are as follows: the readout of bit line current or voltage can be achieved using a transimpedance amplifier or integrator to convert analog quantities into digital quantities. Due to the large array size, the influence of parasitic resistance and capacitance must be considered, and differential readout or calibration techniques can be used to improve accuracy.

[0070] The control logic unit can be a digital circuit (such as a microcontroller or FPGA) or a mixed-signal circuit. It stores a preset polynomial form and executes a gradient descent algorithm. The learning rate in the algorithm can be dynamically adjusted to accelerate convergence.

[0071] In summary, this invention utilizes the dual-mode characteristics of standard CMOS transistors to construct a highly reliable, low-power, and easily integrated parallel polynomial gradient calculation system, solving the problems of poor process compatibility, low yield, and large device differences faced by memristor solutions.

[0072] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A parallel polynomial gradient calculation system based on CMOS neural synaptic transistors, characterized in that, include: Multiple basic computing units, each containing at least one CMOS transistor, the CMOS transistor being configured to exhibit dual-mode characteristics of neuronal and synaptic behavior in a floating substrate operation mode; A voltage pulse programming circuit, connected to each basic computing unit, is used to apply voltage pulses to the CMOS transistor to adjust its conductivity state, thereby programming the basic computing unit to simulate synaptic weights. The parallel array architecture organizes the multiple basic computing units into an array for parallel execution of polynomial gradient calculations. The system also includes a control logic unit, which implements the iterative process of the gradient descent algorithm. First, the coefficients for which the gradient needs to be calculated are determined based on a preset polynomial function. In one iteration, the control logic unit sets the initial synaptic weights of the basic computational units through row drivers and voltage pulse programming circuits. The initial synaptic weights correspond to the polynomial coefficients. Then, the variable values ​​of the polynomial are applied as input signals to the word lines of the array. Through the multiplication and accumulation operations of the parallel array, the column readout circuit obtains the outputs of each bit line. These outputs represent the values ​​of the polynomial under the current input and the intermediate results related to the gradient. Based on these outputs, the gradient of the current coefficients is calculated using the gradient descent algorithm, and updated weights are generated. Subsequently, it controls the row drivers and voltage pulse programming circuits to generate corresponding voltage pulses, adjusts the synaptic weights of the corresponding basic computational units, completes one weight update, and repeats the above process until the gradient converges or the predetermined number of iterations is reached.

2. The parallel polynomial gradient calculation system based on CMOS synaptic transistors according to claim 1, characterized in that, The floating substrate operation mode refers to floating the substrate terminals of a CMOS transistor and controlling its threshold behavior by adjusting the gate voltage and substrate bias to achieve the excitation and recovery process of neurons.

3. The parallel polynomial gradient calculation system based on CMOS synaptic transistors according to claim 2, characterized in that, In floating substrate operation mode, the CMOS transistor can enhance and suppress synaptic weights by applying voltage pulses of different amplitudes, widths or frequencies, and supports short-term and long-term plasticity.

4. The parallel polynomial gradient calculation system based on CMOS synaptic transistors according to claim 1, characterized in that, Each basic computing unit is a 2-transistor unit, including a first CMOS transistor and a second CMOS transistor. The first CMOS transistor is configured to simulate a neuron, and the second CMOS transistor is configured to simulate a synapse. The substrates of both the first CMOS transistor and the second CMOS transistor are floating.

5. The parallel polynomial gradient calculation system based on CMOS synaptic transistors according to claim 4, characterized in that, The gate of the first CMOS transistor is connected to the input signal line, and the drain and source are connected to the power supply and ground, respectively. The neuron threshold can be adjusted through a substrate resistor adjustment circuit. The gate of the second CMOS transistor is connected to the pulse programming line, and the drain and source are connected to the input and output lines, respectively. Its channel conductance is changed by the gate voltage pulse to store synaptic weights.

6. The parallel polynomial gradient calculation system based on CMOS synaptic transistors according to claim 1, characterized in that, The parallel array architecture is a cross array structure, in which each basic computing unit is located at the intersection of word lines and bit lines. Word lines are used for input signals or control signals, and bit lines are used for output signals, enabling multiple units to perform multiplication and accumulation operations simultaneously.

7. The parallel polynomial gradient calculation system based on CMOS synaptic transistors according to claim 6, characterized in that, The system also includes a row driver and a column readout circuit. The row driver is used to provide a voltage pulse sequence to the word line, and the column readout circuit is used to acquire the bit line current or voltage and convert it into a digital signal to complete the gradient calculation.

8. A parallel polynomial gradient calculation method based on CMOS synaptic transistors, applied to the parallel polynomial gradient calculation system based on CMOS synaptic transistors as described in any one of claims 1 to 7, characterized in that, Includes the following steps: Step S1: Configure multiple basic computing units as CMOS transistors operating in floating substrate mode and organize them into a parallel array; Step S2: Apply initialization voltage pulses to each basic computing unit through a voltage pulse programming circuit to set the initial synaptic weights; Step S3: Apply the polynomial variable values ​​as input signals to the rows of the array, and obtain the polynomial values ​​and their gradient-related intermediate results by parallel computing of the array for each unit. Step S4: Read the column output of the array, convert it into a digital signal through the readout circuit, and calculate the gradient according to the preset gradient descent algorithm; Step S5: Based on the calculated gradient, the control logic unit generates an update pulse, and the synaptic weights of the corresponding basic calculation units are adjusted through the voltage pulse programming circuit. Step S6: Repeat steps S3 to S5 until the polynomial gradient satisfies the convergence condition or reaches the predetermined number of iterations.

Citation Information

Patent Citations

  • Heterogeneous neural network acceleration device based on cooperation of DSP and memristor

    CN120046677A

  • Neuromorphic memory circuit and method of neurogenesis for an artificial neural network

    EP4334850A1