Basic structure and circuit system of a neural network large-scale parallel computing circuit and method for parameter adjustment setting operation thereof
Patent Information
- Application Number
- CN202511437759.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-08-21
- Filing Date
- 2025-10-09
- Publication Date
- 2026-09-22
AI Technical Summary
设计复杂、网络设计不灵活、数据精度低、网络规模小、性能也低
Smart Images

Figure CN122797643A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a computational dedicated circuit system and method in the field of artificial intelligence, particularly the basic structure of a large-scale neural network using a hybrid digital and analog circuit, the composition and implementation of its hybrid digital and analog circuit system, its layout and operation method, and the parameter adjustment and setting method within the circuit system. More specifically, it describes how to implement circuit systems and methods for multiply-accumulate calculations and linear and nonlinear activation of neural networks, and further, how to construct circuits, layouts, and methods for large-scale neural network systems. The hardware and software technologies of this invention can be applied to printed circuits or chips / integrated circuits. Background Technology
[0002] With the rapid development of artificial intelligence (AI) technology, neural networks, especially ultra-large-scale neural networks, have demonstrated remarkable capabilities in various fields such as image recognition, natural language processing, and speech recognition. DeepSeek and other technologies have proven that by employing specific techniques and optimization strategies, it is possible to train and apply large-scale models at a relatively low cost. Countries worldwide have invested trillions of dollars in this area in hopes of achieving AGI; however, current implementation methods still have significant limitations.
[0003] I. Solution based on computing chip + GPU: First, existing implementations of ultra-large-scale neural networks primarily rely on server clusters and expensive GPUs. Essentially, they involve digital simulations of complex computational processes, leading not only to high hardware costs but also increased system energy consumption and physical size. High energy consumption raises operating costs and negatively impacts the environment; while the massive system size limits their usability in mobile devices or space-constrained applications. Furthermore, because these systems typically require continuous high-performance computing support, their power supply requirements are very stringent, further limiting the flexibility of their application scenarios.
[0004] Secondly, while technological advancements by companies like DeepSeek have significantly reduced the cost of training and deploying large-scale neural networks, their implementation still heavily relies on complex and sophisticated software programming techniques. This means that development teams must possess deep expertise and skills to effectively design, train, and optimize these models. For many small and medium-sized enterprises or research institutions with limited resources, such requirements constitute a significant barrier to entry, limiting the speed of technological innovation and adoption.
[0005] Finally, existing solutions still have several orders of magnitude of room for improvement in performance and energy efficiency compared to the human brain (which consumes approximately 20 watts). In particular, current technology has not yet reached its ideal state in terms of concurrency, processing speed, response time, and energy efficiency. For example, in certain real-time applications, such as autonomous vehicles or instant translation services, rapid and accurate decision-making is crucial, and existing solutions based on servers and computing chips + GPUs often fall short of meeting these requirements.
[0006] II. In-Memory Computing (IMC). As a cutting-edge technology, IMC has demonstrated enormous potential in the field of neural networks. However, despite its advantages and the abundance of solutions available, IMC still faces challenges in practical applications, including durability and reliability issues, low data accuracy, small network size, and incomplete activation functions, making it difficult to effectively support the training and inference of complex models. These problems have hindered the widespread adoption of IMC solutions in large-scale neural network applications.
[0007] III. Analog Circuit Implementation. Using analog circuits similar to operational amplifiers to implement neural networks currently offers no advantages over digital circuits. The design is complex, the network design is inflexible, the data accuracy is low, the network size is small, and the performance is also low.
[0008] In summary, current implementation schemes for ultra-large-scale neural networks all suffer from significant drawbacks and challenges. While the computing chip + GPU-based approach offers high computing power, its high cost, high energy consumption, and physical space constraints limit its widespread application. In-memory computing excels in reducing data transmission latency and power consumption, but issues with the durability and reliability of storage media, insufficient data precision, and difficulty in supporting large-scale model training and inference hinder its large-scale adoption. Analog circuit implementations face numerous problems such as design complexity, poor flexibility, low data precision, and poor performance, making it difficult to compete with digital circuits. Furthermore, in terms of network construction, traditional accumulation and transfer methods have failed to adequately integrate with other hardware technologies, resulting in performance and energy efficiency far below expectations. These issues collectively indicate that existing technologies and methods are insufficient to fully meet the practical needs of ultra-large-scale neural networks, necessitating innovative architectural designs and circuit implementations to overcome these limitations. Summary of the Invention
[0009] In view of the above challenges, this invention proposes a basic structure for a large-scale neural network suitable for parallel execution using a hybrid digital and analog circuit, as well as the composition, implementation, computational principles, and working methods of the hybrid digital and analog circuit, and the parameter adjustment and setting methods in its circuit system. This technology lays the technical foundation for future large-scale neural network hardware and points to a direction for software and hardware technology that surpasses the human brain. It is a major breakthrough with historical influence, achieving a long-term goal in the field. It utilizes existing chip technology (such as CMOS) to solve the long-standing problem of efficiently implementing parallel large-scale neural network computation using analog circuits, far exceeding traditional analog and digital solutions. Design simplification and manufacturing friendliness: simple, practical, parallelizable, and easy to standardize manufacturing; high tolerance for defects, easy to expand and partition, significantly reducing chip process requirements. Huge performance leap: achieving optimizations of 10^n times in multiple dimensions (especially energy efficiency). It possesses high computing power per unit area, high energy efficiency, efficient storage read / write, and can be set to high-precision computation, making it suitable for neural network training with great potential. Profound impact: It enables large models to be decoupled from servers / GPUs, and ontological intelligence to be embedded in terminal devices, ushering in the era of "AI everywhere"; it provides a systematic framework for overcoming software arbitrariness and technological bias, and developing matching software and hardware.
[0010] This invention discloses a basic structure of a neural network massively parallel computing circuit and a method for adjusting and setting its parameters.
[0011] This invention discloses a large-scale parallel computing circuit system for neural networks. The basic neuron unit includes: an input current modulation control module, a weight current modulation control module, a capacitor control module, a capacitor module, and a neuron output module. The capacitor control module and the capacitor modules are divided into positive and negative terminals. The positive capacitor control module has a circuit for charging and discharging the positive capacitor module, and the negative capacitor control module has a circuit for charging and discharging the negative capacitor module, capable of independently charging and discharging the corresponding capacitor module to a specified voltage. It has one or more reference time source signals. The input current modulation control module, based on an external signal, or the signal from the neuron output module, or the reference time source signal and... The input item fast access unit loads the value and generates the input item current modulation control signal; the weighted item current modulation control module, based on the value of its own weighted item fast access unit and the corresponding input item current modulation control signal, further selects, trims, or inserts a delay to generate a control signal for the modulation current, or adjusts the equivalent resistance of the current channel; multiple current channels with corresponding neural network links are capable of charging and discharging the positive and negative capacitor modules according to settings; each current channel has a switch for generating the modulation current, controlled by the control signal for the modulation current, to generate the modulation current calculated by the analog neural network for charging and discharging the capacitor modules; the input item fast access... The unit and the fast access unit for weighted terms jointly control whether to charge or discharge the positive or negative capacitor module; the capacitor control module has a current channel for detecting charging and discharging, and the channel contains an equivalent detection resistor, which can charge and discharge the corresponding capacitor module; the capacitor control module can precharge the capacitor module to a specific voltage to prepare for the simulated neural network calculation; after the modulation current of the simulated neural network calculation charges and discharges the capacitor module, the equivalent detection resistor continues to charge and discharge the capacitor module; the capacitor control module contains a voltage comparator, which makes the voltage of the capacitor in the capacitor module equal to the reference voltage during the energization of the equivalent detection resistor. The voltage comparison module compares the voltage of a capacitor within the capacitor module with that of other capacitors of the same sign to generate a voltage comparison signal, and shuts off the current channel for detecting charging and discharging based on this signal. The positive and negative capacitor control modules each output an equivalent time pulse based on the equivalent time from the on to the off state of the positive and negative current channels for detecting charging and discharging. The pulse width of the equivalent time pulse is the ratio of the duty cycle to the minimum duty cycle of the circuit in pulse width current modulation, the ratio of the pulse width time to the minimum pulse width time of the circuit in single-pulse pulse width modulation, and the number of switching capacitor operations in switched-capacitor unit current modulation. The neuron output module generates the level of the equivalent time difference or its digitally converted value based on the positive and negative equivalent time pulses.The neuron output module generates a sign level based on a comparison of positive and negative equivalent time pulses, or a comparison of the voltages of positive and negative capacitor modules.
[0012] This invention discloses that the capacitor module is internally a combination of capacitors and switches; the capacitor module contains a main capacitor and a detection capacitor; wherein a certain main capacitor and a certain detection capacitor are charged and discharged together through a switch combination; the remaining disconnected or uncombined detection capacitors are charged and discharged independently or used for voltage comparison.
[0013] This invention discloses a multi-layer neural network, with an analog layer as the middle layer. The analog layer uses equivalent single-pulse pulse width modulation (EPPWM) output. It includes a multi-layer circuit for detecting charge and discharge. The first layer generates a time difference pulse to control the charge and discharge of the detection capacitor in the next layer. The analog layer generates linear or non-linear activation outputs based on the selected charge and discharge mode and depth of the detection capacitor. The EPPWM output is subsequently used as input to its own layer or to other layers capable of receiving EPPWM. In switched-capacitor type unit current operation, the EPPWM is current modulation based on the number of operations.
[0014] This invention discloses a layered activation function numerical source that generates numerical information with equivalent time as the independent variable based on the setting of the activation function. The numerical information with equivalent time as the independent variable includes time values, activation function values, derivative values, intermediate calculation values, or combinations thereof. The neuron output module has a fast activation function access unit. The neuron output module selects the corresponding symbol's numerical information with equivalent time as the independent variable based on the symbol level. When the level expressing the equivalent time difference of capacitor charging and discharging begins to change or ends to change, the output value of the activation function of the corresponding neuron basic unit at the current time is obtained based on the numerical information with equivalent time as the independent variable for the corresponding symbol.
[0015] This invention discloses a computing circuit system that dynamically serves as one or more layers of a neural network to be computed; the input current modulation control module loads values from memory or external circuitry, or shares or loads values from neuron output modules, for the input fast access unit; the weight current modulation control module loads values from memory or external circuitry, or shares or loads values from the corresponding weight fast access unit, for the corresponding weight; at the end of the computation, the value of the activation function fast access unit of the last layer is output to memory, its own input, external circuitry, or other modules.
[0016] This invention discloses that the equivalent detection resistor of the capacitance control module for a computational neuron is an equivalent resistance network with multiple equivalent resistance values, which can accelerate the detection of charging and discharging. A reference neuron is included, whose output is the pulse output of the capacitance control module, used to control the charging and discharging time of the equivalent detection resistor. The reference neuron can control the pulse output value of the capacitance control module according to parameter settings. The capacitance control module of the computational neuron has multiple reference voltages for comparison, and the voltage range of the detection capacitor is determined based on the comparison results. The control logic switches and selects the corresponding equivalent detection resistor value according to the real-time voltage range of the detection capacitor to control the instantaneous current of accelerated charging and discharging. Simultaneously, the pulse output of the corresponding reference neuron is selected to control the duration of accelerated charging and discharging. The output of the computational neuron after accelerated charging and discharging by the positive and negative capacitance control modules obtains an equivalent time difference pulse. After the equivalent time difference pulse is digitized, the preset output value of the reference neuron is added according to the sign to obtain the multiplicative sum value before neuron activation.
[0017] The present invention discloses an array-type memory used as a cache; the array-type memory has a row-selection write circuit to write the output of a neural network layer to the selected row; the array-type memory has a transposed block-selection read circuit to write the data of a selected block in a column to the input or weight of a neural network layer.
[0018] The present invention also discloses one or more underlying high-frequency reference pulse signals; the input current modulation control module has a pulse selection or shielding circuit, which selects or shields the underlying high-frequency reference pulse signal according to the input current modulation control signal to generate a short pulse input current modulation control signal, or selects one of the underlying high-frequency reference pulse signals.
[0019] The present invention discloses that the input item fast access unit is segmented bit by bit, and each segment independently generates one or more segment input item modulation signals; the weight item fast access unit is segmented bit by bit; each segment of the weight item fast access unit independently generates multiple modulation currents according to its own segment weight value and the corresponding segment input item modulation signal; the multiple modulation currents simultaneously charge and discharge the capacitor module of the corresponding segment.
[0020] This invention discloses a system with multiple neurons, whose weight term current modulation control module has a shared weight term fast access unit; each neuron's input term current modulation control module generates multiple sets of modulation signals based on the shared weight term fast access unit and its own input value; multiple switches in the neuron's link current channel generate multiple sets of modulation currents based on the multiple sets of modulation signals; multiple neuron basic units work simultaneously or sequentially.
[0021] This invention discloses a dynamically alternating two-layer neural network, with one layer serving as the input layer and the other as the output layer, which are dynamically and alternately used as the various layers of the entire neural network. The dual-layer neural network has a bidirectional intermediate module that combines the input current modulation control module and the neuron output module. One layer of the bidirectional intermediate module has links and current channels for charging and discharging the capacitor modules of the corresponding neuron basic units through multiple weight current modulation control modules. The neuron basic unit has channels for loading values of weight fast access units and input fast access units from memory or external circuitry, and also channels for outputting values of activation function fast access units to memory or external circuitry. The output layer in the dynamically alternating dual layers is used as the layer being computed in the entire neural network by loading the values of the aforementioned fast access units.
[0022] This invention discloses a layer activation function numerical source, which generates activation function numerical information based on the setting of the activation function and sends it to the neuron output module of the corresponding layer; the neuron output module generates a new charge-discharge equivalent time based on the activation function numerical information; the neuron output module generates a new activation function value based on the new charge-discharge equivalent time and another activation function numerical information.
[0023] This invention discloses a multi-segment neuron output module that can convert the equivalent time signal input from the positive and negative capacitance control module of the corresponding segment into the value of the fast access unit; the neuron summary calculation output module can perform shifting, addition and subtraction on the value of the fast access unit of each segment neuron output module according to the corresponding bit range of its segment, and write the result value into its fast access unit.
[0024] This invention discloses a network summarization layer; the network summarization layer has network summarization layer neurons that connect to the basic computational units of neurons in the current layer, and uses the activated values of each neuron in the current layer as the input of the network summarization layer neurons; a layer computation module receives the output of the network summarization layer neurons and generates a layer parameter value; the layer computation module, based on the layer parameter value, regenerates activation function numerical information with equivalent time as the independent variable and sends it to the neuron output module of the corresponding layer.
[0025] This invention discloses a computational method for forward propagation of a neural network: Step 1, selecting a charging / discharging mode based on the computational mode of the neural network layer, and selecting a specific initial voltage and a driving voltage for charging / discharging based on the charging / discharging mode; Step 2, pre-charging and discharging the capacitor module to the specific initial voltage; Step 3, generating a modulation current through a modulation control signal based on the correction values of the input item fast access unit and the weight item fast access unit; Step 4, charging and discharging the capacitor module through the modulation current. Step 5: The capacitor module continues to charge and discharge through the capacitor control module; Step 6: The voltage of the capacitor module is compared with the reference voltage or other capacitor voltages, and a detection signal is output; Step 7: The positive and negative detection signals from the capacitor control module are combined to obtain information on the voltage comparison detection between the positive and negative capacitor modules.
[0026] This invention discloses a method for detecting signals using a neuron activation function within a comprehensive positive and negative capacitor control module: Step 1, acquiring voltage comparison detection information between the positive and negative capacitor modules; Step 2, a single neuron output module performs pre-alignment based on the level signal expressing the equivalent time difference between capacitor charging and discharging, i.e., at the moment the level of the equivalent time difference between capacitor charging and discharging begins to change, when this signal occurs, the detection of the charging and discharging process is interrupted; Step 3, waiting for all neuron output modules within the layer to complete the XOR gate signal pre-alignment; Step 4, all neuron output modules within the layer continue to detect the charging and discharging process; Step 5, based on the timing of the edge trigger signal at the end of the level signal expressing the equivalent time difference between capacitor charging and discharging, the corresponding neuron output module obtains the final function value information from the layer activation function value source and stores it in the activation function fast access unit; Step 6, acquiring the output function value information of the activation function fast access unit.
[0027] This invention discloses a dynamic working method: Step 1, the basic neuron unit loads the values of the weight term fast access unit and the input term fast access unit from the memory or external circuit; Step 2, the neural network propagates forward, i.e., the process of calculating the multiplicative sum and activation value; Step 3, in parallel, if the entire neural network has not reached the calculation of the last layer, the basic neuron unit loads the values of the weight term fast access unit of the new layer of the neural network from the memory or external circuit; Step 4, the basic neuron unit generates the output value, i.e., the capacitor detects charging and discharging, and if there is an activation function, it executes the activation function; Step 5, the current basic neuron unit generates the value of the output term fast access unit or the activation function fast access unit, and copies it to the input term fast access unit (if a shared fast access unit is used, the copying process is omitted); Step 6, if the entire neural network has reached the calculation of the last layer, the value of the activation function fast access unit or the output term fast access unit is output to the memory or external circuit, and the loop ends; Step 7, the above steps 2, 3, 4, 5, and 6 are repeated cyclically.
[0028] This invention discloses a method for mini-batch training: In mini-batch training, data is stored in each row, with forward propagation and data stored at the same node written sequentially in adjacent rows; Step 1, during forward propagation, the array-like storage caches the output of the activated layer in each computation of the neural network, row by row; Step 2, during backward propagation, the array-like storage caches the product of the partial derivatives of the activated output nodes of each layer in each computation of the neural network, and the derivative of the activation function, i.e., the multiplicative partial derivatives before activation; Step 3, the block-based reading circuit reads the multiplicative partial derivatives of each backward propagation node in the mini-batch corresponding to the block into the layer's input, and reads the multiple output values of each neuron's output node from the forward propagation of the previous layer into the multiple connection weights of one neuron corresponding to the current computation circuit; Step 4, the multiplicative sum is calculated and multiplied by the learning rate, the product being the adjustment value of the corresponding connection weight; Step 5, the corresponding connection weight is updated.
[0029] This invention discloses neurons sharing a common computational core and a common memory; the layer activation function source writes activation function information into the common memory; the computational core obtains activation function values from the common memory through table lookup, copying, and calculation; a method for synthesizing multiple rounds of output values is provided: Step 1, perform forward propagation to obtain the level signal expressing the equivalent time difference of capacitor charging and discharging and the level signal expressing the sign of the energy relationship of each capacitor; Step 2, update the parameters of the layer activation function value source and continuously output the time value, the derivative change value of the activation function over time, and the current value to the neuron output module, or output only the time value and obtain other values through table lookup; Step 3, if the activation function needs to be used, obtain the activation function value of each neuron through interpolation calculation of the derivative change value and the current value; Step 4, if it is necessary to add layer output values, do not reset the output item fast access unit or the activation function fast access unit, and add the current activation function value with the previous round activation function value, repeating steps 1, 2, 3, and 4; Step 5, obtain the final neuron output value.
[0030] This invention discloses a backup fast access unit for the weight term current modulation control module; the output value of the neuron output module is preloaded into the backup fast access unit in blocks or batches; the backup fast access unit has a switching circuit that can switch the effective weight values in the weight term current modulation control module; the pre-generated output value of the neuron output module is directly provided to the input term current modulation control module as a real-time input value; the pre-loaded weight values in this layer of the neural network are enabled and re-output for matrix multiplication operations.
[0031] This invention discloses a method for normalizing layer output values: Step 1, the network summarizing layer neurons use the values of the activation function fast access units of the current main computing layer as input to the network summarizing layer neurons; Step 2, the network summarizing layer neurons calculate and generate layer parameter values, i.e., the process of powering on, detecting, and activating; Step 3, the layer parameter values are given to the layer activation function value calculation chip, and the layer activation function value source regenerates normalized activation function value information with equivalent time as the independent variable based on the layer parameter values; Step 4, the activation function value information with equivalent time as the independent variable is sent to the neuron output module of the current main computing layer; Step 5, the current main computing layer regenerates the normalized activation function value based on the original activation function value and stores it in the activation function fast access unit.
[0032] This invention discloses a method for implementing a gated network: Step 1, the output value or activation function value of the first forward propagation is written into the fast access unit of the weight term of the linear weight term current modulation control module; Step 2, the output value or activation function value of the second forward propagation is written into the fast access unit of the input term current modulation control module; Step 3, based on the values of the fast access unit of the input term and the fast access unit of the linear weight term, a modulation current is generated to charge and discharge the capacitor module, that is, the multiplication of the two is realized through the charging and discharging of the capacitor. Step 5: The capacitor module is detected to charge and discharge through the capacitor control module; Step 6: The voltage of the capacitor module is compared with the reference voltage, and a detection signal is output; Step 7: The positive and negative charge and discharge detection signals of the capacitor control module are combined to obtain the level signal or activation function information expressing the equivalent time difference of capacitor charge and discharge.
[0033] This invention discloses a method for generating comparison detection signals P and N by comparing the capacitor voltage of the capacitor module with a reference voltage; the neuron output module logically synthesizes the comparison detection signals P and N, and generates an XOR gate signal based on the comparison detection signals P and N, i.e., whether the comparison detection signals P and N are at the same level; the neuron output module generates a positive or negative signal SGN when the XOR gate signal starts, i.e., when P and N become unequal; if the comparison detection signal P has a longer duration, SGN is positive, otherwise it is negative.
[0034] This invention discloses multiple sets of phase-staggered, non-overlapping underlying reference pulse signals; the input current modulation control module of each neuron generates multiple sets of current modulation signals based on the phase-staggered underlying reference pulse signals and its own input value; the switching of the link current channel of the neuron generates multiple sets of modulation currents based on the multiple sets of current modulation signals.
[0035] This invention discloses a method with two-layer dynamic alternation: Step 1, the basic neuron unit loads the values of the weight term fast access unit and the input term fast access unit from the memory or external circuit; Step 2, the dynamically alternating two-layer neural network, with one layer as the input layer and the other as the output layer, performs forward propagation calculations; Step 3, if the entire neural network has reached the calculation of the last layer, the activation function fast access unit value is output to the memory or external circuit; Step 4, if the entire neural network has not reached the calculation of the last layer, the basic neuron unit loads the values of the weight term fast access unit of the new layer from the memory or external circuit; Step 5, the original input layer of the dynamically alternating two-layer neural network is used as the output layer, and the original output layer is used as the input layer, performing forward propagation calculations; Step 6, steps 3, 4, and 5 are repeated cyclically.
[0036] This invention discloses an equivalent resistance network that generates a specified equivalent resistance value based on the value of the fast access unit of the weight term, thereby adjusting the instantaneous current of the corresponding current channel. This invention also discloses an input current modulation control module that generates a single pulse level based on the value of the fast access unit, used to generate activation function values. Furthermore, this invention discloses that each basic neuron unit has a bias link, which is collectively linked to a bias link input current modulation control module; the modulation current ratio of the bias link input current modulation control module is 100%; the modulation current ratio of the bias link is controlled by the fast access unit of the weight term of the bias link weight term current modulation control module.
[0037] This invention discloses a detection process, which includes a comparator for the positive and negative capacitor voltages, the result of which generates an enable signal to control the operating state of other circuits.
[0038] The present invention also discloses that the capacitor control module has a comparator that uses a piecewise function to divide the reference voltage at the point. The comparison result divides the discharge detection process, and then generates pulses corresponding to different activation functions and finally outputs them.
[0039] This invention discloses a backup fast access unit for the weighted term current modulation control module; based on a fully connected neural network, a method for generating a self-attention QK matrix and performing multiplication is provided: Step 1, read the Q matrix parameters from an external source or storage to the fast access unit of the input term current modulation control module; Step 2, in parallel, read the values of the input vector to the fast access unit of the weighted term current modulation control module; Step 3, perform forward propagation calculation, and the neuron output module generates a round of Q matrix result values; Step 4: Copy the result value from the neuron output module to the spare fast access unit of the corresponding weight term current modulation control module; Step 5: Repeat steps 2, 3, and 4 until the spare fast access units of each weight term current modulation control module are filled or all Q matrix parameters are used; Step 6: Read the K matrix parameters from external storage or the spare fast access unit to the fast access unit of the input term current modulation control module; Step 7: In parallel, read the input vector values to the fast access units of the weight term current modulation control modules; Step 8: Perform forward propagation calculation, and the neuron output module generates one round of K matrix result values; Step 9: The neuron output module transfers the K matrix result values to the fast access units of the input term current modulation control modules; Step 10: The main and spare fast access units of each weight term current modulation control module are exchanged, i.e., the result values of the Q matrix are used. Step 11: Perform forward propagation calculation. The neuron output module generates a round of QK matrix multiplication result value and saves the round of QK matrix result value to external storage or memory, or other idle and readily accessible units. Step 12: Repeat steps 6, 7, 8, 9, 10, and 11 until all K matrices have been calculated. Based on the calculation results of the previous steps, a method for generating self-attention QKV results is as follows: Step 13: The QK matrix multiplication result is normalized and activated, then stored in other idle spare fast access units of the weight term current modulation control module. This process is repeated multiple times until all QK activation result values are obtained. Step 14: The V matrix parameters are read from external storage, memory, or spare fast access units and stored in the fast access unit of the weight term current modulation control module. Step 15: In parallel, the QK matrix multiplication result value of one round, i.e., one round of attention weights, is read and stored in the fast access unit of the input term current modulation control module. Step 16: Forward propagation calculation and activation are performed. The neuron output module generates one round of QKV matrix multiplication result value and saves it to external storage or memory. Step 17: Steps 14, 15, and 16 are repeated until all QKV matrices have been calculated. Step 18: If the input vector and layer input are inconsistent, the input vector is segmented and grouped. The final result is obtained through a combination of multiple groups and multiple rounds of calculation. If position encoding is required in the above steps, the data can be modified externally to insert the position encoding, or the vector after inserting the position encoding can be generated by setting the bias term weight parameters of the weight term current modulation control module.
[0040] This invention discloses a method for network propagation calculation using single-pulse pulse width: Step 1, the weight term current modulation control module reads a preset weight value and sets the equivalent resistance of the neuron link current channel; Step 2, the input term current modulation control modules each receive a single-pulse signal and symbol, and control the switching of the neuron link current channel to charge and discharge the pre-charged capacitor module accordingly; Step 3, the positive and negative capacitor control module controls the charging and discharging of the equivalent detection resistor and generates positive and negative detection pulse signals; Step 4, the neuron output module generates a single-pulse output and a symbol output based on the positive and negative detection pulse signals from the positive and negative capacitor control module; the output single-pulse pulse width is proportional to the sum of the product of the input single-pulse signal pulse width and the equivalent resistance of the current channel; Step 5, the output single pulse is charged and discharged twice according to the settings or execution, and a new output pulse is regenerated. Attached Figure Description
[0041] Figure 1-9 These are all circuit schematics. Figure 10 This is a diagram illustrating the working process. Figure 11-12 It is a circuit schematic. Figure 13-14 This is a logic level diagram. Figure 15-17 It is a circuit schematic. Figure 18 It is a diagram of a mathematical function. Figure 19-21 This is a diagram illustrating the working process. Figures 22a-22c It is a circuit schematic. Figure 22d It is a diagram of a mathematical function. Figures 23a-23b It is a circuit schematic. Figure 24 This is a functional layout diagram. Figure 25 This is a data flow diagram. Figure 26 It is a circuit schematic. Detailed Implementation
[0042] Unless otherwise specified, the following embodiments are basically parallel operations, that is, a large number of basic neuron units are computed in parallel, and a large number of links within a single neuron unit are also computed in parallel.
[0043] Simplest embodiment* such as Figure 1 It is a hardware implementation of a neural network with two input units in the first layer and only one neuron in the last layer. It includes two network link input current modulation control modules (21)(19), weight current modulation control modules (5)(6)(17)(18), capacitor module (7)(22), capacitor control module (8)(16), and neuron output module (14).
[0044] The network link input current modulation control module and the weighted current modulation control module jointly control the charging and discharging current of the capacitor module. In this embodiment, the network link input current modulation control module is a circuit module that dynamically controls the current and is used to generate the modulation current. (Known forms of modulation current include, but are not limited to: pulse width modulation (PWM), pulse density modulation (PDM), pulse position modulation (PPM), single pulse level, etc.; the modulation current is converted into charging and discharging current by the control signal through the switching circuit of the current channel, and its implementation includes direct clock generation, multi-channel signal mixing, reference signal trimming or delay insertion, etc. The current waveform also includes, but is not limited to, centrally symmetrical or edge-aligned, rectangular wave, triangular wave or sine wave, etc. The modulation current in other embodiments is the same.) The modulation current plus the switching of the internal resistance of the weighted current modulation control module (5)(6)(17)(18) realizes the control of the current. In this embodiment, the weighted current modulation control module is a small uniform resistor. In other embodiments, it can be a complex resistor network or an equivalent circuit. In this embodiment, the effect of the current flowing into the capacitor controlled by the network link input current modulation control module and the weight current modulation control module is equivalent to the link parameter w * input value x in the neural network. The network link input current modulation control module contains a fast access unit (the form of the fast access unit includes, but is not limited to: register, register file, flip-flop, latch, various SRAM / DRAM cells, array, and various new storage technology access units. Other embodiments are the same), which is set by the control information channel (1) (the control information channel and related fast access units also include functions such as enable, capacitor link port high impedance control, positive and negative setting of link parameter w and input value x, etc.), and generates current modulation (PWM, etc.) by comparing with the timing information of the timing information channel (2). Here, a part (bits) of the current modulation fast access unit represents the network input value x, and a part represents the link parameter w (which works together with the weight current modulation control module). The driving voltage channel (3) is linked to a dynamically selectable capacitor charging and discharging driving voltage source shared by the entire neural network layer. For example, the driving voltage a = 0 volts during discharge and the driving voltage a = 2 volts during charging. The specific voltage value depends on the actual needs; this is just for ease of description. The above control information channels, timing information channels, and even resistance / current channels are not limited to a single physical power-on link. The specific number of physical power-on links is determined according to the complexity of the data and actual needs.
[0045] (As is well known, a timer has multiple channels, and each channel has a fast access unit. Comparing the current with the timer's fast access unit generates multiple corresponding current modulations, which is a basic function of a microcontroller. In addition to being driven by a crystal oscillator clock, a timer can also be implemented using a delay chain (such as an inverter).) In this embodiment, the current is basically controlled by current modulation and a standard resistor network. In addition to expressing x, current modulation is also used to express w. Sometimes, in order to save chip area, some embodiments (specific value resistor embodiments) use different specific resistor values to express the link parameter w, or even use a specifically generated stable voltage source to replace current modulation.
[0046] The capacitor modules of the last layer neurons are divided into positive and negative capacitor modules. In this embodiment, the left capacitor module represents a positive value (7), and the right capacitor module represents a negative value (22). The network link input current modulation control module will select to charge or discharge the positive or negative capacitor module according to the sign bit of the fast access unit set by the internal link parameter w and the sign bit of the fast access unit set by the input value x. That is, it determines which capacitor module to charge or discharge according to the sign of (w*x).
[0047] The capacitor control modules (8) and (16) correspond to their respective capacitor modules. During the charging and discharging process of the capacitor modules, the capacitor modules are pre-charged and discharged through the initial voltage source channel (9). In the charging mode, the capacitor modules are pre-charged to the initial voltage (in this embodiment, the initial voltage b = 1 volt); in the discharging mode, the capacitor modules are pre-charged to the initial voltage (in this embodiment, the initial voltage b = 2 volts). In the current embodiment, the initial voltage source can be set to an initial voltage b = 1 volt or b = 2 volts through the layer parameters of the neural network (1 or 2 volts is just an example reference, and the specific voltage value should be set according to the actual situation). If the neural network is more complex and there are different activation functions in one layer, or other more complex requirements, then multiple initial voltage source channels (9) with different voltages are needed. After the pre-charging and discharging steps, the input current modulation control module and the weight current modulation control module are linked through the network, and the neural network is charged and discharged, which is the actual working process of the neural network. After the neural network charging and discharging steps, the capacitor discharge detection begins. The capacitor module discharges to GND through the standard resistor R inside the capacitor control module until the detection voltage c = 1 volt (the specific voltage value is set according to the actual situation; in this embodiment, it is 1 volt). The capacitor control module contains a voltage comparator. When the voltage of the discharged capacitor module is lower than the detection voltage c, the output port (15) of the capacitor control module flips, displaying a level representing the comparison result (the level is set as needed; in this embodiment, it is set to high). All operations of the capacitor control module are manipulated by its own capacitor control channel (23). (Voltage comparison is a basic function of the chip, and its circuitry is common knowledge.) The neuron output module (14) logically merges the output signals of the capacitor control modules representing positive and negative values (here, the symbol P / N represents the output of the positive and negative capacitor control modules, i.e., the positive module output is P and the negative module output is N). The layer output (13) of the neuron output module (14) represents the information of the duration difference of the discharge. In this embodiment, its information is logically equal to the level duration of P xor N. The layer output (12) of the neuron output module (14) represents the comparison of the magnitude of the positive and negative capacitor voltages or the comparison of the magnitude of the positive and negative capacitor discharge durations. In this embodiment, it is logically equal to the level of P output by the positive module before P xor N flips last. (It is known that logical operations such as xor are basic functions of digital circuits.) Note that in this embodiment, what needs to be output is the duration difference information, not just the duration difference level itself. The duration difference information has multiple forms of expression. Although this embodiment outputs the duration difference level for processing by the microcontroller and other modules, some embodiments (such as the current modulation layer output embodiment) require converting the duration difference information into a current modulation signal for output. That is, the neuron output module (14) has a time-to-digital converter circuit for detecting the duration difference. The layer output control signal channel (11) controls the time-to-digital converter circuit of the neuron output module to capture the Pxor N and obtain its duration through the timing information input through the timing information channel (10) and store it in the internal fast access unit. Finally, based on the value of the internal fast access unit and the timing information input through the timing information channel (10), the current modulation signal is output. (It is known that detecting the square wave pulse width is a basic function of the microcontroller. At low frequencies, a counter can be used to count clock pulses, and at high frequencies, a delay chain signal / multi-phase multi-line signal is used.) The following are the steps of the working method of the analog-digital hybrid neural network circuit in this embodiment (some steps can be performed in parallel according to actual needs, and the specific voltage values are set according to actual needs): 1. Initialization settings: Set the fast access unit values of the network link input current modulation control module, including enable, high impedance control of capacitor link port, positive and negative settings of link parameter w and input value x, etc.
[0048] 2. Select appropriate resistance values or resistor networks as needed.
[0049] 3. Generate an appropriate current modulation signal.
[0050] 4. Capacitor pre-charge and discharge: Use the initial voltage source channel (9) to pre-charge and discharge the capacitor module (7)(22) to a specific initial voltage (charging mode b=1 volt, discharging mode b=2 volts). Pre-charge and discharge can be achieved by directly connecting to the initial voltage source or by detection through the voltage comparator (of the capacitor control module).
[0051] 5. Perform charging and discharging. Based on the positive and negative values of w and x set internally, select to perform charging and discharging operations on the positive capacitor module (7) and the negative capacitor module (22).
[0052] 6. The charging and discharging current of the capacitor module (7)(22) is jointly controlled by the network-linked input current modulation control module (21)(19) and the weighted current modulation control module (4)(20)(5)(6)(17)(18).
[0053] 7. After the charging and discharging process, during detection, the charging and discharging continues until the detection voltage is reached. (Regarding energy efficiency, the detection process can be further optimized by setting multiple reference comparison voltages and selecting different resistance values for the detection discharge resistor based on the voltage range of the positive and negative capacitors during detection, thereby accelerating the detection discharge speed. Because the result is a difference and the positive and negative values switch synchronously, it does not affect the final result.) 8. Output detection result information / signal.
[0054] 9. Layer output processing: The neuron output module (14) combines the output signals (P and N) of the positive and negative capacitance control module and uses the time difference information to represent the final output. The time difference information and sign are the level and signal pulse, or the signed digital value after further time-to-digital conversion.
[0055] 10. In some embodiments, it may be necessary to convert this duration difference into a current-modulated output or other form of electrical information, which involves capturing the level duration of P xor N and outputting a current-modulated signal based on this information.
[0056] As is well known, the conversion from time pulse to digital quantity described in step 9 can be achieved through a time-to-digital converter (TDC). Specific implementations include, but are not limited to, counters, interpolators, inverter delay chains, vernier signal structures, time amplifier circuits, etc. In the case of a multi-layer neural network, if the input current modulation uses a single-pulse level, then the time pulse can be directly used as the input to the next layer of the neural network, or the entire network can use single pulses. Note that mathematically, it can be verified that if the PWM duty cycle variable in the derivation below is replaced with the pulse width of a single pulse, i.e., the pulse duration, the resulting formula, even in the counterintuitive case where the single pulses are not aligned in time, still holds true. That is, the output time difference is proportional to the product of the total pulse width and conductance of each input (the time integral of the total conductance), which still holds true due to the exponential multiplication effect. The capacitor discharge process is an exponential multiplication, with the discharge voltage V(t) = V0 * e^(-t / (RC)). e^a * e^b = e^(a+b), which is independent of the order, length, or even overlap of the single pulses corresponding to a or b. (That is, it only depends on the time integral of the total positive / negative conductance; the derivation is omitted). However, appropriate pulse repetition in PWM can average out various uncontrollable factors such as interference and inconsistencies, although it also increases power consumption. Single pulses, PWM, capacitor charge transfer, or other more complex time- and conductance-based similar forms are essentially all about the ratio of conduction time and current amplitude between various links and detection channels. The final result can be obtained by applying similar formulas to achieve the proportional relationship.
[0057] The following is a computational description of the working principle of the aforementioned hardware, and an explanation of how to generate the connection parameters w and input value x of the hardware. The calculations below are based on ideal components, neglecting leakage current, voltage and temperature variations, and other interference factors. Therefore, the calculations are approximate results. In actual neural network training, if a test statistical characteristic table of the hardware circuit is needed, the software layer will use a lookup table and interpolation method to obtain the actual input / weight / output values / derivatives of the circuit. The following calculations use the simplest PWM current modulation form as an example, but its essence is to control the current ratio of each current channel through switching; therefore, other forms of modulation are equivalent. Furthermore, in the following calculations, the pulse width, resistance, and capacitance do not require precise values; what is needed are precise and stable ratios.
[0058] The value x represents the duty cycle of the input portion in current modulation (PWM pulse width modulation; in single-pulse modulation, it's the pulse duration width, i.e., the ratio to the minimum pulse duration width). The duty cycle affects the voltage difference between the equivalent input voltage source and the capacitor voltage. Current input x = Xn = pulseXn The value w is the duty cycle of the portion representing the equivalent resistance in current modulation (pulseRn; in single-pulse modulation, this can be omitted and set to 1; if this weighted item exists in single-pulse modulation, it represents the scaling factor of the pulse duration or current intensity), divided by the corresponding resistance Rn and capacitance cap. Current weighted current limiting w = Wn = pulseRn / Rn / cap SSP(...), SumSelectPositive means selecting all values greater than or equal to 0, summing them, and taking the absolute value.
[0059] SSN(...), SumSelectNegative means selecting all values less than 0, summing them, and taking the absolute value.
[0060] V[t] represents the time function of the capacitor voltage, V'[t] is its derivative, and e is the natural constant. 1.1. In discharge mode, the initial voltage is 'a', and the discharge drive voltage is GND voltage 0. After the neural network discharges, it enters the detection process to continue discharging, with the stop / detection voltage c=b=1. The discharge resistor used in the detection process is R. t is a time variable, which is a standard time length that can be set and controlled.
[0061] Solve the difference equations respectively (b <= V[t] <= a). Dsolve[{V'[t]==SSP(...,Xn*(-V[t])*Wn), V[0]==a}, {V[t]}, t] DSolve[{V'[t]==SSN(...,Xn*(-V[t])*Wn), V[0]==a}, {V[t]}, t] Solution results The positive capacitance V[t] is given by gp1 = a * e^(-t*SSP(...,Xn*Wn)). The negative capacitance V[t] is given by gn1 = a * e^(-t*SSN(...,Xn*Wn)). During detection, the positive capacitor discharge time function tp1 = cap*R*ln[gp1 / b]=cap*R*(ln[gp1]-ln[b]) During detection, the discharge time function of the negative capacitor is tn1 = cap*R*ln[gn1 / b]=cap*R*(ln[gn1]-ln[b]). The difference between the two, tp1-tn1, is calculated according to the logarithmic rule. ln[a * e^(-t*SSP(...,Xn*Wn))]=ln[a]+ln[e^(-t*SSP(...,Xn*Wn)] = ln[a]-t*SSP(...,Xn*Wn) ln[a * e^(-t*SSN(...,Xn*Wn))]=ln[a]+ln[e^(-t*SSN(...,Xn*Wn)] = ln[a]-t*SSN(...,Xn*Wn) tp1 = cap*R*(ln[a]-t*SSP(...,Xn*Wn)-ln[b]) tn1 = cap*R*(ln[a]-t*SSN(...,Xn*Wn)-ln[b]) tp1-tn1 = -cap*R*(X1*W1+X2*W2+...+Xn*Wn)*t As can be seen from tp1 and tn1, its activation function is a linear function. Substituting this into Wn = pulseRn / Rn / cap... tp1-tn1 = -R*(X1*pulseR1 / R1+X2*pulseR2 / R2+...+Xn*pulseRn / Rn)*t Here, R and R1...Rn are all proportional. The value of pulseR1*R / R1 is the value of the weighting term. If all resistors are standard resistors, that is, R1...Rn equals R, then... tp1-tn1 = -(X1*pulseR1+X2*pulseR2+...+Xn*pulseRn)*t If we flip the positive and negative signals of the neuron's output module, then the result of tp1-tn1 is... R*(X1*pulseR1 / R1+X2*pulseR2 / R2+...+Xn*pulseRn / Rn)*t (X1*pulseR1+X2*pulseR2+...+Xn*pulseRn)*t, In simple terms, if the modulated current is generated using PWM, the total duty cycle is duty_n = pulseXn * pulseRn. The input value Xn = t * pulseXn, representing the duty cycle of the input portion. The weight value Qn = pulseRn * R / Rn, where pulseRn is the duty cycle representing the weight portion of the neuron's connection, Rn is the resistance of the weight term, and R is the standard resistor for detecting discharge. In essence, the calculation is achieved by adjusting the ratio of the standard resistor to the weight term resistance, the duty cycle of the weight term and the input term, and the total discharge time to make them equivalent to the input and weight values in the software.
[0062] Additionally, there is a time t. If t is not the base time, assuming t=8, the final discharge detection time or input value should be subtracted from the time t. Since t=8 is a multiple of 2, the discharge detection time or input value can be directly shifted by 3 bits.
[0063] If there are deviations in the production process, resulting in one capacitor having a larger capacitance than the other (cap1 and cap2), then the above formula becomes... tp1 = cap1*R*(ln[a]-t*SSP(...,Xn*Wn)-ln[b]) tn1 = cap2*R*(ln[a]-t*SSN(...,Xn*Wn)-ln[b]) tp1-tn1 =R*(cap1-cap2)*ln(a / b)-R*(X1*pulseR1 / R1+X2*pulseR2 / R2+...+Xn*pulseRn / Rn)*t The calculation error caused by the capacitance difference: err_cap = R * (cap1 - cap2) * ln(a / b) Because the capacitance difference is an independent term, its error value can be easily obtained through calculations with a 0% duty cycle. For digital circuits (with digital-to-analog conversion), this error can be subtracted from the final calculation result. For analog circuits (without analog-to-analog conversion), the error can be corrected by adjusting the neuron bias term connections or by setting additional bias term connections.
[0064] The standard resistor R and other resistors R1, R2...Rn may have manufacturing errors. However, we only use the resistance ratio R / Rn here, which is relatively accurate. In addition, manufacturing errors, operating temperature, voltage, line inductance, frequency, leakage current, and switching delay will all affect the calculation results. Various errors can be adjusted using appropriate fine-tuning circuits or by inserting a delay into the PWM. These adjustment circuits and how to reduce manufacturing errors are very specific design tasks, which will not be discussed in detail here. A relatively simple method for adjusting the weights is to obtain the average deviation of the weights for each link through multiple calculations based on specific parameters (such as 0), and then correct the final result by modifying the ideal weight values (before using the weights).
[0065] 1.2. If the detection process begins, and the device recharges from gn1 or gp1 to a=2, the stop / detection voltage is b=a. The discharge resistor used in the detection process is R. The charging drive voltage is c=a+1, then... The positive capacitor discharge time function during detection is tp3 = cap*R*ln[(c-gp1) / (ca)] ; tp3 = cap*R*ln[3 - 2 * e^(-t*SSP(...,Xn*Wn))] The discharge time function of the negative capacitor during detection is tn3 = cap*R*ln[(c-gn1) / (ca)] ; tn3 = cap*R*ln[3 - 2 * e^(-t*SSN(...,Xn*Wn))] It is evident that tp3 and tn3 are some kind of nonlinear functions. 2.1. Now returning to the charging mode in the simplest embodiment, the initial voltage is b=1, the charging drive voltage is a=2, and after the neural network charging is complete, it enters the detection process to begin discharging to ground, with the stop / detection voltage c=b=1. The discharge resistor used in the detection process is Rt, which is a time variable. Mathematical software solves the difference equations (b <= V[t] <= a). DSolve[{V'[t]==SSP(...,Xn*(aV[t])*Wn), V[0]==b}, {V[t]}, t] DSolve[{V'[t]==SSN(...,Xn*(aV[t])*Wn), V[0]==b}, {V[t]}, t] Solution results For a positive capacitor, V[t] is given by gp2 = a+(ba) * e^(-t*SSP(...,Xn*Wn)); for a negative capacitor, V[t] is given by gn2 = a+(ba) * e^(-t*SSN(...,Xn*Wn)). The discharge time function for the positive capacitor during detection is tp2 = cap*R*ln[gp2 / b]; the discharge time function for the negative capacitor during detection is tn2 = cap*R*ln[gn2 / b]. The difference between the two, tp2-tn2, can be substituted into b=1 and a=2. tp2 = cap*R*(ln[2-e^(-t*SSP(...,Xn*Wn))]); tn2 = cap*R*(ln[2-e^(-t*SSN(...,Xn*Wn))]) tp2-tn2 = cap*R*(ln[2-e^(-t*SSP(...,Xn*Wn))] - ln[2-e^(-t*SSN(...,Xn*Wn))]) As can be seen from tp2 and tn2, their activation function is the difference between two nonlinear functions tp2 and tn2. The curves of tp2 and tn2 in the first quadrant are similar in shape but have better training performance than the cap*R*ln(2)*tanh function. tanh is a mature neural network activation function with good performance. The maximum values of tp2 and tn2 are cap*R*ln(2).
[0066] In other words, when nonlinear activation is required in this embodiment, the activation function used is to group the Xn*Wn groups according to their positive and negative signs, sum the sums of each group, take the absolute value (val), and then perform a nonlinear transformation of cap*R*ln(2-e^(-t*val)), and then directly take the value or take the difference between the two groups after the nonlinear transformation.
[0067] 2.2. If the detection process begins, the capacitor is recharged from gn2 or gp2 to a=2, and the stop / detection voltage a is reached. The charging resistor used during the detection charging process is R. The charging drive voltage is c=a+k, k=1. Then, the positive capacitor charging time function during detection is tp4 = cap*R*ln[(c-gp2) / (ca)] = cap*R*ln[a+k - (a+(ba) * e^(-t*SSP(...,Xn*Wn)))] = cap*R*ln[1 + (ba) / k * e^(-t*SSP(...,Xn*Wn))] = cap*R*ln[1 + e^(-t*SSP(...,Xn*Wn))] The charging time function of the negative capacitor during detection is tn4 = cap*R*ln[(c-gn2) / (ca)] = cap*R*ln[ a+k - (a+(ba) * e^(-t*SSN(...,Xn*Wn)))]= cap*R*ln[ 1 + (ba) / k * e^(-t*SSN(...,Xn*Wn))] = cap*R*ln[ 1 + e^(-t*SSN(...,Xn*Wn))] As can be seen from tp4 and tn4, their activation function is the difference between the same two nonlinear functions tp4 and tn4. 3.0. Example of Positive and Negative Capacitor Detection and Comparison* Similarly, the capacitor is first pre-charged to a predetermined voltage 'a', then the network discharges for forward propagation of the neural network, and then the detection process is executed. The voltage comparator directly compares the voltages of the positive and negative capacitors. The capacitor with the larger voltage discharges to ground through a standard resistor R during the detection process until the voltages of the two capacitors are equal. The discharge time is the time difference. Its sign represents the result of comparing the magnitudes of the positive and negative capacitors at the start of detection. Its unsigned value is consistent with |tp1-tn1|, and is also a linear function. However, because the voltage comparator has an offset voltage, and various leakage currents may be inconsistent, this will affect the accuracy and efficiency of the calculation results.
[0068] 3.1. Example of Dual-Capacitor Detection and Comparison* (Each capacitor has positive and negative terminals, i.e., 4 capacitors) Here, we supplement another example of detection and discharge comparison (dual-capacitor detection and comparison), where an identical capacitor is added to the capacitor module. This capacitor does not participate in the charging and discharging of the neural network; it only plays a role during detection (the network discharges, and detection also discharges). Before detection, this capacitor is pre-charged to 'a', and then it discharges to ground (0 voltage) through 'R'. The stop / detection voltage is set to the voltages of the capacitors in the capacitor module that participate in the charging and discharging of the neural network, i.e., gp and gn, i.e., the voltages of two capacitors with the same sign are compared using analog voltage comparison. During detection, the positive capacitor discharge time function tp0 = cap*R*ln[a / gp] = cap*R*(t*SSP(...,Xn*Wn)) During detection, the discharge time function of the negative capacitor is tn0 = cap*R*ln[a / gn] = cap*R*(t*SSN(...,Xn*Wn)). diff0 = tp0 - tn0 = cap*R*(X1*W1+X2*W2+...+Xn*Wn)*t; From tp0 and tn0, it can be seen that its activation function is also a linear function. 3.2 In the dual-capacitor detection process (where the network is charging and detection is also charging), the capacitor is pre-charged to b=1V before detection, and then charged through R with a driving voltage of a=2V (the specific voltage value depends on the actual setting; this is only for reference). The stop / detection voltage is set to the voltages of the capacitors participating in the neural network charging and discharging in the capacitor module, i.e., gp2 and gn2, that is, the voltages of two capacitors with the same sign are compared using analog voltage comparison. During testing, the positive capacitor charging time function is tp5 = cap*R*ln[(a-gp2) / (ab)]. = cap*R*ln[(a-(a+(ba) * e^(-t*SSP(...,Xn*Wn)))) / (ab)] =cap*R*ln[e^(-t*SSP(...,Xn*Wn))] The charging time function of the negative capacitor during detection is tn5 = cap*R*ln[(a-gn2) / (ab)]. = cap*R*ln[(a-(a+(ba) * e^(-t*SSN(...,Xn*Wn)))) / (ab)] = cap*R*ln[e^(-t*SSN(...,Xn*Wn))] tp5-tn5 = -cap*R*(X1*W1+X2*W2+...+Xn*Wn)*t As can be seen from TP5 and TN5, their activation functions are also linear functions. The optimal method for training model parameters is to calculate them using the statistical characteristics of the circuit through table lookup and interpolation. After obtaining the final model parameters through backpropagation training, these parameters are then input into the circuit of this embodiment to achieve neural network output (for inference). For parameters exceeding the range in the software, they can be divided into multiple parameters and multiple inputs, or the hardware circuit can be improved to allow the weighted current modulation control module to select more different resistance values.
[0069] In the computational part, as summarized above, tp5-tn5, tn1-tp1, tp0-tn0, etc., are ideally equivalent to the summation of product terms commonly used in software neural networks (the constant term is simply setting the input of one of the product terms to 1). tp4-tn4, tp3-tn3, tp2-tn2, etc., can also be used in neural networks, but there has been no attempt or publicly available information from the software and academic communities.
[0070] Looking at the derivation process of tp5-tn5, tn1-tp1, tp0-tn0, for example, tn1, tp1, and tn1-tp1, the calculation results are independent of the capacitance value of the charging and discharging capacitors. This greatly facilitates the design and production of chips or circuits, because the capacitance value is difficult to determine precisely. Our invention's calculation results do not depend on the capacitance value, which brings a huge advantage to mass production applications. In addition, because tn1-tp1 calculates the difference, the errors in tn1 and tp1 caused by the leakage current of the positive and negative capacitor control module and the capacitor module can be mutually canceled to a certain extent (of course, the leakage current should be minimized in the design). The calculation errors caused by the resistance value deviation and the errors caused by the current modulation time deviation are relatively small, and their accuracy is high. The capacitor and resistance deviations (only related to the ratio) can be corrected by obtaining consistency parameters during testing and written into the internal memory of the neuron output module. The neuron output module automatically adjusts the XOR time difference when it is working.
[0071] As can be seen from tp5-tn5 and tp0-tn0, although the dual capacitors use twice the amount of capacitors, they can still perform the calculation of accumulating positive and negative product terms whether charging or discharging.
[0072] This embodiment implements the essential multiply-accumulate, linear activation, and nonlinear activation functions required by neural networks. Other activation functions are not strictly necessary for neural networks and can be subsequently handled by computing chips / GPUs / microcontrollers, or see more embodiments below. The multi-channel multiply-accumulate parallel circuit described above reduces the number of CMOS transistors used by at least an order of magnitude compared to a single floating-point multiplier in existing chips.
[0073] In some embodiments, grouping operations are not desired (in embodiments with single-group nonlinear activation*). Instead, it is preferable to use the traditional multiply-add form w1x1 + w2x2... followed by nonlinear transformation activation. This is consistent with the implementation of traditional neural networks. While there are hardware workarounds in this invention, they will increase system latency. Details are as follows: The layer network uses a discharge mode. Each neuron output module has an additional equivalent conversion capacitor, pre-charged to c=1 volt. During the effective level of the P xor N output (as explained above, this duration is linear), the conversion capacitor is charged with a driving voltage a=2 through a resistor R (the resistor value here is set to the same value as the detection discharge resistor). Then, similar to the detection discharge process above, the charged conversion capacitor is discharged, also through the equivalent resistor R to ground. The discharge stops when the conversion capacitor voltage is less than or equal to the detection voltage b=c=1 volt. Similarly, the duration of the conversion capacitor detection discharge is a nonlinear transformation function with a shape similar to cap*R*ln(2)*tanh first quadrant curve. That is, f(y) = cap * R * (ln[2 - e^(-y)]), (y = tp0 - tn0 or y = tp1 - tn1). In effect, it's equivalent to adding a non-linear layer with one-to-one neuron connections on top of a linear activation layer.
[0074] Some embodiments may require implementing a nonlinear activation function (ADC detection type embodiment*) in the capacitor control module (8)(16) by using a high-speed ADC to read the capacitor voltage into a fast access unit, where a=2, b=1. The positive capacitance V[t] is given by: gp2 = a + (ba) * e^(-t*SSP(...,Xn*Wn)) = 2 - e^(-t*SSP(...,Xn*Wn)). The negative capacitance V[t] is given by: gn2 = a+(ba) * e^(-t*SSN(...,Xn*Wn)) = 2-e^(-t*SSN(...,Xn*Wn)). Using op-amps to perform subtraction (gp2-1, gn2-1), and then importing the result into an ADC, the voltage result from the ADC is stored in a fast access unit (the value in the fast access unit). Similarly, a nonlinear function with a zero-crossing extreme of 1 and a tanh-like shape can be obtained in the first quadrant. That is 1-e^(-y), y=t*(SSP or SSN)(...,Xn*Wn) Example of a positive and negative capacitor linkage switch* Figure 9 The current modulation control unit (80) of the neuron selects the specific charging capacitor through the positive and negative capacitor selection switch (81). When the number of neurons is small, the capacitor selection switch can be placed between the capacitor module (83) and the weight term current modulation control module (82). The resistor network has a control channel (84) to control and manipulate its resistance value, which, together with the PWM input term current modulation control module at the other end, controls the current flowing from the input unit to the capacitor module of the next layer of neuron unit.
[0075] To achieve precise current control and finer precision, a PWM switching micro-delay module can be added. This module contains various semiconductor component combinations and controls the micro-delay before and after the PWM switching action via registers, thus enabling more accurate current control. The switching delay time of the PWM switching micro-delay module is often implemented internally by hardware, without needing to consider the base frequency timing signal. Generating more precise PWM through micro-delay or other methods is a common practice in the industry.
[0076] (If there are many neurons, the capacitor selection switch can be combined and placed at the other end of the resistor network. That is, as in the previous embodiment, the resistor network is also divided into positive and negative, and the positive and negative capacitors are each connected to the corresponding input term's positive and negative resistor network. This reduces the number of positive and negative capacitor selection switches, but the resistor network is doubled because it is divided into positive and negative.) The simplest embodiment of a single-layer neural network with 4 inputs and 2 outputs* is a simple extension of the simplest embodiment above. Its working principle is the same as the simplest embodiment above, but it has more inputs and outputs. Figure 2 The neural network has only one layer; it has four inputs, resulting in four input current modulation control modules (30) and four corresponding weighted current modulation control modules (31); it has two outputs (32), with two groups of weighted current modulation control modules, each group containing two resistors (positive and negative), for a total of 2*2*4=16 resistors. Each of the two outputs contains information about the discharge time difference and the sign. The switch for turning off the current channel is integrated into the input current modulation control module.
[0077] The simplest embodiment of a two-layer neural network with 4 inputs, 2 intermediate layers, and 1 output*. This embodiment is an extension of the previous embodiment, as follows. Figure 3 Besides the input section, it has two layers. The middle layer is located between the input and the last layer. The middle layer module (35) combines the functions of the layer input current modulation control module and the neuron output module in the simplest embodiment above. Because this embodiment only has one output, the middle layer module (35) has only one set of two resistors (positive and negative resistors), which are respectively connected to the positive and negative capacitors and the capacitor control module of the last layer. The function of the neuron output module in outputting positive and negative signs is directly used in the middle layer module to select the positive and negative capacitor modules and the positive and negative resistors.
[0078] The neuron output module (34) has an output channel (39) for positive and negative signals. Meanwhile, the last layer in the embodiment is designed as a single linear input (single pulse modulated current), and then the level signal output by the last layer is nonlinearly activated. The time difference information of the capacitor discharge of P xor N mentioned above is given to the capacitor module and capacitor control module at the end through the form of high and low level control of the charging and discharging switch (the driving voltage is the voltage source or ground selected by the layer), that is, to charge the capacitor module. The voltage is also selected according to the actual situation. Here, the initial voltage after pre-charging is set to c=1 volt, the driving voltage for formal activation is a=2 volts (with a standard resistance inside), the voltage after activation is assumed to be gp, and then it is discharged from gp to ground through the standard resistor. The comparison voltage to stop detection is set to b=c=1 volt. The duration of discharge in the final detection stage is expressed using a continuous level (the specific expression can be continuous level / voltage expression / current modulation / digital readout channel of fast access unit / other options, designed according to actual needs). As previously calculated, this conversion process is a nonlinear function f(y)=cap*R*(ln[2-e^(-y)]), (y=tp0-tn0 or y=tp1-tn1), with its first quadrant shape resembling tanh. Therefore, the final output is nonlinear.
[0079] Example of ADC detection-type nonlinear activation* Figure 3 The final output duration is used to charge the capacitor of the additional connected layer at the end, and then the capacitor-controlled module (37) is detected by the ADC, and finally the detection result (38) is output in some form. The detection result of the capacitor voltage is another nonlinear function of the original output duration.
[0080] An example of implementing arbitrary activation functions in the layer*, where we have previously disclosed how to obtain a linear transition (discharge mode), as shown in the following formula: tp0-tn0 = cap*R*(X1*W1+X2*W2+...+Xn*Wn)*t; tp1-tn1 = -cap*R*(X1*W1+X2*W2+...+Xn*Wn)*t; The following explains how to achieve a faster, more accurate, and more flexible activation function based on the linear output of the discharge mode. In this embodiment, as... Figure 2 The neuron output module (43) has a control circuit (41) for the capacitor control module (40), which can control whether the capacitor control module discharges and can temporarily interrupt the detection of the discharge process.
[0081] During the detection process, when the capacitors in the positive and negative capacitor control module discharge, they output a level representing the comparison between the capacitor voltage and the detection voltage to the neuron output module. At the beginning of the discharge process, because the residual voltage from the previous pre-charging and neural network discharge process is much greater than the detection voltage (the discharge mode is designed to ensure that the capacitor voltage is greater than the detection voltage by shortening the network discharge time; to ensure accuracy, even after the neural network discharges, the voltage of the capacitor and the detection voltage should have a sufficient amplitude, because if the voltage after the network discharges is too close to the detection voltage, various delays, leakage currents, and other interference factors will impair the calculation accuracy), the outputs P and N of the positive and negative capacitor control module are at the same level, so P xor N is logically equal to 0. Then, because the voltages of the positive and negative capacitors are generally not equal, the capacitor with the lower voltage will reach the detection voltage first during the detection discharge process, so its P xor N will logically become 1 first.
[0082] XOR Level Pre-alignment: Assuming there are many neurons in a layer, we need to use the edge signal where Pxor N logically switches to 1 to align all XOR levels. That is, the detection process first discharges until Pxor N logically switches to 1. During the switch, the neuron output module (control signal to the capacitor control module) temporarily stops detecting and discharging. The XOR level pre-alignment process is complete when all Pxor N values inside all neurons become 1 (or after a reasonable waiting time when all Pxor N values are expected to become 1). Because the discharge of each capacitor control module is temporarily stopped, the capacitor voltage remains unchanged for a short period, and all Pxor N values remain at logic 1.
[0083] The activation function value source (42) is shared by the entire layer, which usually uses the same activation function. After XOR level pre-alignment, when the activation function is executed, the neuron output module sends a unified signal to the capacitor control module to continue detecting discharge. When the neuron output module of a certain neuron changes its P xor N logic, it will automatically record the information of the activation function value source of the next layer, that is, the relevant value (and sign) of the (non)linear function with duration as the independent variable.
[0084] If the activation function is positive and negative, meaning the absolute value of the activation function differs depending on the sign of the duration, and the activation function curves of the positive and negative quadrants of the independent variable have different shapes and signs, then the layer activation function numerical source needs to output two values (positive and negative) and two signs (the signs of the activation function values of the positive and negative quadrants of the independent variable). Based on the values and signs recorded by itself, as well as its own sign, the neuron output module ultimately selects the charging and discharging capacitors for the next layer and outputs the modulation current (PWM, etc.) and the resistance value of the resistor network.
[0085] The so-called simultaneous output of two values (positive and negative) and two signs, more specifically, means having two sets of channels, one positive and one negative, to simultaneously generate non-linear values and output them to the neuron output module.
[0086] There are various forms of numerical sources for layer activation functions: If the layer activation function source is single-line, it releases a non-linear pulse (pulse count Count = f(t), where t is time). The neuron output module contains a counter, preset to 0 before each activation. While Pxor N is logically 1, it counts the non-linear pulses. When a single neuron's Pxor N becomes 0 again, the counter stops. The number of pulses stored in the counter's internal fast access unit represents the output of each neuron's activation function, which can then be used to generate current modulation, sustained-level signals, or be directly read.
[0087] In another form, the layer activation function value source, in the case of multi-line output, when Pxor N becomes 0 again, can directly copy the value of the fast access unit of the timing source into the internal fast access unit of the neuron output module through its dedicated data bus. Of course, at the beginning of activation, the layer activation function value source is either set to 0 or copied to its initial value (when Pxor N becomes 1), and can also be (non)linearly added over time. If the value of the layer activation function value source is copied during both Pxor N 0 / 1 bidirectional shears, and the layer activation function source is linear, then the difference between the two values is the value of the time pulse duration. This linear value can be used later for memory lookup and interpolation calculations. (If pre-alignment is not performed, then copying the two differences and subsequent processing are necessary.) (The encoding form of the multi-line / multi-bit value is not limited to binary and is compatible with: Gray code (reducing bit transition interference), one-hot code (simplifying comparison circuits), symbolic numeric code (optimizing symbol processing), etc.; it can also be grouped, with signal lines of different hardware types, bit counts, encoding forms, rates, loads, and power consumptions divided into different groups; the encoding selection is determined based on performance requirements, circuit noise characteristics, and power consumption requirements. Other embodiments are similar. For example, to achieve higher frequencies and resolutions, conventional numerical signals are used for high-bit / low-frequency transitions, while multi-channel multi-phase signals are used for low-bit / high-frequency volatile parts. Even when the number of bits is small, a delay chain is used instead of a crystal oscillator clock for high-frequency signals.) The layer activation function numerical source module uses a high-performance microcontroller or DMA to directly read pre-stored values, or a combination of both, or other specially designed methods to generate and output the fast access unit values of the layer activation function numerical source. This is common knowledge and software technique, and will not be elaborated here. As for linearly scaled functions, it is only necessary to change the fundamental frequency of the timing source or (if any) change the value of the output independent variable.
[0088] Dynamic alternation-layer propagation implementation example* Figure 5 / Figure 6 ( Figure 6 It's about the design layout. Figure 5 yes Figure 6 (Logical expansion). The bidirectional intermediate module (50) combines the functions of the neuron output module and the network link input current modulation control module, and also integrates the function of the (non)linear activation function of the previous embodiment. Thus, together with the weight term current modulation control module, the capacitor module, and the capacitor control module (51), the entire basic computing unit is formed. The basic computing unit can be the basic neuron unit of any layer. From the perspective of the physical structure of the circuit, the entire network has two layers, dynamic odd layers (52)(55) and dynamic even layers (53)(56). The bidirectional intermediate module has parallel lines connected to the storage module, which can load data from the storage module into the fast access unit inside the bidirectional intermediate module in parallel. The storage module can be any form of digital storage circuit, such as the well-known mature array memory SRAM, DRAM, flash, etc. (of course, it also includes other new storage devices such as RRam / MRam / PCRam). In this embodiment, the weighted current modulation control module (54) in the figure is drawn as a pattern of a single equivalent resistance for ease of drawing. The weighted current modulation control module in this embodiment is also a network resistor (i.e., a combination resistor with multiple selectable resistance values). The current control bus (49) and (57) have current channels or multiple charge / discharge control signal lines connected to the network resistor module. The current switch for modulating the current in the figure is set in the bidirectional intermediate module, but it can be set at any position in the current loop as needed.
[0089] Dynamic odd and even layers are not fixed relative to the overall structure of the neural network; rather, they dynamically alternate as layers. Of course, the number of neurons in dynamic odd and even layers needs to be greater than the number of neurons in the largest layer of the neural network (if less, multiple rounds of calculation are required before merging the results to obtain the activation value). Unused neurons and circuits can be temporarily shut down to conserve energy. For example, a switch can be added to the charging and discharging circuit of the capacitor module (e.g., a switch between the capacitor and ground, controlled by the capacitor control module); or the duty cycle of the current modulation can be set to 0, meaning the network link input current modulation control module does not charge or discharge (high impedance state); or, the weight current modulation control module is a resistor network, and the resistance value can be controlled to be set to non-conductive.
[0090] The specific working process is as follows: 1. The dynamic odd layer serves as the first input layer of the entire network, while the dynamic even layer serves as the first computation layer, i.e., the layer below the input layer. The dynamic odd layer loads the input data of the entire network into the fast access unit (network weight parameters / current modulation fast access unit) inside the bidirectional intermediate module. Of course, other data will also be preloaded into the fast access unit, including fast access units for various control settings, which are loaded in parallel from the storage module or directly from the outside. The dynamic even layer resets its fast access unit, and its capacitor module performs pre-charging and other operations.
[0091] 2. Initially, the odd and even layers of the network are activated by power-on. The dynamic odd layer generates a modulated current based on the fast access unit values, which charges and discharges the capacitor module of the dynamic even layer through the weight resistor (the charging / discharging mode is selected according to the network design). Simultaneously, the capacitor of the current input layer, i.e., the dynamic odd layer, begins pre-charging, and the fast access units related to capacitor control are reset. The dynamic even layer loads data into its internal pre-read fast access unit (for neural networks, this is the network weight parameters of the next layer). This continues until the dynamic even layer's capacitor discharge detection and the execution of the activation function are completed. The result of the activation function is stored in the output fast access unit of the dynamic even layer (i.e., the input of the next layer, the fast access unit of the activation function result + the fast access unit of the next layer's network weight parameters; the data from both together generate the current modulation).
[0092] 3. The odd-layer network is activated by power-on, meaning the dynamic even-layer generates a current modulation based on the fast access unit value, which charges and discharges the capacitor module of the dynamic odd-layer through a resistor (the charging / discharging mode is selected according to the network design). Simultaneously, the capacitor of the current input layer, i.e., the dynamic even-layer, begins pre-charging, and the fast access units related to capacitor control are reset. The dynamic odd-layer begins loading data into its internal pre-fetch fast access unit (for neural networks, this is the network weight parameters of the next layer). This continues until the capacitor discharge detection and activation function execution are completed, and the activation function result is stored in the output fast access unit of the dynamic odd-layer (i.e., the input of the next layer).
[0093] … The above two-thirds steps are executed alternately, dynamically changing to represent the various layers of the entire neural network. The network computation is performed layer by layer, and when the final layer is reached, the final computation result is stored in the storage module or output externally.
[0094] It's important to note that the "odd" in "dynamic odd layer" is relative to "dynamic even layer." This is for ease of description and to conform to language conventions; it doesn't limit the number of layers in the network to an odd number, but simply describes an alternating propagation process. Furthermore, if the neural network layers to be computed are parallel layers at the same level or parallel neural network modules (if there's only one parallel layer, it's treated as a single-layer neural network), then the current computation results need to be stored in the storage module, and the input and network weight parameters need to be read again from the storage module or externally to compute the parallel layers one by one.
[0095] Current modulation includes current modulation of the network input layer and instantaneous current adjustment representing the weight parameters of the resistive network (the weights can be implemented in the form of an equivalent resistive network, modulated current, or a combination of both; this embodiment only uses a resistive network). The current modulation weights representing resistance occupy the largest number of transistors in the fast access unit (HAU). In a fully connected network, the number of HAUs is proportional to the square of the number of neurons N. For example, assuming 10,000 neurons, each neuron has 10,000 connections. Assuming a 32-bit HAU with 6 CMOS transistors per bit, this would require 19.2 billion CMOS transistors. Therefore, neural network design should reduce the size of a single layer by increasing the number of parallel model layers and the number of layer levels to adapt to the limited resources of a single chip.
[0096] The bias term explains that there are bias links in the neuron's network. When the current modulation of the input layer of the network is 100% (i.e., it is energized at all times), then the contribution of the entire resistance link to the layer output depends entirely on the resistance value and the current modulation representing the weight parameters of the resistance network.
[0097] Regarding the modulation frequency, when the current modulation frequency is high, the waveform deviates from the square wave, or even becomes a triangular wave, or the modulation waveforms of each input item are not centrally symmetrical or synchronous, meaning that the time or actual current of each input item deviates from the expectation. Therefore, when applying the parameters trained by the software to the fast access unit of current modulation, an equivalent relational expression or conversion relational table should be obtained based on the measurement statistical characteristics. After numerical conversion, it should then be applied to the fast access unit.
[0098] Example of a Dynamic Alternating Convolutional Neural Network* Compared to fully connected networks, convolutional neural networks have fewer and shared connection parameters, but more network inputs and more neurons.
[0099] Compared to the dynamic alternation-layer propagation embodiment above, the neurons (the basic computational units mentioned earlier) in this embodiment are arranged in a two-dimensional array within a layer. They are also divided into dynamic odd layers and dynamic even layers. According to its definition and settings, the two-dimensional coordinates of the neurons (outputs) and the coordinates of the neuron link inputs are explicitly corresponding. This embodiment is no exception; the connections are local convolutional kernel connections, and these connections have explicit positional coordinate relationships. That is, each neuron link with corresponding coordinates has a specific number of inputs and corresponding coordinates (convolutional kernel connections), and these connection circuits are fixed during the design phase.
[0100] Furthermore, since the (linked) hardware circuitry is fixed, the size of the convolutional kernel in the hardware circuitry should ideally be larger than that of common convolutional kernels. The kernel size can be dynamically adjusted by changing the resistor network within the weighted current modulation control module (which is designed as a resistor network) and disabling unused links. Convolution is essentially a simplification of fully connected circuitry by incorporating coordinate information. Because convolution is not fully connected, circuit routing is easier, and chip resource consumption is much lower.
[0101] Because convolutions share parameters, the fast access units (FLUs) for current modulation of network weight parameters can be shared, significantly reducing the need for a large number of FLUs. However, FLUs for current modulation of network layer inputs are still required. Since parts of a convolutional neural network are consistent (only the two-dimensional coordinates differ), the convolutional layer input can be divided into multiple parts for separate computation, reducing the demand on chip transistors.
[0102] In summary, based on the previous embodiment, two-dimensional encoded information is added to the neurons, fully connected connections are replaced with convolutional connections, and shared fast access units (based on two-dimensional coordinates) are used. Then, the network circuits are dynamically and alternately powered on / detected / activated layer by layer until the computation of the convolutional network is completed, and the data is stored in the storage module or output externally.
[0103] Self-attention matrix operation neural network example* The self-attention model calculates Q, K, and V by multiplying the input X by the parameter matrix. The calculation process is similar to the calculation process of tp0-tn0 or tp1-tn1 in the fully connected linear activation above. Although it often uses matrix mathematical operations to express and describe it, there is no difference in essence. Both involve multiplying the input terms with multiple weight terms and finally summing the multiple product terms.
[0104] Then, positional encoding information needs to be added to the Q or K matrices. This encoding information can take many forms, but the most common approach is to add a calculated positional encoding value to each value of the input vector X before generating the QKV matrices. This calculation is performed before multiplication and addition. Therefore, it can be performed by common digital circuit modules (CPU, FPGA, DMA, lookup table circuits, etc.). However, the circuit of this invention can also be used for calculation because in many subsequent embodiments, the neuron output module has a digital adder / CPU, and the general-purpose circuits for internal adders for neuron input and output functions can be shared. In other words, this part of the circuit can be directly used for positional encoding addition. Alternatively, positional encoding information can be directly generated at the interface when reading the input vector and added to each element value of the word vector, thus enhancing the computational capability of the reading interface.
[0105] In some large model designs, positional encoding information is inserted into the generated Q / K / V matrix. This insertion of positional encoding information also utilizes the bias parameters of the original neural network. As mentioned earlier, when the current modulation of the network input layer is 100% (i.e., energized at all times), the contribution of the entire resistive link to the layer output depends entirely on the resistance value and the current modulation representing the weight parameters of the resistive network. Therefore, during the generation of Q / K / V, the bias parameters can be set to add the corresponding positional encoding information to the output values of the relevant neurons. In other words, when generating the Q / K / V matrix, positional encoding information (specific numbers loaded from the storage module or externally) is added to the output values before activation using the bias parameters.
[0106] Then, when multiplying multiple QK matrices, the multiplication of a single QK matrix (including matrix transpose, which is essentially the sum of multiple product terms) requires writing the value of one matrix into the input term current modulation control module's register representing the input term, and the value of the other matrix into the weight term current modulation control module's fast access unit representing the weight parameters. Then, forward propagation / power-on detection activation is performed according to the process. That is, in one round of calculation, one set of input values is multiplied and added with multiple sets of weight values respectively, and the output value is stored in the fast access unit. The result of matrix QK multiplication is obtained through multiple rounds of calculation.
[0107] Therefore, the process of generating and multiplying the QK matrix in the self-attention mechanism, as well as the final QKV matrix operation, can be calculated using the circuitry of our already disclosed embodiments. The steps are as follows: (According to the definition of the self-attention model, Q / K / V refer to the intermediate value matrix generated by multiplying different parameter matrices with the input vector matrix). First, during the generation of the QK matrix, the system performs multiplication and addition calculations for each row of the matrix through multiple rounds of operations. Specifically, the system reads a portion of the Q matrix parameters (one row or one column per round) from external memory or internal cache and loads it into the fast access unit of the input current modulation control module. Simultaneously, the position-encoded input vectors (i.e., word vectors, which may be multiple) are loaded in parallel into the fast access unit of the weight current modulation control module. Then, the neural network performs a forward propagation calculation, outputting the Q matrix result value for the current round (if the number of words exceeds the number of neurons, it needs to be divided into multiple batches and multiple rounds of calculation; also, the parameter matrix may have multiple rows, requiring several rounds of calculation), and temporarily stores it in the spare fast access unit of the corresponding weight module. This process is repeated until all Q matrix results are filled into the spare fast access units of the corresponding weight modules. In addition to multiple rounds of calculation, the output vector can also be calculated in blocks within a single round, simultaneously calculating QKV, with the results of each block written to the corresponding storage circuit.
[0108] Next, the system enters the K-matrix processing stage. At this point, the system loads some parameters of the K-matrix (one row or one column per round) into the fast access unit of the input module, and simultaneously loads the position-encoded input vectors (multiple word vectors) into the main fast access unit of the weight module in parallel. The neural network performs forward propagation calculations, generating the corresponding K-matrix result value, and passes this result back to the fast access unit of the input module. Then, the contents of the main and backup fast access units of the weight module are swapped (i.e., the Q-matrix result value is read), so that the Q-matrix results of each round are reused as weight values. Afterward, forward propagation is performed, and the neuron output module obtains the result of a round of QK matrix multiplication and addition operations, saving it to external memory or an idle fast access unit. This process is repeated multiple times until the entire QK matrix multiplication calculation is completed.
[0109] (If the positional encoding is inserted during the generation of Q / K / V instead of in the word vectors, then during the calculation of Q / K / V, by setting the positional encoding in the bias term parameters of the neural network, the output of the network calculation will include positional information.) After completing the QK matrix, the system further processes it using the SoftMax normalized activation function (see related examples below), and temporarily stores the activation results externally or in memory. This process also gradually generates all QK activation results through multiple rounds of calculation.
[0110] During the V matrix generation phase, the position-encoded input vectors corresponding to multiple terms are loaded into the fast access unit of the weight term current modulation control module. One round of V matrix parameters is then loaded into the fast access unit of the input term current modulation control module. In the forward propagation computation, the neuron output module generates the V matrix result and outputs it to the spare fast access unit of the weight term current modulation control module or external / memory. This process is repeated until the spare fast access units for all weight terms are filled.
[0111] Finally, the system enters the QKV matrix generation stage. The weight term current modulation control module switches to the backup fast access unit, which contains the V matrix values corresponding to each term. The result of multiplying the QK matrix (originally stored in memory or externally) by the SoftMax-normalized QK matrix (i.e., the attention weights) is read into the fast access unit of the input term current modulation control module in multiple rounds. Then, forward propagation calculation is performed to obtain the final QKV matrix calculation result. This result is then used for further calculations.
[0112] It's worth noting that, to accommodate the processing needs of large-scale input vectors, the system supports grouping input vectors for processing. That is, each batch processes only a portion of the input data, and through multiple rounds of computation, the complete output result is finally concatenated (using addition). This approach not only improves the system's flexibility but also effectively reduces hardware resource consumption. Word vectors don't necessarily have to be read into the weights; they could be placed in the inputs as well. However, word vectors are significantly more numerous than parameter matrices, and weights require more circuit resources. Therefore, they are placed in the weights, while the parameter matrices are placed in the inputs because of their smaller number.
[0113] In summary, this embodiment implements the core computational process of the self-attention mechanism through hardware circuit design, featuring high efficiency, low power consumption, and strong scalability, making it suitable for the deployment and application of Transformer-type models on low-power edge devices or dedicated AI chips.
[0114] SoftMax Implementation Example*: Calculate the activation function SoftMax, which is defined according to a well-known definition: SoftMax = e^(Yi) / SumAll[e^(Y1), e^(Y2),...e^(Yn)] This involves dividing the exponent of the current neuron's natural base by the sum of the exponents of all neurons. While this calculation process can certainly be performed using an external chip module for combined internal and external computation, our implementation offers better parallel performance. For example... Figure 7In this embodiment, dynamic odd layers (61) and dynamic even layers (62) are alternately powered on and executed. According to the above embodiment of implementing arbitrary activation functions of the layer, the function e^(Yi) can be calculated, that is, the PWM output of this layer is equivalent to e^(Yi) (proportional to e^(Yi)). Then, in the next layer (66), the linear sum of this layer, that is, the part of S=SumAll[e^(Y1), e^(Y2),...e^(Yn)], is calculated. Here it is called the S value. The function of SoftMax's whole layer normalization is mainly implemented through this part. Since SumAll[e^(Y1), e^(Y2),...e^(Yn)] only requires one neuron, then in hardware, an additional neuron unit with a special bidirectional intermediate module (67) can be designed as the network summarization layer (66). The register of the output of the corresponding neuron can be directly used by the layer calculation module (64) through the data path (65). After the layer calculation module obtains the S-value parameter, it uses its reciprocal to generate a fast access unit value (usually generated by a high-frequency microcontroller) (modulated by PWM, etc.).
[0115] Then, this value is simultaneously written to each neuron through the parallel write channel (69) connecting the entire layer to the fast access unit of the PWM part representing the weight value of the output PWM. This value is W=a / SumAll[e^(Y1), e^(Y2),...e^(Yn)], and the value of the fast access unit of the PWM representing the output value of the current layer is e^(Yi). The product of the two is proportional to SoftMax (= a*e^(Yi) / SumAll[e^(Y1), e^(Y2),...e^(Yn)]). The proportional coefficient a needs to be adjusted according to the actual needs based on the resistance and capacitance value through the base frequency. The link to the next layer here is a one-to-one connection. Because it is a one-to-one connection, the functions of the current layer and the next layer can be merged, and the system can charge and discharge itself. That is, the current bidirectional intermediate module (63) charges and discharges the capacitor (68) of its own neuron unit through the internal resistor using PWM, and then discharges and compares the discharge time to obtain the discharge time. Then, linear activation is performed (that is, e^(Yi) is scaled proportionally and normalized according to the S value) to update the output register value. Then, the entire neural network is calculated alternately as in the dynamic alternating-layer propagation embodiment disclosed above. Because the system charges and discharges itself, the sign is known or only the absolute value is known, so the capacitor (68) is selected.
[0116] In this way, after the operation is completed, a set of multiple SoftMax activation probabilities is obtained. If multiple sets are computed in parallel, the results of multiple SoftMax activations can be obtained simultaneously.
[0117] Paired Residual Neural Network Example* A residual neural network is essentially two neural networks: a linear neural network and a nonlinear neural network. The output values of the output layers are added pairwise according to their coordinate positions (the arrangement of neurons) to form a new output. This example is not fundamentally different from the examples described above. What is required is a circuit for pairwise addition. This circuit can be implemented using a digital adder, the capacitor charging and discharging mechanism of the above example, or a time difference alignment method similar to that used in the example implementing arbitrary activation functions in the layers.
[0118] The time difference alignment method involves pre-aligning the output layers of two neural networks. This means that the positive and negative capacitance control modules of the two neurons in each output layer first perform a detection process to discharge until their P XOR N logically switches to 1 (meaning P and N are logically different). Then, the discharge switch of one of the modules is interrupted until the P XOR N of the continuously discharging module logically switches to 0 (meaning P and N are logically the same again). The interrupted module then continues discharging until its P XOR N also becomes 0. The time from the completion of pre-alignment to the point where both modules' P XOR Ns become 0 is the sum of their output times. A simple logic circuit can achieve the function of summing and merging the duration difference between the two neuron computational units. This parallel residual calculation, i.e., the calculation based on the sequence number and coordinates, is achieved through the continuity of the time difference level. Of course, an interconnecting circuit and a switch based on the coordinates are needed to conduct the discharge of the two neurons.
[0119] Dynamic single-layer propagation implementation example*, such as Figure 8 The neuron output module or intermediate module (73) has a register channel (72) that directly transmits the value of the output register to the input current modulation control module (71), sharing or jointly using the value of the fast access unit of the neuron input / output items (the input part of the fast access unit for modulation such as PWM). The value of the fast access unit of the output neuron output item is copied and written to the fast access unit of the input current modulation control module, or the two modules share a fast access unit. Compared with the dynamic alternating-layer propagation embodiment, this embodiment saves a lot of register resources, but CMOS resources are usually mainly consumed in the circuits corresponding to weight items and links, and the resources occupied by the circuits corresponding to input items and output items are very small. The working process of the two embodiments is similar. The basic computation unit of the neuron reads the weight item parameters and resistor selection parameters, as well as the output value of the previous layer network, from the storage module each time, and then uses them to calculate the output value of the current layer. In this process, the entire circuit is dynamically used to calculate each layer of the neural network, and finally outputs the result to the outside or writes it to the storage module.
[0120] The working process of a dynamic single-layer propagation implementation is as follows: 1. Basic neuron units load data from storage modules or external sources into corresponding fast access units. (These include fast access units for various control settings, weighted items, input items, resistance selection, etc.) 2. The capacitor begins pre-charging. (If there are fast access units related to capacitor control, they will also be reset before charging.) 3. (Forward propagation steps 3~7) The network is powered on, that is, the input current modulation control module or intermediate module generates a modulation current according to the value of the relevant register, and charges and discharges the capacitor module through the resistor (select the charging mode / discharging mode according to the network design).
[0121] 4. Capacitor discharge is detected and the activation function (if any) is executed. The output result is stored in the fast access unit of the output item (i.e., the input of the next layer). The layer activation function value source needs to be set before executing the activation function.
[0122] 5. Copy the value of the output item's fast access unit to the input item's fast access unit of the current neuron's basic computational unit. (If shared, omit the copying process; sharing saves CMOS resources.) 6. If there are special calculations involving layer parameters, such as SoftMax, then perform the relevant calculations for the network aggregation layer.
[0123] 7. Once all layers of the neural network have finished computing, write the computation results of the output fast access unit to the storage module or external storage.
[0124] 8. If the calculation is not finished, load other data required for the next layer of the neural network calculation (weight fast access unit, resistor selection fast access unit, related calculation settings, etc.) from the storage module or externally.
[0125] 9. Repeat steps 2 through 8 until all layers of the neural network have been computed.
[0126] Multiple dynamic layers can be arranged together symmetrically in space to balance the unidirectional flow of current in various situations. Figure 4 As shown, when the capacitor discharges, the current flows unidirectionally downwards and then back to the capacitor ground.
[0127] An example of activation after multiple accumulations*, the previous dynamic single-layer propagation example, such as Figure 8 The number of input items in the input layer and the number of output items in the output layer are equal. Although the number of the two may be different in some embodiments described above, it is still fixed. At most, the number of input items in the linked input layer is reduced by closing the links between the preceding and following layers.
[0128] In this embodiment, the output module of the output layer has a counter for the activation function timing pulses. If the fundamental frequency is relatively high, the layer activation function value source can have multiple pulse lines compared to a single-line pulse counter. For example, 1x, 2x, 4x, 8x, etc. pulse lines. For instance, for an 8x line, one pulse increments the counter by 8.
[0129] The microcontroller and DMA (Memory Direct Write) corresponding to the layer activation function value source generate the programmed activation function parameters. The internal circuitry of the layer activation function value source generates the aforementioned multi-line pulses based on these parameters and the base frequency. Activation function parameters include, for example, an interval time register specifying how many base frequency pulses are needed to generate one multi-line pulse; how many base frequency pulses the interval time register increments or decrements by 1 (i.e., the interval time second derivative register); and how many base frequency pulses need to be read from memory again (i.e., the update time register). In short, the chip calculates the function parameters for pre-generating interpolation, and the internal circuitry generates the interpolation multi-line pulses based on these parameters. The layer activation function value source is positive and negative, meaning it generates two sets of multi-line pulses simultaneously, one positive and one negative. The output module selects the positive or negative multi-line pulse for counting (representing a multi-bit increment) based on its own sign.
[0130] like Figure 10 A large input layer (90) is divided into multiple smaller input layers (91)(93)(95)(96). To accumulate the output values of these smaller input layers, the charging and discharging of network capacitors corresponding to the number of neurons in the input layers, as well as detection and activation operations (forward propagation) (92), are required. These operations use a linear activation function y=x to convert the duration difference of the xor level of the positive and negative capacitors into the value of the fast access unit of the counter. The counter adds a multi-bit increment value expressed by the multi-line pulse each time, which is the time increment value of the activation function.
[0131] The forward propagation of the input layers (91), (93), (95), and (96) all use the same circuit modules, but the input and weight parameters they read from the external array storage modules are different. Figure 10 The layers in the middle are divided into parts of the multi-round operation process, and the physical circuits are the same.
[0132] In this embodiment, the output layer (94) and the input layer (91) are powered on and perform the detection and activation processes completely, thus completing the first forward propagation of the segmented layer. This is because the activation function used is a linear function of y=x, and its value is stored in the fast access unit of the activation function timing pulse counter in the output module.
[0133] Then, the second forward propagation is performed, namely the forward propagation of the output layer (94) and the input layer (93). Of course, before the input layer (93) performs the forward propagation, the corresponding input items and weight parameters are loaded from the external memory (or the input module itself has more fast access units to temporarily store the values of the input items in advance). However, it is different when the activation function is executed. In the second forward propagation, the counter of the activation function timing pulse will not be reset before activation. Because it is a linear function of y=x, the activation of the neuron is uniformly increased, and the positive and negative pulse lines are the same. Therefore, it is only necessary to continue to accumulate or decrement the pulse count during the duration of the difference in the duration of the positive and negative capacitor XOR levels when the logic level is 1. Of course, if the sign of the positive and negative capacitor logic synthesis is negative (that is, the output time of the negative capacitor module is longer), then the counter should select the pulse signal of the negative layer activation function value source, and should decrease the count value of the counter. If it is reduced to below 0, then the logic level signal related to the sign of the entire output module needs to be flipped. After the second process is completed, the result of the fast access unit of the counter is the absolute value of the arithmetic sum of all product terms in the two forward propagations.
[0134] Repeat the second forward propagation process, calculate the forward propagation of the output layer (94) and the input layer (95) (96) respectively, and then the result is the arithmetic sum of the products of input terms and weight terms of the forward propagation of each neuron according to all input layers.
[0135] In this embodiment, forward propagation of input items that are a multiple of the number of output neurons is achieved through step-by-step calculation. Furthermore, this embodiment can also implement a residual neural network. Compared to the paired residual neural network embodiment*, this embodiment does not require pairing circuitry based on neuron coordinates or positions in hardware. Its operation is as follows: First, during the first round of forward propagation at the first layer, the appropriate (non)linear activation function is selected.
[0136] Then, in the next layer of forward propagation after the second round, a linear activation function for y=x is chosen.
[0137] The result is that the output value of a single neuron equals the first (non)linear activation value plus subsequent linear activation values. Examples of activation functions with learnable parameter β, such as Swish*, illustrate this. Swish's activation function means that the curve of the activation function for each neuron in a layer is adjustable based on the parameter β. Depending on β, activation functions with specific curve shapes, whether linear or non-linear, can be chosen; that is, the activation function for each neuron in a layer is different. This inconsistency significantly increases hardware complexity. However, to directly apply pre-trained models in software and avoid retraining and adjusting model parameters, a circuit can be designed to execute this activation function. Furthermore, the computational circuitry related to neural network weights consumes considerable circuit resources (n neurons, each with n links, then the consumed circuit resources are n^2), especially for large models. Therefore, even if the activation function-related circuitry provides more complex functions and consumes more chip resources, it is relatively small.
[0138] This embodiment uses multiple layers of activation function numerical sources, each with corresponding positive and negative pulse lines, generating pulses for the Swish activation function with different β parameters. Each neuron's output module contains a counter that selects the corresponding layer activation function numerical source's positive and negative pulse lines based on its own β parameter number. Thus, each neuron selects a different activation function. Of course, before activation, each neuron's output module loads the β parameter number.
[0139] While increasing the number of activation function numerical sources by multiple layers doesn't consume more chip resources than the weight calculation circuitry, the number of layers is still finite. Chip design can limit the number of layers, such as 8 or 16. When training large models with software, the β parameter can be adjusted initially, and later, the β parameter can be gradually clustered so that the β parameter of the trained neurons is ultimately constrained to a limited number of specific values.
[0140] When some β-parameter neurons perform capacitance detection charging and discharging and generate XOR signals with duration differences, the operation can be interrupted first, and the activation of certain β-parameter neurons can be processed first. That is, because the number of layer activation function numerical sources is finite, depending on the β-parameter index, some β-parameter neurons can be activated first, followed by others. By activating neurons with different β-parameters in each round, the multiple sets of layer activation function numerical sources will change their corresponding β-parameters and generate new pulse signals with different function curves. By loading different β-parameter Swish function data into each set of layer activation function numerical sources in different rounds, a multiplication effect of the finite number of layer activation function numerical sources can be achieved. Of course, because the detection process involves multiple rounds of activation, there will be a corresponding delay.
[0141] (In this embodiment and the previous embodiment, a more flexible approach is to use a CPU-like digital calculation method. During the forward propagation discharge detection process, the time value sent by the layer activation function source is first recorded at the start / end of the XOR signal with the duration difference. Then, the digital time value of the duration is obtained by the difference between the two start / end time values, thus obtaining the digital expression of the current round of discharge detection duration. An adder is then used to accumulate the value with the value from the previous round. If activation function transformation is required during the process, a CPU lookup table method can be used to read the function reference value, derivative value / derivative change value, second derivative, function scaling factor, etc., from memory based on the digital time value, and the activation function value is obtained through interpolation calculation.) Examples of SwiGLU and similar gated activation functions* The SwiGLU gated activation function is essentially the product of the output values of two parallel layers in a neural network, paired by the same index n or the same coordinates, where each layer uses a different activation function. For example, suppose the two layers are A and B, and the values of the neurons before activation are (A1, A2...An) and (B1, B2...Bn), respectively. According to the SwiGLU documentation, SwiGLU(n) = An*Swish(Bn). The previous example already described how to implement the Swish activation function in hardware.
[0142] (According to publicly available information, the β parameter of the Swish activation function in SwiGLU is 1 or the β parameter of the entire layer is the same. The activation functions of each neuron in the same layer are consistent, so it is not as complicated as the previous embodiment. Refer to the embodiment of implementing arbitrary activation functions in the layer*) Of course, the output value could be dumped and passed to an external computing chip or FPGA for parallel multiplication, but parallel multiplication performed by a computing chip or FPGA is inefficient. This embodiment implements multiplication without a multiplier, only an adder. To achieve multiplication, the neuron's output module can have a spare fast access unit to temporarily store the output value of the previous activation. For example, the output values of each neuron in A are first stored in this spare register. Then, during the Swish activation of B, a value An is added for each Swish pulse received. That is, B's activation is no longer a counter, but a signed adder that performs an addition for each pulse received. An is signed; if it is negative, subtraction is performed. During B's activation, the timing source pulse line of the activation function receiving the pulse is selected based on its sign (positive or negative). The result is the product of the two.
[0143] (This embodiment can also use the CPU to perform mixed calculations. For example, after the basic circuit for capacitance calculation in this invention obtains the values of An and Swish(Bn), the CPU is used to perform the multiplication of the two.) An embodiment of a 1:1 input-output ratio* Neuron circuit module relationship explanation: When the ratio of input to output terms in a neural network is 1:1, such as... Figure 23a All processes and steps, including enabling each module, writing and reading values, are controlled by a shared controller (260) of multiple neurons.
[0144] The clock source (266) sends the clock carry pulse signal to the global clock (265) and the input current modulation control module (267). The multi-bit cyclic timing signal output by the global clock (265) is compared with the value of the input register in the input current modulation control module to generate the input PWM. If necessary, the input PWM and the direct signal of the clock source can be combined (logical AND) to generate multiple refined PWM pulses. The PWM pulse is sent to the positive and negative complementary signal generator (251) to convert the PWM signal into a complementary PWM (or a complementary signal can be generated globally and then combined with the input PWM) to control the transmission gate charging and discharging switch (252) of the capacitor, which generates the actual PWM modulated charging and discharging current for the capacitor's charging and discharging channel. There are two sets of transmission gates corresponding to the positive and negative capacitors (the transmission gates can balance the parasitic charges injected when the NMOS / PMOS transistors are switched), and one is selected according to the positive or negative value of the input. The instantaneous magnitude of the PWM-modulated charging and discharging current is controlled by the resistor network (250) within the weighted current modulation control module. Its resistance value is controlled by the weight values of the corresponding neural network links. The input current modulation control module and the weighted current modulation control module jointly control the charging and discharging currents of the positive and negative capacitors. This embodiment uses a discharge mode, so the discharge current ultimately flows back to the other end of the capacitor, i.e., the common terminal. The diagram shows only one link; in reality, the neural network has multiple links, meaning the capacitor has multiple discharge channels (253).
[0145] The capacitor control module (259), like the capacitor module, has positive and negative terminals. The positive capacitor control module controls the positive capacitor, and the negative one controls the negative one. Before forward propagation, the capacitor is pre-charged to a certain voltage (256), such as 0.21V, by the pre-charge module (255). Then, the capacitor is discharged through the transmission gate (257) inside the capacitor control module and the standard discharge resistor. The discharge stops when the standard pre-charge voltage, such as 0.20V, is reached. The discharge operation is achieved by the voltage comparator (254) inside the capacitor control module. This voltage comparator can select multiple reference comparison voltages through the voltage selection switch (268) to control the capacitor to achieve multiple voltages. The above is the voltage control process before forward propagation. After the capacitor module discharges during forward propagation, the next step is the discharge detection process. During the detection process, the transmission gate inside the capacitor control module is turned on, causing the capacitor module to continue discharging through the standard resistor until the capacitor voltage reaches the comparison voltage to terminate the discharge detection, such as 0.15V. The capacitor control module contains a fast logic circuit (258). Once the comparator output voltage arrives, the signal level of the entire capacitor control module output to the neuron output module (261) is flipped. The neuron output module contains an XOR circuit, which logically combines the level signals of the positive and negative capacitor control modules to generate a time difference (263) and a positive or negative sign. Finally, the time difference is converted into an activation function value (262) through a multi-bit signal (264) from the activation function value source. This value is then written to a buffer for later writing to memory, transmission to external systems, or rewriting to the input modulation control module, etc., thus completing this round of calculation.
[0146] In other embodiments, the ratio of input, weight, capacitor, and output modules can be adjusted as needed. For example, the operating frequency of the input current modulation control module and the weight current modulation control module may be much higher than the operating frequency of the capacitor module / capacitor control module / neuron output module. Therefore, multiple sets of capacitor modules / capacitor control modules / neuron output modules can be configured, allowing the lower-frequency modules to operate in rotation, thus increasing the overall operating frequency. Generally, the number of weight terms equals the number of input terms multiplied by the number of output terms, and is the largest, so optimizing its operating frequency and area is prioritized. Similarly, multiple sets of input current modulation control modules can also be configured.
[0147] Examples of Using Switched Capacitors Instead of Resistor Networks* Switched capacitor networks, also known as capacitor-resistor networks, can replace conventional network resistors in our circuits, such as... Figure 23bIn a single neuron, a current channel of a certain link is connected to a switching capacitor (423) of different capacitance values. The left and right parallel switches (424 and 425) select one of the channels and then alternately conduct, thereby transporting the charge of the capacitor module (422) to the common terminal of the capacitor module. The effect is equivalent to the fixed pulse width operation of a conventional resistor single pulse, that is, the ratio of the charge release of the two is not affected by the main capacitor voltage. In terms of control, the number of times the switching capacitor is transported is equivalent to the pulse width of a conventional resistor single pulse. The switching capacitor network can be used for network weights, for detecting discharge resistance, or for both. The preferred design is that the neural network input terms (421) and (422) control the number of times the switching capacitor is transported, and the neural network weight terms (426) control the selection of the capacitance value of the capacitor network. The two together control the discharge current of the capacitor module of the corresponding symbol neuron. If the discharge detection is also controlled by the switching capacitor, the pulse width value conversion of the time pulse needs to be changed to the transport count of the switching capacitor. Note: The switches used in the figure are high-frequency switches that can reduce the charge injection (parasitic capacitance).
[0148] Examples of forward propagation in high-frequency networks with more complex current modulation and activation* In some of the preceding embodiments, the fast access unit for current modulation shares the input term portion, while the link-related portion for weight terms is multiple and independent. The actual switching occurs on the circuitry of the weight term current modulation control module (the connections between the preceding and following layers of neurons). However, this results in a large number of register bits for current modulation. For simple PWM modulation, this leads to a lower operating frequency; for other commonly used current modulation methods, it becomes more complex and consumes more chip resources. Because the number of neuron links is much larger than the number of neurons, the chip resources associated with neuron links increase quadratically (number of neurons multiplied by the number of links per neuron). However, if the overall operating frequency and accuracy of the network can be improved, this increase in complexity and resource consumption related to weight terms is a feasible solution that can be considered.
[0149] In this embodiment, as Figure 11The input current modulation control module (102) generates an input PWM (100) based on the layer input value of the input layer neurons of the neural network, i.e., the value of the input register. Based on the logic level of the input PWM, it selects or masks the underlying high-frequency reference PWM (in the form of a logic AND, OR, NAND, or OR selector). The width of this high-frequency standard PWM adopts the standard frequency, standard pulse width, and standard duty cycle as much as possible. (At extremely high frequencies, there can be multiple PWMs with different pulse widths generated by inverter delay chains (centrally symmetrical, in binary values). The input current modulation control module selects the PWM signal based on the input value using a selector to generate a high-frequency input PWM signal.) Then, the reference pulse signal generated based on the underlying high-frequency reference PWM and the input PWM (100), called the input synthesized modulation signal wave (101), is given to the weighted current modulation module (103) of the relevant neuron link. Only the current path for charging and discharging the capacitor module is shown in the figure, but the path for the synthesized modulation signal wave given to the weighted current modulation module exists in the actual circuit. The weighted current modulation module includes a positive / negative selection switch module (104), which selects whether to charge or discharge the positive or negative capacitor module by the positive or negative result of the input item and the weighted item register (i.e., the positive or negative result after multiplying their respective signs).
[0150] The weighted current modulation module includes a weighted fast access unit, a resistor network, and a switching delay circuit, which control the charging and discharging current of the capacitor module (105). Specifically, based on the value of the weighted fast access unit, the actual instantaneous current is adjusted through the resistor network, while the input-term synthesized modulation signal wave is adjusted through the switching delay circuit to generate a new switching signal, ultimately producing a current modulation pulse of the corresponding width. There are many ways to insert the delay; for example, the edge signal of the input-term synthesized modulation signal wave can be used as a trigger signal, and the pulse width can be completely determined by setting the delay to be off. When using the circuit of this embodiment, the delay can also be modified according to environmental factors such as temperature, voltage, and production deviations.
[0151] It is important to note that if the input current modulation control module outputs only modulation wave information, i.e., signal pulses, and not the actual current pulses used, then the resistor module, its associated positive / negative selection switch module, or other modules in the link have separate switches and drive power supply channels for charging and discharging (discharging to common ground or a designated charging drive voltage source). Because the modulation signal pulse and the modulation current pulse can be separated, it is easy to operate MOSFETs, etc., to generate modulation current as long as there is a switch signal. Therefore, the switch can be placed anywhere in the link current circulation channel, not just within the weighted current modulation module, depending on the specific design. Furthermore, the current switch should not be understood as NMOS or PMOS; any other type of switch that keeps up with the frequency is acceptable. This is because, firstly, it depends on the capacitor and its charging voltage; secondly, the chip design must minimize the capacitance value during capacitor module charging and discharging, the voltage difference during pre-charging and detection, and the capacitor's operating voltage. Therefore, the parasitic capacitance current introduced during MOSFET switch operation will inevitably cause interference. In this case, using both PMOS and NMOS in the transmission gate can balance the parasitic capacitance current injected into the MOSFET during switching, making it a very suitable switch. It might even involve a combination of multiple PMOS and NMOS transistors with different performance characteristics, while adjusting the design of the trench length and area inside the PMOS transistor to further optimize accuracy and reduce noise. Additionally, the discharge mode mentioned in the previous embodiments does not necessarily mean discharging to ground, but rather discharging to the other end of the capacitor, or the capacitor's common terminal. If the capacitor's common terminal is connected to a voltage source, then the other end of the capacitor can be connected to a current path. Furthermore, due to manufacturing variations in the performance parameters of the voltage comparator, the so-called pre-charge to a certain voltage is not necessarily the voltage of a uniform reference voltage source. It can also be a pre-charge voltage obtained by comparing the voltage of each capacitor independently with the reference voltage source voltage through the voltage comparator of each neuron capacitor control module (charging to a slightly higher voltage, then discharging to the actual independent pre-charge voltage determined by the comparator, and then performing forward propagation discharge and detection voltage comparison). The other embodiments mentioned above are also like this.
[0152] The weighted current modulation in this embodiment avoids the use of the simple PWM modulation form commonly used in microcontrollers, because this would lead to a decrease in frequency, especially in cases of high precision or a large number of bits in the weighted register, thus giving the entire network a great potential for increasing its operating frequency.
[0153] In practice, there are many equivalent forms of resistive networks, such as capacitive resistors (using capacitors to transport charge), voltage-controlled MOSFET resistors (active resistors), combined resistors, and newer resistive components like memristors, magnetoresistive resistors, and photoresistors. The focus of this specification is on describing the actual effect, i.e., the equivalent resistance, and should not be interpreted as referring to a specific physical form of resistance. If the neural network is small in scale, has few parameters, and low accuracy requirements, then using a single fixed resistor is also feasible.
[0154] This embodiment is also a dynamic single-layer neural network. The layers on the circuit are dynamically used as the various layers of the entire neural network to be computed. Weight fast access units and input fast access units are loaded from memory (or the current modulation module has built-in storage units) or external circuits. The dynamic single layer is used as the layer being computed in the entire neural network by loading the values of the above fast access units and calculating the activation function or output fast access unit values of the current layer itself.
[0155] The steps are as follows: 1. Load the values of the weights and inputs from memory or external circuitry; 2. Power on the neural network during forward propagation, i.e., charge and discharge the capacitor module; 3. If the entire neural network has not reached the final layer of computation, the basic neuron unit simultaneously loads the weight values of the new layer from memory or external circuitry (preloading); 4. Generate an output value, i.e., capacitor discharge detection, and execute the activation function if applicable; 5. Generate the output value of the current basic neuron unit and copy it to the input unit; 6. If the entire neural network has reached the final layer of computation, output the activation function or the output value to memory or external circuitry, ending the loop; 7. If it is not the final layer of computation, repeat the above process.
[0156] This embodiment employs a discharge mode, where the capacitor modules of the entire neural network are first pre-charged to a specific voltage (assuming a=2V, a controllable charge-discharge circuit can be constructed externally using inductors, bootstrap capacitors, and diodes to recover short-circuit energy, improve energy efficiency, and reduce loop oscillation). Then, the capacitor modules discharge to ground (capacitor common ground) through the weighted current modulation control module, with the current being the modulation current generated by the aforementioned weighted current modulation control module. The detection process then begins, where the capacitor modules continue discharging through the (adjacent) capacitor control module (internal standard resistor) until the voltage comparator inside the capacitor control module produces a voltage comparison result, indicating that the capacitor module voltage is less than the reference voltage (assuming b=1V). At this point, the capacitor control module outputs a detection signal to the neuron output module (110), because each neuron has two capacitor modules (positive and negative) and corresponding positive and negative capacitor control modules. Therefore, the neuron output module integrates the detection signals from the positive and negative capacitor control modules to obtain the time difference information of the positive and negative detection signals (logically XOR, that is, when the logic of the positive and negative detection signals becomes inconsistent), as well as the integrated positive and negative sign information (that is, the positive and negative performance of the capacitor control module output detection signal when XOR logic 1, that is, which of the positive and negative detection signals is true).
[0157] If this layer of the network requires the output of the activation function (i.e., different from the original linear output above), then the layer activation function value source (108) of this embodiment generates a series of parameters or values of the activation function (computing chip, DMA, or specially designed pulse generation module), and sends the output function activation increment information (possibly more than one numerical bit) to each neuron output module through the activation function pulse channel (112). The process is as follows: using the time difference information of the positive and negative detection signals above for pre-alignment, that is, when the XOR gate signal of the comparison detection signal P and N becomes true, that is, when P and N become different, the detection discharge process is interrupted; wait for all neurons to complete the XOR gate signal pre-alignment; and then continue the detection discharge process (starting from the already reset 0 value, receiving and accumulating the function activation increment); when the XOR gate triggers the signal from 1 to 0, that is, when P and N become the same again, the corresponding neuron output module locks the fast access unit of the multi-bit increment counter of the activation function. In this way, the output module of this neuron obtains the activation function value information of the current XOR gate time difference of the layer activation function value source. Similar to the previous embodiments, the activation function pulse channel of the layer activation function value source is internally divided into positive and negative signals. The neuron output module selects to receive signals based on their own sign (the sign is determined when P and N are different). When the XOR gate signals of all neuron output modules end, that is, when the P and N of all neurons are converted back to the same value, the values of the fast access units of all multi-bit increment counters are locked. At this point, all neurons have been activated and have obtained their activation function values.
[0158] If multiple neurons share a common computing core (such as a microcontroller CPU), and the number of cores is small—for example, 128 neurons working simultaneously share one CPU, and the entire network has 1024 neurons—then there are 8 computing cores. These 8 cores can then share a fast memory, i.e., a memory with multiple parallel read / write ports. The computing cores first obtain the time difference of the XOR level using a shared timer. Specifically, they record the timer values for the start and end edges of the XOR level, subtract the time difference, and this value is the time difference, equivalent to the linear activation function value mentioned above. Then, the activation function source writes the activation function information table into the shared memory. The computing cores look up the approximate activation function value and its (second) derivative in the shared memory based on the time difference, and then interpolate to obtain the final activation function value. The advantage of this approach is that the computing cores can operate asynchronously, but it consumes more chip resources.
[0159] If the output of this layer needs to calculate the normalized value in the form of multiplication and addition (such as in the softmax layer), then there will be an additional normalization-specific network aggregation layer, which has a single neuron and its neuron output module (111). Similarly, after multiplication and addition are achieved through charging and discharging, the time difference information of the XOR gate of the positive and negative detection signals is obtained. If further function transformation is required, the single neuron dedicated to normalization will perform activation function operation. As with the above method, the layer activation function value source (108) needs to generate a series of parameters or values of the corresponding activation function, and send the increase value information of the output function activation to the single neuron and its neuron output module through the activation function pulse channel (109) of the single neuron. The value is obtained after the XOR signal ends. Finally, the activation function value is handed over to the computing chip (the layer activation function source contains conventional computing circuits such as computing chips, or layer computing modules). If normalization does not require re-performing parallel multiplication and division, then the network in this layer regenerates the XOR time difference (the network re-executes the charging and discharging of the input, or the digital circuit calculation, or the capacitor control module is controlled to charge and discharge through the internal resistor, etc., until a new theoretically equivalent time difference is generated). At the same time, the computing chip performs appropriate numerical calculations and regenerates the activation function. Then, the value of the output register of each neuron output module of the secondary activation in this layer is the required normalized neuron output value.
[0160] If residual addition or multiple output value accumulation is required, the multi-bit adder counter (the counter for multi-bit addition) of the activation function in the neuron's output module should not be reset during linear activation. The multi-bit adder counter continues to receive increments over time, accumulating multiple activation values (subtracting if the signs differ). Because the multi-bit adder counter itself has the properties of an adder, it can be designed as a single, shared adder, directly adding the output values of two activation functions. In other embodiments, a CPU is also a feasible solution, especially when multiple neurons share a specially designed general-purpose computing unit or CPU, allowing addition, subtraction, multiplication, and division to be performed according to its program.
[0161] If gated multiplication is required, this embodiment has a dedicated linear weight term current modulation control module (114), the number of which corresponds to the number of neurons, and it is single-linked, linked to the positive and negative capacitors of neurons with the same index. If necessary, there is also an additional spare output term fast access unit to temporarily store the output value, or to temporarily store it in external memory. Finally, the output values of the two forward propagations to be multiplied are stored in the fast access unit in the linear weight term current modulation control module and the corresponding input term fast access unit in the input term current modulation control module, respectively. Then the entire network performs forward propagation. Because the input and output are single-linked, the result is proportional to the product of the two values above.
[0162] There is a connection (107) between the neuron output module and the input current modulation control module, which includes an output value channel and a time signal channel. If the output of this layer needs to be used for the next layer's calculation, the value is imported into the corresponding input fast access unit of the input current modulation control module through the output value channel of each neuron, in preparation for the subsequent operation of the next layer. If the input value of the input current modulation control module is propagated in a single chain forward through the straight-chain weight current modulation control module, then the XOR time difference in the output module of the neural network is regenerated.
[0163] (While direct-chain multiplication can be achieved (and numerical offset can be achieved with a bias term), the charging and discharging process of the capacitor introduces errors and delays. By directly sending the time difference level information to the neuron output module through the time signal channel, the current modulation and capacitor charging / discharging processes are skipped. The output module then uses this time difference level information to perform activation operations, including linear and nonlinear activation functions. This operation is faster and more accurate than direct-chain forward propagation. The input current modulation control module is designed to generate modulation signals (PWM pulse width modulation signals, etc., which may require modifying the base frequency for linear transformation), so it only needs to send the first generated pulse to the neuron output value module as the time difference information. A signal line is also needed to input the sign of the input value for the time difference level information.) Additionally, if the activation function for the squared difference (Euclidean distance) is needed, the process is the same as for residual calculation. The residual network is the sum of two outputs; in this embodiment, it's the difference between the two outputs, i.e., with different signs, requiring a signed adder. After obtaining the difference, it's used to directly generate the time difference level information and sign, followed by a second activation (the activation function is squared), resulting in the squared difference (Euclidean distance). (Of course, the result can also be obtained directly using digital circuits with addition / subtraction / multiplication / table lookup / interpolation calculations, etc.) The computing chip can also provide PWM adjustment information to the neuron input current modulation control module by adjusting the signal channel (106), thereby adjusting the input value of the entire layer of neurons. In addition to normalization, it can also be corrected to a certain extent in response to temperature, aging and batch error.
[0164] Repeat the above calculation process until the last layer. Because of image size limitations, the number of neurons and links in this embodiment should not be interpreted as limited to what is shown in the image. The circuit wiring of each module is also not limited to what is shown in the image and should be understood functionally.
[0165] Examples of high-frequency, single-layer dynamic networks with shared links* Building upon the previous embodiment, the number of modules other than the weighted current modulation control module is doubled (greater than or equal to 2 times), the number of switch-related circuits used for current modulation is doubled, and the weighted current modulation control module (including the weighted fast access unit, etc.) is shared. This allows for a multiple increase in the computational output frequency of the entire network, resulting in a performance boost. Furthermore, because the number of weighted current modulation control modules in a fully connected neural network is far greater than the number of other modules, this doubling consumes relatively few circuit resources while maximizing computational power. It also balances the inconsistent delays among modules in the circuit and reduces the memory bandwidth requirements of the weighted parameters.
[0166] After doubling, multiple different networks can be computed simultaneously, meaning networks with the same weight parameters and activation functions but different input and output values. If the circuit is specifically designed for convolutional neural networks, multiple convolutional kernels can be computed simultaneously by sharing the weight current modulation control module (such as the weight fast access unit). This does not require additional weight current modulation control modules or related chip resources.
[0167] The number of fast access units for weighted terms in the current modulation control module can be multiplied according to the demand for parameters, which facilitates rapid batch-by-batch parameter replacement, thereby increasing the flexibility for high-frequency calculations.
[0168] like Figure 12 This embodiment features a double design with double the input current modulation control module (120) (or the neuron basic computing unit has two input current modulation control modules). It can generate two sets of input PWM based on the two input fast access units: a first input PWM (121) and a second input PWM (124). Then, based on the two input PWMs, it selects or generates new modulation current signal waves from the first and second bottom-level high-frequency reference PWMs, i.e., simultaneously outputs the first selected reference pulse signal (122) and the second selected reference pulse signal (123). Because the new modulation signal waves selected or generated by the first and second bottom-level high-frequency reference PWMs are sparsely staggered in pulse phase (the input current modulation control module receives two sets of bottom-level reference pulse signals with different phases sent simultaneously by the shared pulse transmitter), the pulse signal phases of the first and second selected reference pulse signals are also staggered. Then, the multi-line channels (125) and (129) of each of the first neurons, which point to the corresponding weight term current modulation control modules, simultaneously send the first selected reference pulse signal (122) and the second selected reference pulse signal (123) (the symbols also have corresponding circuits). The weight term current modulation control modules insert delays and adjust the instantaneous current through the resistor network. As mentioned above, the weight values are shared, so the weight term fast access unit and other parts can be shared. The internal open CMOS switches and other parts can also be partially shared through optimized design. After inserting the delay, the first weight term current modulation control module generates two sets of modulation signals, the first modulation signal and the second modulation signal. The two sets of signals control the first set of positive and negative switches (131) and the second set of positive and negative switches (138) of the first neuron, respectively (the positive and negative are distinguished according to the product of the input and the weight). In addition, the signal sent by the second input term current modulation control module is processed by the second neuron weight term current modulation control module (127) to generate two sets of modulation signals, which control the end of the second set of positive and negative switches of the first neuron and the second set of positive and negative switches (130) of the first neuron, respectively.
[0169] That is to say, the switch of the modulation current in the above current channel generates the modulation current for charging and discharging the corresponding positive and negative capacitor modules (132) and (133) according to the modulation signal. Since the transmitted signals are all pulse signals, the single-chain weight term current modulation control module (126), the first neuron weight term current modulation control module (128), and the second neuron weight term current modulation control module (127) all have channels for the driving voltage source of the capacitor module charging and discharging (discharging to ground or charging to a specified voltage). Since the pulse signals received by the weight term current modulation controller are sparsely staggered, it is possible to generate two sets of modulation currents at the same time. As shown in the figure, the neuron output module (134) and (135) and the positive and negative capacitor control module are multiples.
[0170] The number of neurons (136)(137) in the normalized network summation layer is multiplied. This allows for simultaneous forward propagation to two network layers. This means that a series of computational operations, such as charging / discharging, detecting discharge, and activation, are performed, effectively doubling the computational speed.
[0171] In this embodiment, the multiple and the number of neurons are both 2, which is only for illustrative purposes. In actual applications, the number can be expanded according to the needs.
[0172] An Example of a Delayed Single-Layer Dynamic Network with Segmented Multiline PWM* In the previous example, the multiline modulation signal at the input was used in different neurons (although the weight register was shared). In this example, the multiline modulation signal is used in the same neuron, which can significantly increase the overall operating frequency of the neural network.
[0173] The input fast access unit is divided into multiple segments bit by bit. For example, if the input fast access unit is 16 bits, it can be divided into high 8 bits and low 8 bits (two segments), or it can be divided into high 4 bits, mid-high 4 bits, mid-low 4 bits, and low 4 bits (four segments). Each segment generates PWM independently.
[0174] like Figure 13In this embodiment, the input fast access unit is designed to be divided into two segments: a high 8-bit segment and a low 8-bit segment. The two segments generate PWM independently, namely, a high 8-bit PWM (140) and a low 8-bit PWM (144). Each of the two PWMs is combined with the underlying reference pulse signal (142) and (145) to synthesize or select a new modulation signal pulse signal, namely, a high 8-bit selected pulse signal (142) and a low 8-bit selected pulse signal (146). Both selected pulse signals are simultaneously sent to the connected weighted current modulation control module. Each connected weighted current modulation control module inserts a delay into the two selected pulse signals according to the value of its internal weighted fast access unit, thereby generating the final modulation current switching signal, namely, a high 8-bit modulation current switching signal (143) and a low 8-bit modulation current switching signal (147). The two signals independently control the switching of the modulation current for charging and discharging the positive and negative capacitor modules, namely, positive and negative high 8-bit switches and positive and negative low 8-bit switches. The sign of the product term (input value and weight value) determines whether the capacitor module is charged or discharged positively or negatively.
[0175] If a higher frequency or shorter pulse time is required for the underlying reference pulse signal, a common multi-line multi-phase clock signal can be used (to generate a higher resolution PWM), or a delay insertion method can be used (using delay circuits based on inverters, etc., to generate a modulated current signal such as a single pulse or PWM). It should be noted that the actual frequency and pulse width of the underlying reference pulse signal are not limited to those shown in the figure, and the current modulation is not limited to PWM.
[0176] In this embodiment, the current channels corresponding to the links of the positive and negative high 8-bit switches and the positive and negative low 8-bit switches have different resistance values. The ratio of their resistance values is consistent with the maximum value of each segment. For example, the resistance value of the low 8-bit switches is 2^8 = 256 times larger than that of the high 8-bit switches, representing a smaller 1 / 256th of the current. If the weighted current modulation control module includes an optional resistor network, that is, some bits in the weighted register represent the selection of the resistor network resistance value, then the adjustment of the resistor network resistance value, for the high 8-bit current channel and the low 8-bit current channel, will still result in a ratio of 256 times. This is determined by the number of bits in the input fast access unit. During charging and discharging, the current channels corresponding to the links of the high 8-bit switches and the positive and negative low 8-bit switches converge into a capacitor module of the same symbol.
[0177] The structure of this embodiment, including multiple proportionally sized capacitor modules for charging and discharging modulated current channels, multiple resistor networks, and multiple PWM modulation signals within a single neuron circuit, effectively improves the operating frequency. For example, with 100 neurons and 100 fully connected lines, if it operates at 100M times per second (under 16-bit PWM limitations), then ideally, the number of weight parameters calculated per second is 100*100*100M=1T. If it is divided into two equal segments, then the computing power is 1T*256=256T. Dividing the input register into four segments would achieve even higher frequencies.
[0178] Not only can the input fast access unit be segmented, but the weight fast access unit can also be segmented. Assuming it's also divided into high 8 bits and low 8 bits, if the input current modulation control module sends two selected pulse signals, then the weight modulation control module independently inserts delays in the high and low 8 bits, generating four switching signals from the two selected signals: high 8 + high 8, high 8 + low 8, low 8 + high 8, and low 8 + low 8. These four switching signals control four positive and negative switches (eight switches in total, including positive and negative), and each corresponding current channel has its own corresponding resistance value, i.e., resistances in the ratios of 1, 256x1, 1x256, and 256x256. However, because the number of neuron connections is large, i.e., the number of weight current modulation control modules is large, this part of the circuit needs to be simplified as much as possible.
[0179] The resistance error at the high level is not consistent with that at the low level because it is divided into multi-wire switches and resistors. This inconsistency in error will interfere with the accuracy of the low level. In addition to production and design optimization (such as having a correction circuit inside the circuit), this situation also requires statistical detection or circuit self-testing to generate a numerical conversion table to correct the input value and weight value, so as to obtain an accurate current.
[0180] In some other embodiments, not only can the input modulation control module and the weight modulation control module be segmented, but the capacitor module, capacitor control module, and neuron output module can also be divided into multiple segments. In this case, the charging and discharging input currents of each segment's capacitor module do not need to be aggregated; instead, they are independently input into the capacitor module and aggregated after the output time difference (using digital or other addition circuits). Regarding the activation function, the layer activation function source and the multi-bit incrementing counter can also be improved with reference to this segmented pulse signal design. The corresponding output values of multiple segments are summed, and the activation function value is obtained based on the summed value.
[0181] (The segmentation of the input item fast access unit may include, but is not limited to, the following forms: equal-weighted segmentation (such as every 4 bits or other numerical values), non-equal-weighted segmentation (high bit width is not equal to low bit width), dynamically adjustable segmentation (configured via registers), and the modulation signals generated independently by each segment can drive the capacitor module in parallel. Other embodiments are similar.) An Example of a Single-Layer Dynamic Network for Multi-Segment Multi-Line Conventional PWM* In the previous example, the weighted current modulation control module directly inserted a delay into the sent short pulse signal; this example uses a selection / AND gate method for multi-channel PWM to synthesize the signal, which is technically less difficult.
[0182] Assuming the total number of bits in the fast access unit is 4, divided into 2 segments, such as... Figure 14 A neuron is a basic computational unit. Its input current modulation control module generates A: high 2-bit input PWM (150) and B: low 2-bit input PWM (152) according to the input fast access unit segmentation (sign bit is calculated separately). At the same time, the base frequency of the weight current modulation control module is 1 / 4 of the base frequency of the input current modulation control module. The weight current modulation control module generates C: high 2-bit weight PWM (151) and D: low 2-bit weight PWM (153) according to the weight register (sign bit is calculated separately). The result of their respective logic AND combination is A and C combined (154), A and D combined (155), B and C combined (156), B and D combined (157), that is, (A*4+B)*(C*4+D)=16*A*C+4*A*D+4*B*C+B*D. The numbers in the formula represent the magnitude of the current. After the switching signal is connected to the four current switches, the resistance ratio of AC in each current channel is 1 / 16, the resistance ratio of AD is 1 / 4, the resistance ratio of BC is 1 / 4, and the resistance ratio of BD is 1.
[0183] If the total number of bits in the fast access unit is 16 (excluding the sign bit), divided into two segments, then its input current modulation control module generates A: high 8-bit input PWM and B: low 8-bit input PWM based on the input fast access unit segmentation (sign bit counted separately). Simultaneously, the base frequency of the weighted current modulation control module is 1 / 256 of the base frequency of the input current modulation control module. The weighted current modulation control module generates C: high 8-bit weighted PWM and D: low 8-bit weighted PWM based on the weighted register (sign bit counted separately). (A*256+B)*(C*256+D)=(256^2)*A*C+256*A*D+256*B*C+B*D. Therefore, after connecting the switching signals to four current switches, the resistance ratios of the product terms in each current path are 1 / (256^2):1 / 256:1 / 256:1.
[0184] If the total number of bits in the fast access unit is 12 (excluding the sign), divided into 3 segments, then (a*(16^2)+b*(16^1)+c*(16^0))(e*(16^2)+f*(16^1)+g*(16^0))=a*e*(16^4)+a*f*(16^3)+a*g*(16^2)+b*e*(16^3)+b*f*(16^2)+b*g*(16^1)+c*e*(16^2)+c*f*(16^1)+c*g*(16^0); After the switching signal is connected to 9 current switches, the resistance ratio of each product term in each current path is the reciprocal of the coefficient of each product term. If the switching frequency is very high, two sets of switches can be used alternately, i.e., 9*2 current switches.
[0185] Considering factors such as switching delay, resistance error, and trace capacitance and inductance, it is necessary to convert the actual pulse width or parameter values of the input and weight terms based on actual test statistics. This conversion will not affect the actual execution. In general, the connections between neurons achieve higher operating frequencies by performing parallel charging and discharging operations on the capacitor module through multiple current channels.
[0186] An example of fixed-point computation using a full-resistance network to decompose product terms by bit segment*. Using a resistor network for all weight terms allows for higher frequencies. Reducing the resolution of the input PWM ensures high accuracy with a resistor network (easier to eliminate various interference factors). Decomposing large-bit multiplications into multiple smaller multiplications provides a wider computational range, and also optimizes the speed and energy consumption of capacitor charging and discharging. If factorization involves multiple non-parallel calculations, time can be traded for chip space, because the increased frequency does not actually double the computation time with multiple calculations. Simultaneously, the reduced resolution also lowers the voltage difference during capacitor charging and discharging.
[0187] Assume a neural network has n inputs and n outputs, with each neuron having n links; input values X1, X2, ... Xn; weights (A11, A12... A1n), (A21, A22... A2n)... (Am1, Am2... Amn); output value Y1 = X1*A11 + X2*A12 + ... Xn*A1n, Y2... Yn; the product term of Y indicates the number of links in the neuron; assuming X1 is 16 bits, it is split into X11, X12, X13, X14 by bit; the weights are also split similarly, A11 is divided into A111, A112, A113, A114; A12... and so on; then the first data link of Y1 = (X11 + X12 + X13 + X14) * (A111 + A112 + A113 + A114); the first data link of Y1 = X11 * (A111+ A112+A113+A114)+ X12* (A111+ A112+A113+A114)+X13* (A111+ A112+A113+A114)+X14* (A111+ A112+A113+A114); If the input current modulation control module corresponding to X1 uses two channels of 4-bit PWM concurrently in two rounds, the first round is X11 and X12, and the second round is X13 and X14; each channel is associated with four segmented weight values A111, A112, A113 respectively. A114 performs multiplication, yielding four product values (shifting is not considered here, but will be addressed later). The sum of the shifted product terms represents X1's independent contribution to the output of the first link of the total output Y1. The contributions of the outputs of the other links of the first neuron are calculated similarly. The sum of the product terms of all input links X1 to Xn is the output value of Y1. Two operations with two 4-digit numbers can calculate a 16-digit number; similarly, four operations with two 4-digit numbers can calculate a 32-digit number, and eight operations with two 4-digit numbers can calculate a 64-digit number (integer or fixed-point number).
[0188] This embodiment is applicable to fixed-point numbers of any bit length (if the bit length is sufficient, the fixed-point number can be converted to a floating-point number). However, in the following design, the input item fast access unit value is 16 bits, divided into 4 segments, such as... Figure 17In the first round of forward propagation, the first input current modulation control module (170) generates the high 4 bits (171) and the middle high 4 bits (172), the second round middle low 4 bits and the second round low 4 bits. Here, the high and low bits are determined by the binary bit representing the maximum value. The weight item fast access unit also has 16 bits, divided into 4 segments: high / middle high / middle low / low, which are respectively applied to the first input high 4 bit weight item current modulation control module (173), the first input middle high 4 bit weight item current modulation control module (174), the first input middle low 4 bit (175) weight item current modulation control module, the first input low 4 bit weight item current modulation control module (176), and the second input high 4 bit weight item current modulation control module, the second input middle high 4 bit weight item current modulation control module, the second input middle low 4 bit weight item current modulation control module, and the second input low 4 bit weight item current modulation control module. The input current modulation control module corresponding to the second input PWM cannot be drawn due to image size limitations, but its existence should be easy to understand. Since the input value has two segments in one round and the weighted value has four segments, each neuron of the weighted current modulation control module should have eight. The dot symbol (184) in the figure indicates that there are other weighted current modulation control modules that are not drawn but actually exist.
[0189] Since each segment of the weighted terms has only 4 bits, or 16 values, it can be easily implemented using a resistor network. The input PWM should adopt a waveform center-aligned mode, so that each pulse is centered, resulting in small interference error and achieving consistent and precise current control with fewer pulse repetitions (to accommodate changes in capacitor voltage during charging and discharging). In this round, the high 4 bits and the middle high 4 bits of the PWM are sent from the input multiple signal channel (177) to the 4*2=8 weighted term current modulation control modules mentioned above. In the neuron of this embodiment, the neuron output module, positive and negative capacitor module, and capacitor control module are also divided into 4*2=8 sets according to the number of segments, just like the weighted term current modulation control module. During forward propagation, they operate independently and synchronously according to each segment, and finally output the XOR time difference, storing the value and sign of the time difference in the fast access unit. The conversion of the time difference value is the same as in the previous embodiment, which uses the timing information of the layer activation function to convert the time difference value into a digital signal value in the fast access unit of the segment output module. However, the values of the segment output modules (179) corresponding to each segment in a neuron have bit differences. Because there are bit differences between the segments of the input value and the weight value, the calculation results naturally also have bit differences. Therefore, there is a neuron summarization calculation output module (180) to obtain the output values of each segment and perform addition with shift and sign. The figure also illustrates that the neuron summarization calculation output module can read or copy the segment output values from the data channels (183) of the eight segment output modules. In this embodiment, the function of a digital adder is obviously added. The operating frequency of the digital adder is much faster than the addition function through the time source of the layer activation function in other embodiments. At the same time, it has shift and offset functions, which can quickly realize certain linear transformation functions. The neuron summarization calculation output module, like the neuron output module in the previous embodiment, has various activation and normalization functions.
[0190] The second round of calculation is the same as the first, except that the PWM value represented by the input item is the lower 4 bits of the second round and the lower 4 bits of the second round. The neuron summation and calculation output module needs to sum the values from the two rounds in bit order to obtain the final accumulated value.
[0191] Regarding the activation function, previous embodiments used time differences to obtain the activation function. However, in this embodiment, the capacitor XOR time difference has been converted into the value of the fast access unit and summarized bit by bit, resulting in a complete digital signal. This embodiment uses a comparison-based method. It can be envisioned that the layer activation function source outputs (positive and negative paths) time values and activation function values (multi-line data, with the time values and output module values having the same number of bits). The time values of the layer activation function source are compared with the values calculated by the neuron output module; if they are equal, the activation function value is copied. However, this operation is not compatible with the operating frequency of other modules. Because bit segmentation reduces resolution, the PWM frequency is significantly increased, the voltage difference between capacitor charging and discharging is significantly reduced, and the speed is naturally significantly increased. Therefore, the envisioned activation function implementation method becomes a speed bottleneck. The solution is a piecewise linear interpolation method for digital circuits, matched with the number of calculation rounds, as follows: like Figure 18 The activation function curve is divided into multiple segments at equal intervals (assuming 10,000 segments), containing positive and negative activation function values (192)(190). Each segment has an initial value (191), which is the value at the intersection of the segment line and the activation function curve. The layer activation function source sends the segment time value, initial value, and derivative value (even the second derivative value) sequentially at a certain rhythm, sending positive and negative values simultaneously. If the aforementioned values are refreshed at a frequency of 1G, for 10,000 segments, the total time is 0.01ms, which is a relatively large delay, but acceptable. When the result of comparing the accumulated output value of each neuron's calculation output module with the sent segment time value is changed, the initial value and derivative value are copied down. Then, linear interpolation is performed. Its serial multiplier multiplies the derivative value based on the difference between its own accumulated output value and the segment time value, and the final result is the activation function value of each neuron. This module already has the function of shifting and adding, so adding the function of serial multiplier will not consume too many resources and time.
[0192] (Activation function value generation methods include: simple addition and subtraction counting based on data sent by the layer activation function source, direct copying, time-based accumulation / subtraction, pre-calculation lookup table (with pre-stored function-related values in the table), real-time calculation (CPU dynamic calculation), piecewise linear interpolation (output slope + intercept), second-order interpolation, etc. Derivative information can be obtained through multi-line signal circuits or tables with pre-stored derivative values. Other embodiments are similar. In the bit-segmented embodiment, the linear values of each segment calculation result can be obtained first based on the capacitor discharge time difference, and then the final activation function value can be obtained based on the combined segment output values through table lookup and interpolation calculation or other methods.) If necessary, this can be coordinated with the number of rounds of segmentation. In the first round of calculation, we have already obtained the high-order and mid-high-order calculated values. We extract the largest part, i.e., the highest few bits, which will not be affected by the accumulation in the second round. At this point, we can perform preliminary function segmentation, assuming there are n preliminary segments. The layer activation function source simultaneously distributes n sets of previous time values, initial values, and derivative values, i.e., subdivided segmentation or secondary segmentation. The neuron aggregation calculation output module selects the corresponding data channel of the preliminary function segment based on its own largest value part, and then copies the relevant subdivided segment values when switching according to the comparison results mentioned earlier. This can speed up the communication. Finally, linear interpolation within the segment is performed, i.e., serial multiplication, to obtain the final value. If there are more rounds of calculation and more total bits, the segmentation can be finer and more convenient; if there are fewer rounds and fewer total bits, the calculation can be performed after the total accumulated value is obtained. If multiple neural networks with shared weight parameters are computed simultaneously, as in the example of a high-frequency single-layer dynamic network with shared links*, then the activation function can be computed in a pipelined manner for each layer of the neural network. That is, although the computational latency remains unchanged, the computational power is not affected by the pipelined approach.
[0193] It can be seen that the functionality of the neuron aggregation and output module is already close to that of a very simplified microcontroller CPU, and further development could achieve more complex functions. Moreover, multiple neuron aggregation and output modules can share many internal circuits, as can other modules; many internal circuits can be shared and optimized. Simultaneously, multiple neural networks or multiple neurons in a single network can also share the CPU or the forward propagation link circuitry (i.e., a large number of weighted current modulation control modules). For example, the CPU can manage the selection and copying of data in various fast access units (registers or SRAM, etc.); it can directly convert time values to activation function values through memory lookup tables; it can linearly transform time differences through direct multiplication; it can manage the workflow sequence of multiple neural networks, neurons, or input / output modules using the CPU and program, optimizing performance based on the delay of each process / step; and it can manage communication between larger modules and communication within and outside the chip. We will only provide an overall description of the logic and functionality here; specific details will be omitted.
[0194] The circuit of this invention uses this factorization form to perform multiplication and addition operations. It needs to control segmentation errors (when shifting and adding, the relatively low bits of the higher-order segments become the higher bits of the final number, thus amplifying the shifting and adding errors of these relatively low-order segments). If the number of neuron links is small, such as in a convolutional neural network, then accuracy is easily guaranteed; however, if it is fully connected with many neuron links, then to ensure accuracy in the middle and lower digits, there are requirements for the voltage difference between capacitor charging and discharging and the voltage comparator of the capacitor control module, which must ensure the accuracy of the calculation results. Method 1 involves using different options for capacitor pre-charging and comparison voltage difference when performing segmented calculations in a fully connected neural network. Higher-order segments have larger voltage differences, longer PWM times, and lower calculation frequencies; lower-order segments have smaller voltage differences, shorter PWM times, and higher frequencies. This approach also improves the accuracy of the resistor network, enhances the voltage comparator, and reduces component leakage current to ensure controllable errors.
[0195] Method 2 involves changing the single-layer fully connected neural network into a multi-branch tree structure, i.e., a tree structure, and using linear activation functions in the intermediate layers (i.e., using linear activation functions at the multi-branch points). Method 3 involves calculating the multiplication and accumulation of a finite number of links each time (multiple sets of finite-numbered connections can be calculated simultaneously), and then summing the results at the end. No network design modifications are needed. Although it's a fully connected network, the output is summarized in the form of an addition tree. This is a multi-branch tree of addition (addition can be calculated using multi-layered networks in our invented circuits, or digital circuits, with the values of each output module used for data transfer and shifting between modules in a tree structure). This also brings the advantage of network flexibility; it can function as both convolutional neural networks and fully connected neural networks, with high accuracy under controlled conditions, and the total number of bits can be flexibly defined. Furthermore, this software design does not affect the circuit function of the basic computational units of neurons; it only depends on the number of neuron connections—that is, it's an adjustment to the computational program. However, when using Method 2, it's important to understand that in a fully connected network, the input terms are shared, but the computation is performed with a finite number of links; in a convolutional network, the input terms are independent, but the weights are shared. For fully connected networks, such as... Figure 19 / Figure 20 / Figure 21The segmented neuron input current modulation control module group (200) with a limited number of members has neuron links, that is, it is linked to other modules (202) (205) of the output of a single neuron through the weighted current modulation control module corresponding to each neuron. These modules include capacitor modules, capacitor control modules, segment neuron output modules, and neuron summary calculation output modules. The multiple inputs and their multiple input terms PWM of the first weighted current modulation control module group are sent to the corresponding weighted current modulation control module of the neuron. The weights of each link in each group are different. For example, the weights of each segment in the first link group (201) are different from those in the second link group. After two rounds of calculation, it is the turn of the second weighted current modulation control module group (204) to perform calculation, and then the third, fourth, and so on. With 4 input groups, each group containing n inputs, and each input undergoing 2 rounds of calculation in 2 segments, the current modulation control module is divided into 4*2=8 weighted terms across 4 segments. Through 16 capacitor modules (positive and negative 8x2=16) and a capacitor control module, the final result is obtained through a summation calculation output module (one neuron per group), totaling 4 neuron summation calculation output modules. This completes the calculation for that neuron block. Afterwards, the input values of the first neuron block (210), the second neuron block (207), and subsequent neuron blocks (208) are cyclically transferred in their corresponding fast access units (206). This calculation is repeated, with the final result based on the previous calculation plus the accumulated value from subsequent calculations, until the loop ends. The final summation value of each neuron summation calculation output module, after activation, is the final output value of the fully connected layer.
[0196] Because the input is also 4-bit segmented, the generated PWM only has 16 values, which has a very low resolution. Therefore, it is convenient to use the method of inserting delays by circuit elements to generate PWM instead of according to the base frequency, so that a higher frequency PWM can be obtained.
[0197] This embodiment can simultaneously compute convolutional and fully connected networks through more complex transfer mechanisms. The ratio of neuron inputs to outputs and the number of groups / segments can be set according to actual needs. Directly computing activation functions using a high-performance general-purpose CPU is also a feasible solution, suitable for mixed-precision computation and network training.
[0198] To illustrate the characteristics of arbitrary number of digits, such as Figure 15In this section, the fast access unit has 8 bits of data, divided into two segments: the high 4 bits and the low 4 bits. The first input current modulation control module (162) generates two PWM signals: the first calculated high 4-bit input PWM (160) and the low 4-bit input PWM (161). The two PWM signals are sent separately: the high 4-bit input PWM is sent to the first neuron first chain weight current modulation control module (164) and the second neuron first chain weight current modulation control module; the low 4-bit input PWM is sent to the first neuron second chain weight current modulation control module (165) and the second neuron second chain weight current modulation control module.
[0199] The value of the weight term fast access unit of the weight term current modulation control module is an 8-bit value. The resistance value of the corresponding resistor network is selected and generated, but it needs to be multiplied with the input term PWM. Therefore, the first neuron contains two sets of weight term current control modulation modules, positive and negative capacitor modules and capacitor controllers, segment neuron output modules, etc. The weight values of the two sets of weight term current control modulation modules are the same (multiplied with the two 4-bit input terms PWM respectively). The second neuron operates in the same way. This embodiment implements the discharge mode. The weight term current modulation control module itself has a grounding channel (163). Therefore, the channel from the input term current modulation control module to the weight term current modulation control module is the channel for PWM and other information and control. The channel from the weight term current modulation control module to the capacitor module is the actual discharge current channel. Then, the segment neuron output module generates the time difference value through the charging and discharging time difference and sign of the positive and negative capacitors, as well as the timing signal of the layer activation function source (168). The neuron summary calculation output module obtains the cumulative sum value through shift addition. If normalization is necessary, it is also performed through the network aggregation layer (169), and then normalization and secondary activation are implemented in the same way as other embodiments that require normalization. Here, the first neuron and the second neuron are the same, and the calculation of the 8-bit value can be completed in one round of calculation.
[0200] An embodiment of a multi-layer dynamic network with simulated intermediate layers*. Based on the above embodiment of a dynamic single layer, instead of calculating the weight parameters every time they are read from memory, there are multiple layers, such as... Figure 22a , Figure 22bData is processed like an assembly line between layers. There are two digital layers at the top and bottom, and one or more digital / analog layers in between. Although the computation frequency of each layer remains unchanged, multiple layers are processed at once, so the computational performance is multiplied. Moreover, the intermediate layer between the output of the current layer neuron and the input of the next layer neuron can be optimized. If no training is required and the accuracy requirement is not high (or the network is trained to be insensitive to accuracy), then the first and last two layers (385), that is, the input and output layers of the entire computing circuit, are digital, while the output of the intermediate layer (384) is the XOR single pulse (408) and the sign level generated by the capacitor discharge directly, and the input is a single pulse. Because the intermediate layer omits digital conversion, its operating frequency, area, and energy consumption can be further optimized. In addition to using linear pulse output, the activation function of the intermediate layer can also use other analog nonlinear outputs constructed by the principles of tp4-tn4, tp3-tn3, tp2-tn2, etc., which can limit the maximum output value range and ensure the stability of the single pulse width time range (see Figure 22d (Subsequent related embodiments). During external training, the aforementioned nonlinear function can be used for training, and fluctuation errors are added to the neuron outputs during training to adapt to production deviations in hardware parameters. The network has a residual calculation circuit (383) that can be matched with a residual network. A single intermediate simulation layer has a self-looping information channel (382) from output to input, and can also be used as a dynamic layer for single-pulse simulation, i.e., it can be used as different layers by changing parameters. Furthermore, the connection method of the intermediate simulation layer can be further as follows: Figure 22c There is a more complex network structure (386). The thick lines in the figure represent multi-line signals, that is, single-pulse signals or digital signals of multiple neuron inputs and outputs (380)(381). As mentioned earlier, the capacitance exponential multiplication and addition effect means that even if the single-pulse signals are out of order, the calculation results are not affected. The XOR single pulses and signs of multiple neurons output from multiple intermediate analog layers are sequentially and time-divided into the next layer at the node, so that the outputs of multiple layers can be aggregated into the input of the same layer. Of course, the transmission between layers requires a matching number of neurons. In the complex network structure of this embodiment, single pulses are used as input and output between the upper and lower analog intermediate layers. The working mode is as follows: like Figure 22aThe capacitor module contains two main capacitors (400) and three detection capacitors (401). Switches (402), (403), (404), and (407) control the connection relationships between each capacitor and the neuron's charging / discharging link node (405), the pre-charging node, and the capacitors. This embodiment has three parallel connection states: 1. (Pre-charging) Main capacitor + detection capacitor; 2. (Network discharging) Main capacitor + detection capacitor; 3. (Detection discharging) Detection capacitor. During the parallel processes of pre-charging and network discharging, a main capacitor and a detection capacitor need to be connected together. Pre-charging requires connection to the pre-charging port (408), and network discharging requires connection to the neuron's charging / discharging link node. There is also a parallel state where the detection capacitor works in conjunction with the capacitor control module and the neuron output module (411). Because multiple capacitors work in parallel, the single-pulse output representing the analog signal can be given to the input of the same layer. Furthermore, more detection capacitors and connection switches can be set, presenting different proportions of binary multiples of the detection capacitance value. By selecting the capacitance value, the binary multiple amplification rate of the detection discharge time can be controlled. That is, a shift operation on the XOR single pulse pulse width time output. The comparator of the capacitor control module can be set with multiple reference voltages (406), such as comparing the voltages of the positive and negative capacitors first, so that the control logic (413) obtains the positive and negative signs (410) before the XOR pulse output (409), instead of after. Similarly, the setting and selection of the resistance value of the detection resistor (414) can also control the binary multiple amplification rate of the detection discharge time. The circuit for generating the negative pulse (412) in the figure is similar to that of the positive pulse and is omitted. There is also a difficulty in the analog intermediate layer in chip manufacturing, which is that the analog intermediate layer is very complex to use bit segmentation technology. Therefore, if the binary weighted resistor network in the weighted current modulation control module has a large number of bits, the resistance value increases exponentially bit by bit. Then, resistors with different resistivity and corresponding different manufacturing processes are required. Thus, the resistance value in the chip has a very large error due to the manufacturing deviation of the resistivity. Under multiple tests with specific weight parameters and standard input, we can determine the error resistance and error resistivity of each bit by detecting and statistically analyzing the corresponding number of pulse widths (internal or external) of the linear output of the forward propagation of this layer network. Therefore, before using the weight parameters, we can use the main CPU or external calculations to correct the weight parameter values based on the detection results. Note: Considering chip / circuit area limitations, the functions of the digital and analog layers can also be combined into a single dynamic layer that can be used for both analog and digital applications.
[0201] This embodiment is the best implementation method for neural network inference computation. For specific details, refer to other embodiments according to design requirements.
[0202] Example of training a backpropagation network using hybrid CPU computation* Backpropagation and forward propagation both involve the same process of accumulating product terms. The inverse activation function is a function related to the original activation function and its derivative. It's necessary to simultaneously perform the accumulation calculation of product terms during backpropagation, calculate the change in weight terms, and adjust the values. Backpropagation from the current neuron output value to the weight terms typically requires calculating the partial derivative of the weight terms in conjunction with the learning rate. Other particularly complex backpropagation methods can be found in the following calculation process, using the CPU and the circuitry of this invention for collaborative computation. Basic calculation of backpropagation: [Partial derivative of weight term (single link) DW] = [Partial derivative of neuron output (error gradient) DO] * [Derivative value of activation function DF] * [Input value of neuron link DX]. The learning rate η can be adjusted either during the total loss / error calculation or during parameter updates, depending on the chosen parameter adjustment scheme. (The calculation of second-order backpropagation and second-order partial derivatives for learning rate adjustment during parameter updates, or other adaptive learning rate optimization algorithms, are similar. If the calculation process is too complex, it can be handled by the main control CPU; details are omitted here.) The propagation process for the neuron output partial derivative DO is similar to that of forward propagation, only reversed. The neuron link input value DX is the output value of the previous layer's neuron cached during forward propagation. The product of the neuron output partial derivative DO, activation function derivative DF, and neuron link input value DX should be calculated during backpropagation at this layer. This saves data caching because during forward propagation, only the neuron output partial derivative DO, activation function derivative DF, and neuron output value need to be cached. The previous example of bit-by-bit segmentation has already been described, as it incorporates a shift adder. Following digital circuit principles, multi-cycle multipliers are also based on shift adders; single-cycle tree multipliers also utilize shift adders. In this example, non-multiplicative-accumulator multiplication calculations and function lookup interpolation calculations use digital multipliers, while bit-by-bit segmentation summarization in multiplicative-accumulator calculations also utilizes a shift adder—that is, part of the CPU's computational function (shared by multiple neurons).
[0203] The parameter update method used in this embodiment is mini-batch training. This involves backpropagating multiple sets of independent data to calculate the adjusted weight parameters, caching them first, and then updating the weights after summing and averaging. This embodiment utilizes multiple circuit blocks working collaboratively, with each local circuit operating only dynamically at a single layer. Taking a 64x64 dynamic layer as an example (the number of neurons in the hardware can be expressed in more layers through software control and composite calculations), it has 64 neurons, each neuron has 64 connections, the weight parameters are 64*64 units, and the neuron output value is 64 units. If the number of mini-batch training batches is 64, then the required cache size for the neuron output values is 64x64 units.
[0204] like Figure 25This is the backpropagation data stream of a certain neuron in the backpropagation calculation process. A certain neural network has multiple sets of data, each of which performs first-order backpropagation independently. The partial derivative of the neuron output to a certain neuron in the previous layer, DO, is a multiplicative and cumulative calculation. For example, the backpropagation error data of the first set (320) -> the backpropagation error data of the first set of the next layer (323), the backpropagation error data of the second set (321) -> the backpropagation error data of the second set of the next layer (324)... the backpropagation error data of the fifth set (322) -> the backpropagation error data of the fifth set of the next layer (325)... up to the nth set of backpropagation error data -> the backpropagation error data of the nth set of the next layer. The vertical dashed line represents the n independent backpropagation neuron output values o1, o2... o5... on of a neuron in the next layer; the n independent activation function derivative values j1, j2... j5... jn of this neuron (derived from the independent output values of the forward propagation neurons); and the cached output values x1, x2... x5... xn of the second layer below corresponding to the input value of a link of this neuron. First, we use a digital multiplier to calculate o1*j1, o2*j2...o5*j5...on*jn, which equals k1, k2...k5...kn. Then, we use the circuit of this invention to perform a multiplication-accumulation calculation, i.e., S(328) = k1*x1(326) + k2*x2... + k5*x5(327) + ...kn*xn. In this way, the adjustment value S of the weight parameter of a certain neuron and a certain link in the backpropagation calculation of a small batch of n sets of data is calculated.
New weight value Wnew
Old weight value Wold
Learning rate η
Old weight value Wold
Learning rate η
Old weight value Wold
[0205] like Figure 4This is a simple schematic diagram of a dynamic single-layer calculation circuit. The multiple multiplication and accumulation calculations mentioned above can be performed using this circuit. More complex functions can be combined with other disclosed embodiments. The input current modulation control module (306) has a right-hand signal path (309), a first PWM signal (307) segmented by bit, and a second PWM signal (308), which are sent in two rounds in parallel to multiple rows to the weighted current modulation control modules (305) of the corresponding rows of the array. Input data is read from a fast-access memory. The positive and negative capacitor modules (303) each have current channels (304) that are directed column-wise to multiple weighted current modulation control modules. In discharge mode, the charge in the capacitor module is released back to the common ground of the capacitor module through the multiple weighted current modulation control modules of the corresponding columns, as well as the PWM switches and resistor networks set therein. The resistor network can be set with multiple segmented paths; the circuit cannot be clearly contained in the diagram, so details are omitted. The weight parameters of the resistor network are pre-read from the memory module. There is a positive / negative capacitor control module (302) and a neuron output module (301). Then, as in other embodiments, the neural network data propagation process includes network discharge, detection discharge, and activation, and finally the data (300) is output to a fast-access memory module. The training of the neural network is backward propagation, so compared with forward propagation, the layers of the neural network corresponding to the input and output are reversed.
[0206] like Figure 24Taking the design of fast access array memory (SRAM, etc.) as an example, in order to achieve parallel read and write, if the number of data output by the neurons in the physical layer of the circuit is different from the number of data (bits or bytes, etc.) in a row of array memory, for example, the amount of data in a row of array fast access memory is several times the amount of data in the neurons, or, the output of a neuron in one round of forward and backward propagation is 8 bits, while the array memory is 32 bits, then the known processing is multiple read and write, bit masking / bit selection, selective writing / reading to a specific part of the data address in the selected row (316), and more complex cases may also involve a value distributed bit by bit in multiple rows, the details of which are omitted here. The key point is that there is a certain number of transpose read or block read circuits in the array memory (in the process of multi-chip collaborative computing, the single-chip array memory only needs to meet the computing needs of a small number of layers corresponding to the dynamic layer). In small batch training, it is assumed that the data of each layer in the forward propagation has been filled, for example, the data of several rows of data (313)(314) that are close to each other are the data output of the neurons of the same neuron in multiple batches of forward propagation in the same layer of neural network. The number of rows in this data block, without considering masking and selective read / write, is the number of training batches. The block and row selection module (317) has the function of enabling row or block selection; the column read / write and enable (311) has the functions of enabling, reading / writing by row, and reading / writing by data bit masking. The column read / write enable (312) and block selection (315) together enable the transpose / block read / write function, which can (once or multiple times) read and transpose each batch of data in parallel and apply it to the corresponding neuron input current modulation control module (column to row within the block), that is, perform multiplication and addition calculations according to the data flow process and order described above, that is, read the data into the input of the dynamic layer in sequence according to the requirements or program settings, and save or output the layer output results to the outside. Finally, the entire backpropagation process is completed.
[0207] An example of detecting accelerated discharge* shows that the charging and discharging of neural network links has great potential to increase the frequency, but the comparator of the capacitor control module has a limited operating frequency. At high frequencies, the smallest unit of delay chain time conversion cannot match the resolution of capacitor discharge time detection (the time resolution requirement of the link pool node is much higher than that of the input and weights). As a result, the capacitor, capacitor control module and neuron output module need to be matched with a higher ratio to balance the operating frequency. The neuron output module is already complex, so this takes up a lot of chip area.
[0208] The discharge detection process can be accelerated by adjusting the resistance value of the discharge detection resistor. The principle is as follows (V0 > Vb > Vc): In the first case, there is no acceleration. A capacitor C, with an initial voltage V0, discharges to ground through a resistor R. Discharge stops when the voltage drops below Vc. The detection time is T. In the second case, there are two discharge resistors, R and R / k2, which start discharging in parallel simultaneously. When the detection voltage Vb is reached, resistor R / k2 stops discharging. When the detection voltage Vc is reached, all resistors stop discharging. Therefore, two discharge times, Tb and Tc, are obtained. T can be expressed using Tb and Tc: T = k2 * Tb + Tc. Derivation: Tc = t0 + t1, Tb = t0. The capacitor discharge voltage V(t) = V0 * e^(-t / (RC)), and the two-stage discharge V(t1) = V(t0) * e^(-t1 / (RC)) = (V0 * e^(-t0 * (k2+1) / (RC))) * e^(-t1 / (RC)) = V0e^(-(k2 * t0 + t0 + t1) / (RC)). Therefore, T = k2 * t0 + t0 + t1, T = k2 * Tb + Tc. Based on the property of exponential multiplication, it can be deduced that the relative delay or time shift of the discharge of the two resistors and their respective lengths do not affect the formula. We construct one (or more) reference neurons (352), control their inputs, let the single-pulse time standard output of their single-symbol capacitor control module be Tu, and let the number of parallel detection discharge resistors be j, then the actual pulse width time is Tu / j. This time is used to control the discharge time of the preceding resistor R / k2, i.e., Tb = Tu / j, T = k2 * Tu / j + Tc. Assume the target (binary number) T = 0B1001. If Tu = 0B0010, then control the switching of the two resistor networks, setting k2 / j = 4, so Tc should be 1. Tc is the actual time value obtained from the accelerated discharge detection in case two. T is the value we actually need to calculate; what we actually need to do is calculate T using Tc and the preset time Tu. Vb is a concept we set for ease of understanding; in reality, there should be a Va greater than Vb. We check if T is greater than 0B1000 + margin; if it is large enough, we execute accelerated discharge detection. Now, extending to the case of accelerated discharge in n stages, the actual discharge resistance of each stage is R / (1+k2+k3+ … +kn), … ,R / (1+k2+k3),R / (1+k2),R; similarly, it is easy to obtain T=kn*T1 / j1+k[n-1]*T2 / j1 … k2*T[n-1] / j[n-1]+Tn / jn.If the comparator determines that the current detection capacitor voltage is sufficiently large, then by using multiple standard neurons to generate standard output pulse widths (e.g., 0B000001, 0B000100, 0B10000...) through standard inputs, and controlling the number of parallel connections of kn, discharge can be accelerated in multiple stages with different limit safety resistance values. After obtaining the positive and negative time differences Tn, the logic circuit then adds back the predetermined standard value according to the sign based on the accelerated discharge operation of the existing positive and negative capacitors. This allows for obtaining the numerical conversion value of the time difference with high accuracy and frequency.
[0209] This embodiment uses a single-pulse pulse width-time input, such as... Figure 16The main capacitor (365) in the capacitor module can be turned on by one of the multiple detection capacitors (366) and (367) through switches (363) and (364). In some embodiments, the detection discharge can be operated in parallel with the pre-charging (361) or the charging and discharging process of the network node (362), increasing the overall operating frequency. In this embodiment, multiple detection capacitors are used independently for voltage comparison without interfering with each other, improving the accuracy and performance of the comparator, and setting the priority comparison order. The capacitor control module with the same sign has multiple sets of multiple voltage comparators (358) (different sets can be comparators of different types and performance), and each comparator can be connected to different reference voltages (359) through switches. During the design, the reference voltage can be calculated back from the binary discharge time of the capacitor voltage to determine whether the initial voltage is greater than a certain voltage corresponding to a binary output pulse width value. The first set of comparators and the first spare detection capacitor can obtain the approximate voltage and sign before the second set of operations, and submit the results to the control logic (360) and the control module (357). The second set of comparators and the second spare detection capacitor can be compared with multiple reference voltages in real time during the actual detection discharge process. This design enables the control logic (360) (357) to determine the minimum resistance or maximum discharge current of the resistor network (356) during each accelerated discharge stage of the detection discharge process, accompanying changes in the voltage of the detection capacitor before or during discharge. The comparison results can also be used to assist other circuits. Parallel resistors (355) (354) with discharge switches (356) are typically used conveniently for the resistor network. The operation time for detecting each resistance value of the accelerated discharge comes from the single-symbol pulse-width time output (353) of the reference neuron. The reference neuron can have some components removed as needed compared to the normal neuron. The reference neuron (352) has a structure roughly the same as other neurons, with a capacitor module (351) and the same input link nodes (350), but it does not require positive or negative capacitors (positive and negative can be set to reduce errors), and it directly outputs a single pulse from the single-symbol capacitor control module. Its neurons use independent standard inputs and standard weights. There are multiple reference neurons, each corresponding to a different common standard pulse-width time output. The output of the single-symbol capacitor control module of the neuron, after being adjusted by the control logic based on the pulse width and time of the reference neuron's output (i.e., after accelerated discharge), is then, as in the previous embodiment, calculated to obtain the positive and negative time difference, i.e., the digital value of the XOR pulse width. Then, according to the accelerated discharge strategy, the known preset time value that was reduced during the acceleration process is added back. If necessary, activation transition is calculated, and finally, the output value is generated.
[0210] For analog layers with single-pulse input and output, a multi-channel output accelerated discharge design can be used. This involves a capacitor controller with multiple independent resistor networks (equivalent to shift operations, corresponding to different bits) of varying resistance values for detection and discharge. Similarly, the neuron output module has multiple independent operation modules. The output pulses of the capacitor controller (with different shift operations) undergo positive and negative XOR logic synthesis, and finally, the pulses (with different shift operations) are output independently, meaning multiple lines (with different bits) synchronously output single pulses. Correspondingly, each input item and neuron connection in the next layer also has multiple independent current channels (and independent equivalent resistances) to cooperate with the operation of the current layer. The resistor networks in the next layer corresponding to different shift positions for input / weight / detection processes can have their resistance values or time lengths modified as needed to perform reverse shift corrections, or pulses of different bits can be operated in separate rounds. The so-called round operation means that the output pulses of the upper layer (equivalent to different shifts) are output in batches, and the weighted resistor network of the link current channel of the lower layer is also shifted according to the batch or read from the fast access memory and reloaded after shifting.
[0211] A more complex example of a nonlinear activation function for analog output*, when using a single pulse width time as the analog output, if it is a linear activation function with an output range like ReLU6, then this invention only needs to simply mask the negative time difference pulse output portion. More complex function outputs can be achieved through a two-stage detection charging and discharging process. That is, using the single pulse generated by the time difference level detected by the positive and negative capacitors, the (other) detection capacitor is recharged and discharged through a standard equivalent resistance (only one sign is needed, one of the positive and negative capacitors), and then a new detection charging and discharging process is initiated, generating a new single pulse corresponding to the activation function output value that conforms to software conventions. tp4-tn4, tp3-tn3, tp2-tn2, etc., have been proven through formula derivation. By setting the charging and discharging steps and adjusting the voltage parameters, simple capacitor charging and discharging can generate various forms of nonlinear analog single pulse activation functions. Furthermore, segmented activation functions can be generated through voltage comparison logic and combinations of charging and discharging of multiple detection capacitors. Based on the exponential multiplication characteristic of capacitor charging and discharging (e^a)*(e^b)=e^(a+b), even if there are truncations or out-of-order sequences between the segments of the output piecewise function single pulse, it will not affect the subsequent calculation results.
[0212] like Figure 22d, taking tp2-tn2 as an example, tp2-tn2 = cap*R*(ln[2-e^(-t*SSP(...,Xn*Wn))] -ln[2-e^(-t*SSN(...,Xn*Wn))]), if there is no input for SSN, then tp2-tn2 = cap*R*(ln[2-e^(-t*SSP(...,Xn*Wn))]. This is a function similar to tanh, which passes through the origin and has an origin derivative of cap*R. Let x=-t*SSP(...,Xn*Wn) and a=cap*R; construct f(x)=a*x, x<=b;h(x)=a*(ln[2-e^(-(x-b)))]+a*b,x>b; these two functions can form a piecewise function g(x) = IF[x<b, f(x), h(x)] in the first quadrant, and through selection of sign settings, there can be origin-symmetric curves in the first and third quadrants, or coordinate axis-symmetric curves, or curves in a single quadrant can be masked / selected through sign setting (450).
[0213] as Figure 26The piecewise function hardware circuit includes a more powerful neuron output module (430). It includes the function of generating time pulses by charging and discharging the original integrated positive / negative capacitor control module, that is, outputting XOR time difference pulses (434) to the signal synthesis selection logic circuit (439). It also includes a secondary charging and discharging conversion circuit, the working process of which is: when XOR is logically 1, that is, one of the positive and negative capacitors has stopped discharging, and the remaining positive / negative capacitor is still discharging (discharging mode). The positive and negative capacitor control module has multiple voltage comparators for voltage comparison of multiple standard reference voltages (435), one of which is the standard reference voltage of the piecewise function division point (greater than the comparison termination voltage), which divides the discharge process of the remaining positive / negative capacitor into two parts, linear segment discharge and nonlinear segment discharge. Of course, if when XOR is switched to 1, the voltage of the remaining discharging positive / negative capacitor is less than the standard reference voltage of the piecewise function division point, then it only discharges in the linear segment. That is to say, the comparison result (438) with the standard reference voltage is also sent to the signal synthesis selection logic circuit in real time. The signal synthesis and selection logic circuit, based on the comparison result, divides the XOR time difference pulse into pulse time outputs (442). The linear segment outputs directly, while the nonlinear segment controls the connection of the charging and discharging switch (assuming pre-charging has been completed) to the activation capacitor (432) and the charging and discharging voltage source (431), and performs nonlinear conversion through the activation capacitor control module (445). Similar to the standard positive and negative capacitor process, this stage also has a reference capacitor (433) and a reference capacitor control module to counteract various effects such as leakage current and charge loss injection. The time difference (444) (443) between the two in the nonlinear stage is also input to the signal synthesis and selection logic circuit, which then outputs the pulse time (442). As mentioned in other previous embodiments, due to the calculation process of the exponential multiplication of capacitor charging and discharging, the division and disorder of single-pulse charging and discharging do not affect the calculation of the time difference result of dual-capacitor discharge detection. Therefore, the pulse time outputs of the linear segment and the nonlinear segment are equivalent outputs of the piecewise function pulse time. Similar to other embodiments, this neuron output module also outputs a positive or negative level (441) based on the comparison between the positive and negative capacitor voltages. Mathematically, the mathematical relationship between the values before and after activation can be chosen similarly. Figure 22d The forms of tp2-tn2 and symmetric variants based on sign settings in multiple quadrants. Note: For large-scale model applications, single-clock simulated outputs can also use additive attention methods to replace the attention calculation process based on digitization conversion and matrix dot product in the previous examples.
[0214] Furthermore, because the key circuit is an analog signal circuit based on single-pulse pulse width time calculation, it cannot adjust the positive and negative deviations (originating from deviations such as production, voltage, and temperature) through the results of digital time conversion and numerical detection deviations (stored in registers, etc.) as in other embodiments. However, it can adjust the positive and negative deviations using positive / negative input adjustment nodes (436). The input adjustment node and the positive / negative current channel node (437) of the neuron link input are the same circuit (similar to the neuron bias term, multiple links can be set according to adjustment needs).
[0215] Similar to a biological spiking neural network*, biological neural networks are constantly operating. Charging current and leakage current simultaneously act on the neuron capacitors. When the neuron capacitor voltage exceeds a threshold, it is activated, and the output is determined by the frequency of the activated neuron's output and the number of excited neurons. Our circuit, after modification, can achieve the same function. It compares the voltages of the positive and negative capacitors, and the comparison result is output to the signal synthesis and selection logic. Simultaneously, the positive / negative capacitor control module no longer outputs discharge pulses but instead uses a predetermined voltage source connected to a large resistor as active leakage current. The neuron links charge and discharge the positive and negative capacitors respectively (in the opposite direction to the leakage current). When the positive voltage minus the negative voltage exceeds a specific voltage threshold, the circuit emits a fixed-width positive pulse signal (conversely, when the negative voltage exceeds a certain threshold of the positive voltage, a fixed-width negative pulse signal is emitted). If training makes the input range of the corresponding layer greater than or equal to 0, and only outputs a fixed-width positive pulse signal, then it is closer to a biological neural network. This signal can be connected to the single-pulse analog input embodiment disclosed herein (in the form of a ratio based on the total pulse length or the ratio of the number of capacitor handling operations, etc.), meaning that the input and output of two circuits with different operating principles can be connected (it can be used for a brain-computer interface).
[0216] The aforementioned circuit, which expresses output based on frequency modulation and the number of excited neurons, resembling a biological neuron, is inefficient compared to other embodiments using single-pulse pulse width-time simulation. It suffers from large errors, low operating frequency, high computational unit / resource consumption, weak activation function, and difficulty in training. Its only redeeming feature is sparse network computation; however, this function can also be achieved using gating circuits. The enable-type gating circuit described below differs from embodiments such as the gated activation function SwiGLU*, and does not generate a gating signal in subsequent calculations of detection and activation. Figure 26The positive and negative voltage comparator (446) compares the positive and negative capacitor voltages after the two devices of the neuron network are charged and discharged and before the detection phase begins. At high frequencies, it is actually synchronized with the detection, that is, it quickly obtains the sign of the neuron input multiplication and accumulation value for use as enable and gating. On the one hand, it can quickly enable the subsequent circuits such as internal (445), (444), and (431) to reduce energy consumption; on the other hand, it can be used as enable and gating (440) of external circuits. For example, the gating signal is output to other neurons to shield the output of external neurons or turn off their working state. Furthermore, the gating and enable signals can be stored in the fast access unit, and the working state of subsequent neurons can be controlled according to the read gating and enable signals.
[0217] A hybrid precision hybrid CPU collaborative computing embodiment* proposes a hybrid precision scheme to balance efficiency and accuracy: the neural network is divided into a small portion requiring high-precision computation (handled by traditional CPUs / GPUs / NPUs, etc.) and a large portion requiring low-precision computation (e.g., computations below 32-bit fixed-point numbers, handled by the differential capacitor circuit of this invention). The precision division is determined based on the network structure and experiments. Simultaneously, the circuit of this invention can also perform overall floating-point computation by shifting the entire layer through methods such as changing the detection equivalent resistance and detection capacitor or frequency conversion of the input portion.
[0218] In summary, the above embodiments have fully illustrated the implementation of neural network circuits, especially the hardware parallel computing implementation of various types of neural networks within currently popular large-scale models. The sub-modules within each module in the above embodiments represent equivalent functions and do not imply that sub-modules cannot be spatially or physically recombinated with other modules. In the above diagrams, thick lines represent multi-line bundles. Terms such as comparator and fast access unit in the above embodiments are only used to describe their circuit functions and do not specify specific circuit forms (because there are many implementation forms). The CPU mentioned in this document refers to a traditional program execution module and is not limited to a specific form or structure. The above embodiments have also fully illustrated methods for expanding neural networks, including inter-layer and intra-layer expansion, i.e., how to achieve large-scale parallel neural networks through simple replication and expansion. Therefore, the number of neurons should not be limited by textual and image descriptions. The specific calculated values can be easily adjusted according to the specific design through fundamental frequency and standard time t, resistance, capacitance, and even shifting. The circuit for synchronous charging and discharging of positive and negative dual capacitors proposed in this paper, if the positive and negative capacitors are operated in steps, each charging and discharging independently at different times, and the time difference is calculated at the end, then this degraded design should also be considered within the scope of this invention. It is known that various delay circuits, signal relay amplifier circuits, and adjustments to component and line parameters can be made to achieve the expected signal delay and strength. It is also known that multiple switches can be used to save energy, reduce leakage current, or improve driving capability. It is known that the transmission of various signals and pulses may use differential lines, lines with return paths, or other anti-interference designs. This invention is not limited to the application of certain delay, amplification, low-power design, anti-interference, or anti-static measures. This invention is based on circuit principles and functions and is not limited to specific layout designs or specific connection lines. This invention has various potential or mixed embodiments and is not limited to the content described in the partial text.
Claims
1. A neural network massively parallel computing circuit system, wherein the basic neuron unit includes: Input current modulation control module, weight current modulation control module, capacitor control module, capacitor module, neuron output module; The capacitor control module and the capacitor module each have positive and negative terminals; The positive capacitor control module has a circuit for charging and discharging positive capacitor modules, and the negative capacitor control module has a circuit for charging and discharging negative capacitor modules, and can independently charge and discharge the capacitor modules with the corresponding symbols to a specified voltage. There is a reference time source signal with one or more bits; The input current modulation control module generates an input current modulation control signal based on an external signal, a signal from a neuron output module, or a reference time source signal and the value of its loaded input fast access unit. The weighted current modulation control module further selects, trims, or inserts a delay to generate a control signal for the modulated current, or adjusts the equivalent resistance of the current channel, based on the value of its own weighted fast access unit and the corresponding input current modulation control signal. Multiple current channels with corresponding neural network links are capable of charging and discharging the positive and negative capacitor modules respectively according to settings. The current channel has a switch for generating a modulated current, which is controlled by the control signal for the modulated current to generate a modulated current calculated by an analog neural network for charging and discharging the capacitor module. The signs of the input item fast access unit and the weight item fast access unit together control whether the capacitor module is charged or discharged to a positive or negative value. The capacitor control module has a current channel for detecting charging and discharging, and there is an equivalent detection resistor in the channel, which can charge and discharge the corresponding capacitor module through the equivalent detection resistor. The capacitor control module can precharge the capacitor module to a specific voltage to prepare for the analog neural network calculation; after the modulation current of the analog neural network calculation charges and discharges the capacitor module, the equivalent detection resistor continues to charge and discharge the capacitor module. The capacitor control module contains a voltage comparator. During the energization of the equivalent detection resistor, the voltage of the capacitor in the capacitor module is compared with the reference voltage, or the voltage of the capacitor in the capacitor module is compared with the voltage of other capacitors of the same sign, generating a voltage comparison signal, and the current channel for detecting charging and discharging is turned off according to this signal. The positive and negative capacitor control modules output equivalent time pulses respectively, based on the equivalent time from the on to the off of the positive and negative detection charging and discharging current channels. The pulse width and time of the equivalent time pulse are, in pulse width current modulation, the ratio of duty cycle to the minimum duty cycle of the circuit; in single pulse width modulation, the ratio of pulse width time to the minimum pulse width time of the circuit; and in switched capacitor unit current modulation, the number of switching capacitor operations. The neuron output module generates the level of the equivalent time difference or its digitally converted value based on the positive and negative equivalent time pulses. The neuron output module generates a symbol level based on the comparison of positive and negative equivalent time pulses, or the comparison of the voltages of positive and negative capacitor modules.
2. The computing circuit system according to claim 1, characterized in that: The capacitor module contains a combination of capacitors and switches; the capacitor module includes a main capacitor and a detection capacitor; one of the main capacitors and one of the detection capacitors are charged and discharged together through a switch combination; the remaining disconnected or uncombined detection capacitors are charged and discharged independently or used for voltage comparison.
3. The computing circuit system according to claim 2, characterized in that: There is a multi-layer neural network, with the middle layer being an analog layer; the analog layer uses equivalent single-pulse pulse width modulation output; there is a multi-layer circuit for detecting charge and discharge, with the first layer generating a time difference pulse to control the charge and discharge of the detection capacitor in the next layer; The analog layer generates a linear or nonlinear activation output based on the selected charging / discharging mode and charging / discharging depth of the detection capacitor. The equivalent single-pulse pulse width modulation output is subsequently used as input to its own layer or to other layers that can receive the equivalent single-pulse pulse width modulation. In switched-capacitor type unit current operation, the equivalent single-pulse pulse width modulation is current modulation based on the number of operations.
4. The computing circuit system according to claim 1, characterized in that: There is a layer of activation function numerical source, which generates numerical information with equivalent time as the independent variable according to the setting of the activation function; The numerical information with equivalent time as the independent variable is the time value, the activation function value, the derivative value, the intermediate calculation value, or a combination thereof; The neuron output module has a fast access unit for activation functions; The neuron output module selects the numerical information of the corresponding symbol with equivalent time as the independent variable according to the symbol level. When the level of the equivalent time difference of capacitor charging and discharging begins to change or ends, the output value of the activation function of the corresponding neuron basic unit at the current time is obtained based on the numerical information of the corresponding symbol with the equivalent time as the independent variable.
5. The computing circuit system according to claim 4, characterized in that: The computing circuit system is dynamically used as one or more layers of the entire neural network to be computed; The input current modulation control module loads the value of the input fast access unit from the memory or external circuit, or shares or loads it from the neuron output module. The weighted term current modulation control module loads the value of the corresponding weighted term fast access unit from the memory or external circuit, or shares or loads it from the neuron output module. When the calculation is complete, the value of the activation function fast access unit in the last layer is output to the memory, its own input, external circuitry, or other modules.
6. The computing circuit system according to claim 4, characterized in that: The equivalent detection resistance of the capacitance control module of the neuron used for computation is an equivalent resistance network with multiple equivalent resistance values, which can accelerate the detection of charging and discharging. There is a reference neuron whose output is the pulse output of the capacitor control module, which is used to control the charging and discharging time of the equivalent detection resistor. The reference neuron can control the pulse output value of the capacitor control module according to the parameter settings; The neuron used for calculation has a capacitor control module with multiple reference voltages for comparison, and the voltage range of the capacitor detection capacitor is determined based on the comparison results. The control logic switches and selects the corresponding equivalent detection resistor value according to the real-time voltage range of the detection capacitor, thereby controlling the instantaneous current for accelerated charging and discharging. Simultaneously, the pulse output of the corresponding reference neuron is selected to control the duration of accelerated charging and discharging; The output of the neurons used for computation is obtained by accelerating the charging and discharging of neurons using positive and negative capacitor control modules to obtain equivalent time difference pulses. After the equivalent time difference pulse is digitized, the preset output value of the reference neuron is added to the sign to obtain the multiplicative sum value before the neuron is activated.
7. The computing circuit system according to claim 5, characterized in that: An array of memory is used as a cache; The array-type memory has a row-selective write circuit that writes the output of a neural network layer to the selected row; The array-type memory has a transposed block-by-block read circuit that writes the data of selected blocks in a column into the input or weights of a neural network layer.
8. The computing circuit system according to claim 5, characterized in that: There is one or more underlying high-frequency reference pulse signals; the input current modulation control module has a pulse selection or shielding circuit, which selects or shields the underlying high-frequency reference pulse signal according to the input current modulation control signal to generate a short pulse input current modulation control signal, or selects one of the underlying high-frequency reference pulse signals.
9. The computing circuit system according to claim 5, characterized in that: The input item fast access unit is segmented bit by bit, and each segment independently generates one or more segment input item modulation signals; The aforementioned fast access unit for weighted items is segmented bit by bit; Each segment of the weighted item fast access unit independently generates a multi-channel modulated current based on its own segment weight value and the corresponding segment input item modulation signal; The multi-channel modulated current simultaneously charges and discharges the capacitor module of the corresponding segment.
10. The computing circuit system according to claim 8, characterized in that: There are multiple neurons, and their weight term current modulation control module has a shared weight term fast access unit inside; Each neuron's input current modulation control module generates multiple sets of modulation signals based on the shared weight term fast access unit and its own input value. Multiple switches in the link current channel of the neuron generate multiple sets of modulated currents based on the multiple sets of modulated signals; Multiple basic neuronal units work simultaneously or sequentially.