A neural network data processing system and method based on a smooth dynamic piecewise activation function

CN122797633APending Publication Date: 2026-09-22黄自升
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610061944.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-10-02
Filing Date
2026-01-16
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

世界各国已经对此投入了千万亿的资金以期实现AGI,然而,当前的人工智能神经网络实现方式仍存在显著的局限性

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122797633A_ABST
    Figure CN122797633A_ABST
Patent Text Reader

Abstract

The application relates to the design of an activation function used by a neural network computing device in the field of artificial intelligence, and thus can optimize the model of the neural network, that is, optimize the data calculation and processing steps of the calculator (FPGA, CPU, GPU, TPU, Npu, in-memory calculation, analog calculation accelerator, etc.) in the training and actual inference application process of original vector / tensor data representing a graph sound meaning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the design of an activation function used in neural network computing devices in the field of artificial intelligence, and thus to the optimization of neural network models, namely, the optimization of data calculation and processing steps of calculators (FPGA, CPU, GPU, TPU, NPU, in-memory computing, analog computing accelerators, etc.) during the training and actual inference of original vector / tensor data representing the meaning of graphs. Background Technology

[0002] More than 20 years ago, the sigmoid scaling activation function of the form a*tanh(b*x), abbreviated as atanh, was already used in alphanumeric recognition networks such as LeNet-5.

[0003] With the rapid development of artificial intelligence (AI) technology, neural networks, especially ultra-large-scale neural networks, have demonstrated remarkable capabilities in various fields such as image recognition, natural language processing, and speech recognition. DeepSeek and others have proven that by employing specific techniques and optimization strategies, it is possible to train and apply large-scale models at a relatively low cost. Countries worldwide have invested trillions of dollars in this area in hopes of achieving AGI (Artificial Intelligence of Things); however, current methods for implementing artificial intelligence neural networks still have significant limitations.

[0004] While the ReLU function significantly improved computational efficiency in activation, it also led to training failures, neuron death, and oscillations in training parameters caused by sudden changes in derivatives at hard transitions. Furthermore, its range was unrestricted, and it required normalization / scaling of the neuron's output (ReLU value) using methods like normalization, resulting in high computational costs. While approximation functions such as Leaky ReLU, ELU, SELU, Silu, and GELU offer some optimizations, they haven't completely resolved the aforementioned problems.

[0005] Second, the activation functions used in training and inference are inconsistent. During training, perfect mathematical functions are used to improve accuracy, but in the inference process, degenerate or simplified functions are used to reduce the amount of computation, or low-precision functions or interpolation tables are used, or soft transitions in functions are changed to hard transitions, which causes a decline in inference quality or ability, making it inferior to that during training.

[0006] Third, the use of nonlinear and asymmetric functions leads to insufficient (or uneven) adjustments at the lower layers of the neural network during training, resulting in decreased training accuracy or preventing the training of deep networks. This necessitates the use of residual layers in segments at the network level, i.e., at the computational level, and the output of these residual layers requires standardization / normalization scaling of the overall output data. This increases the complexity of the network design and leads to too many macroscopic computational steps.

[0007] Fourth, for the reasons mentioned above, it is difficult to simplify accelerator circuits. Even if arbitrary activation functions can be executed through hardware lookup tables and interpolation, subsequent processing such as standardization / normalization scaling of activation function data still relies on general-purpose computing core circuits (CPU, GPU, etc.), which severely limits the development of computing accelerators.

[0008] Fifth, residuals, hc residuals, mhc residuals, and the standardization / normalization of output data before and after residuals, and even linear dimensionality expansion, are computationally complex, but they severely limit the overall capabilities and output quality of neural networks.

[0009] Despite decades of development and trillions of dollars invested globally, from small-scale models to large-scale ones, and various attempts, the problems remain unsolved. These issues indicate that existing technologies and methods are insufficient to fully meet the practical needs of ultra-large-scale neural networks, and innovative architectural designs and circuit implementations are urgently needed to overcome these limitations.

[0010] Humanity must not continue with these foolish designs and methods. Summary of the Invention

[0011] In view of the above challenges, this invention proposes a dynamic and variable activation function with a hybrid linear and nonlinear piecewise structure, and an optimized neural network structure, namely, an optimization of the neural network data computation and processing steps. It eliminates the need for output standardization / normalization scaling, significantly reduces the computational load of the subsequent nonlinear activation part after multiplication and addition, optimizes the neural network model structure, achieves greater model configuration adaptability, further reduces computational power requirements, decreases system latency, and improves the overall capability of the neural network and the quality and accuracy of the output results. This invention overcomes the prejudice that linear and nonlinear activation functions cannot be freely combined, ending a decade of misconceptions, and unifying and laying the foundation for the micro-construction of future neural network layers. Moreover, the activation functions used for training and inference, software and hardware, linear and nonlinear functions, and different value ranges can be unified. Furthermore, certain specially adapted hardware circuits can be used in neural networks, improving computational performance or energy efficiency by several orders of magnitude, i.e., thousands of times. Even without dedicated circuits, simply adjusting the activation function lookup table or calculation formula and values ​​to the scheme of this invention can optimize computational power consumption, increase the depth of trainable layers, and improve the quality of network output.

[0012] The present invention provides an artificial neural network data processing system, wherein a computing circuit performs computation of a multi-layer neural network, the input comprises multi-dimensional vectors / tensors of graphics / audio / semantics, and a feature vector / tensor as a result is finally output; wherein at least one layer of neurons of the neural network performs numerical conversion on neuron outputs based on a multiply-accumulate value of each connection of a previous layer according to an activation function; the activation function is a piecewise function; in the numerical conversion, the multiply-accumulate value is compared with a reference input value at a piecewise junction of the piecewise function, and a numerical conversion computing circuit, a computing subprogram or function calculation parameters corresponding to a corresponding function piece are selected according to a comparison result; for the piecewise function G(x), there is a linear segment c*x + v in the middle, and at least one end of the two ends has a non-linear segment; the non-linear segment comprises a sub-function F(x), if there is an upper non-linear segment, a region where x>k is c*u*F((-k + x) / u) + c*k + v, if there is a lower non-linear segment, a region where x<j is c*u*F((-j + x) / u) + c*j + v; the sub-function F(x) passes through the origin [0,0]; a first derivative value F'(0) of the sub-function F(x) at the origin is 1; the sub-function F(x) monotonically increases with decreasing increment in the first quadrant; the sub-function F(x) has an upper limit in the first quadrant; c, v, u, k, j are settable fixed values, or modifiable during training, or learnable parameters.

[0013] The present invention also provides that the sub-function F(x) is F(x)=tanh(x).

[0014] The present invention also provides that the sub-function F(x) is F(x)=q*(1 / (1+e^(-x))-0.5); q=4.

[0015] The present invention also provides that the sub-function F(x) for the upper non-linear segment is F(x)=ln[q+1-q*e^(-x)] / q; the sub-function F(x) for the lower non-linear segment is F(x)=-ln[q+1-q*e^(x)] / q; q is a settable parameter.

[0016] This invention also proposes that the piecewise function G(x) is single-quadrant, bisegmented, passes through the origin [0,0], and uses the origin as the starting point; it has a dedicated neural network computing circuit, whose activation-related part includes linear segment numerical transformation and nonlinear segment numerical transformation; the dedicated neural network computing circuit outputs an analog quantity through its activation-related part; the computing circuit can generate numerical transformations symmetrical to the first quadrant in other quadrants by changing the signs of the input and output; the transformation function corresponding to the nonlinear segment numerical transformation is the subfunction F(x); the unsigned analog quantity, which serves as the multiplication and accumulation value of the neuron link input, is compared with the reference input analog quantity corresponding to the intersection point of the piecewise function, and the nonlinear segment numerical transformation-related circuit is enabled based on this result, i.e., the nonlinear segment numerical transformation-related circuit is activated when the input value exceeds the linear segment; the computing circuit generates the linear analog quantity of the linear segment numerical transformation; the nonlinear segment numerical transformation-related circuit is activated to generate the nonlinear analog quantity of the nonlinear segment numerical transformation; the linear analog quantity and the nonlinear analog quantity are output independently through different lines or in different time periods.

[0017] This invention also proposes a method for switching between gradual hard transitions. Step 1: During training, set k, j, and u in G(x) as layer-shared parameters, and let k+u equal the upper limit value minus v, and j+u equal the lower limit value minus v; Step 2: Train the neural network for a certain number of rounds, and then adjust u and the corresponding k and j; Step 3: Repeat step 2 until training is complete; Step 4: Change the activation function G(x) to a hard transition, i.e., a fully linear piecewise function; Step 5: In the inference process, the trained model and parameters are used to change the activation function G(x) to a hard transition, i.e., a fully linear piecewise function.

[0018] This invention also proposes a method for calculating G(x): Step 1: The analog quantity x, which is the input of the piecewise function G(x), is divided into two analog quantities according to the corresponding input analog quantity x0 at the intersection point of the first quadrant of the piecewise function, namely, the linear analog quantity part x1 and the nonlinear analog quantity part x2, i.e., x = x1 + x2. When x>x0, the linear analog component is x1=x0, and the nonlinear analog component is x2=x-x0; When 0 <= x <= x0, the linear analog component is x1 = x, and the nonlinear analog component is 0; Step 2: The analog quantities of the linear and nonlinear outputs are calculated separately using analog circuits based on the linear subfunction y1=c*x1 and the nonlinear subfunction y2=c*F(x2). Step 3, the output of G(x) is the sum of the two analog outputs, that is, G(x) = y1 + y2.

[0019] This invention also proposes a method for training dynamic activation functions in the same layer: Step 1: Construct a neural network, which includes dynamic layers; Step 2: In the dynamic layer, v, k, j, u in G(x) are set as neuron-specific and trainable parameters, and k, j, u are set to establish a linear linkage relationship. Step 3: During the training of the relevant layers of the neural network, adjust v, k, j, u so that the range of values ​​of the activation function output of the gated layer neurons is differentiated into positive and negative activation, or no activation, or negative activation, or intermediate states.

[0020] This invention also proposes a method with dynamic gating: The linear linkage relationship described herein is attributed to a single trainable parameter, meaning that all other parameters depend on this single trainable parameter. The dynamic layer is a dynamic gated layer. The output of a neuron in the dynamic gated layer is multiplied by the output of a neuron in another layer before being input into the next layer.

[0021] This invention also proposes a method for constructing physical neural networks: Step 1: Based on the analog quantity conversion calculation function of the physical calculation circuit of the inference terminal that meets the requirements of the F(x) subfunction, construct the G(x) activation function; Step 2: On the training end, construct a multi-layer neural network model based on G(x) and complete the training. Step 3: In the physical computing circuit at the inference end, use the parameters of the multilayer neural network model constructed based on G(x).

[0022] The present invention also proposes a multi-core computing system; a shared memory containing a unified F(x) or F'(x) function and a shared query interpolation data table for the function; a multi-layer neural network system having a G(x) form function with different parameter settings other than u; the calculation of the G(x) form function with different parameter settings other than u uses the shared query interpolation data table. Attached Figure Description

[0023] Figures 1-12 For the function graph, Figure 13 This is a circuit diagram. Figures 14-17 This is a graph of the function. Detailed Implementation

[0024] The simplest piecewise tanh implementation: F(x) = tanh(x) (non-linear); The characteristics of F(x) are that the curve passes through the origin [0,0], the first derivative F'(0) at the origin is 1, it is monotonically increasing in the first quadrant with decreasing increments, and there is an upper limit in the first quadrant. The c, v, u, k, j below are function parameters of neurons in a neural network that can be set to fixed values ​​or modified and learned during training.

[0025] like Figure 1 The single-quadrant piecewise activation function y = G(x) = Segment 1: c * u * F((-k + x) / u) + c * k + vx>k Section 2: c * x + vx >= 0 Where v=0, k=1, c=1, u=1, this function is only in the first quadrant, 0<=x, and the output value range is [0,2]. The piecewise function is composed of two functions: the lower half is 800 y=x, the upper half is 801 tanh(x-1)+1, the intersection point of the splicing is 802, and it terminates at the origin 803.

[0026] like Figure 2 The full-quadrant piecewise activation function y = G(x) = Segment 1: c * u * F((-k + x) / u) + c * k + vx>k Section 2: c * x + vj <= x <= k Segment 3: c * u * F((-j + x) / u) + c * j + vx <j Where v=0, k=1, j=-1, c=1, u=1, the function is symmetric at the origin and its output range is [-2,2]; Segment 1804: tanh(x-1)+1; Segment 2: x; Segment 3805: tanh(x+1)-1; The intersection points 806 and 807 of the piecewise functions are the junctions of the linear function passing through the origin and the two nonlinear segments; like Figure 3 The full-quadrant piecewise activation function y = G(x) = Segment 1: c * u * F((-k + x) / u) + c * k + vx>k Section 2: c * x + vj <= x <= k Segment 3: c * u * F((-j + x) / u) + c * j + vx <j Given v=2, k=1, j=-1, c=1, u=1, this function is symmetric at the point [0,2], and its output value range is [0,4]. like Figure 4 The full-quadrant piecewise activation function y = G(x) = Segment 1: c * u * F((-k + x) / u) + c * k + vx>k Section 2: c * x + vj <= x <= k Segment 3: c * u * F((-j + x) / u) + c * j + vx <j Where v=0.5, k=0.25, j=-0.25, c=1, u=0.25, the function has a symmetric output value range of [0,0.5] and can replace the sigmoid function; it can also be used for gating.

[0027] like Figure 5 The full-quadrant piecewise activation function y = G(x) = Segment 1: c * u * F((-k + x) / u) + c * k + vx>k Section 2: c * x + vj <= x <= k Segment 3: c * u * F((-j + x) / u) + c * j + vx <j Given v=1, k=0.25, j=-0.25, c=2, u=0.25, this function is symmetric about the point [0,1], and its output value range is [0,2]. The following formula has been modified to include the parameter n: like Figure 6 The full-quadrant piecewise activation function y = G(x) = Segment 1: c * u * F((-k + x) / u) + c * k + vx>k Section 2: c * x + vj <= x <= k Section 3: c * n * F((-j + x) / n) + c * j + vx <j Where v=0, k=2, j=-0.5, c=1, u=1, n=0.25, the function is asymmetric, and the output value range is [-0.5,3]; segment 1 810 and segment 2 811 are asymmetric curves.

[0028] like Figure 7 The full-quadrant piecewise activation function y = G(x) = Segment 1: c * u * F((-k + x) / u) + c * k + vx>k Section 2: c * x + vj <= x <= k Section 3: c * n * F((-j + x) / n) + c * j + vx <j Where v=0, k=2, j=0, c=1, u=1, n=0.25, this function is asymmetric, the lower intersection point 812 coincides with the origin at [0,0], and the output value range is [-0.25,3]. Combination Figure 1-7 Clearly, this piecewise activation function exhibits smooth transitions between linear and nonlinear segments, with no abrupt changes or jitters in either the derivative or the function itself. Most notably, this piecewise activation function is highly flexible; by setting parameters, it can be modified and replaced by most existing activation functions. Furthermore, because it inherently has an output range, there's no need for normalization or standardization of the output values, thus saving significant computational resources in the normalization / standardization process. Additionally, these parameters can be set to be learnable to adapt to different input distributions, improving network capabilities and the accuracy of results. If the slope of the activation function is set to 1 (c=1) and it is set to be origin-symmetric (v=0), then the calculations of c multiplication and v addition in the function can be saved, and other parameters can be adjusted accordingly to reduce computational load. In inference computation, the activation function used for training is often changed from a more precise arc-shaped curve to a hard curve; this activation function brings greater flexibility to this operation and reduces inconsistencies between training and inference.

[0029] (The calculation of tanh and its derivative does not actually require repeated calculation of e^x; by sharing or using shared memory, it only needs to be calculated once.) Calculation steps: The calculation steps for the activation function G(x) and related multilayer neural networks by the computer hardware circuit are as follows: Step 0: Import the metadata of the image / sound / meaning of the feature to be identified / extracted, and organize it into a neural network input tensor / vector, which is a multi-dimensional intermediate data containing image, sound and meaning. Step 1: Calculate the multiplication and summation of each neuron's connection based on the intermediate data; Step 2: If other activation functions or special calculations need to be performed, skip steps 3-5 after execution. Step 3: Calculate whether the circuit belongs to a linear segment or a nonlinear segment; Step 3.1: Determine the sign of the input value. If the range of the input value does not include the origin, then it is necessary to subtract it from the midpoint first, and then determine the sign. Step 3.2: Select the reference input value of the corresponding endpoint of the linear segment according to the sign of the input value. The reference input value is obtained from the reference digital value of the digital register or the fixed circuit, or the reference signal of the analog circuit as required. Step 3.3: Based on the comparison between the input value and the reference input value, determine whether it belongs to a linear segment or a nonlinear segment; (Of the above steps, the second-best option is to obtain the judgment result by comparing it with the reference input values ​​at the two endpoints of the linear segment. However, the comparison process is relatively complex and consumes more energy than the sign determination.) Step 4: If it is a nonlinear segment, the calculation circuit performs f(x) calculation, and the output value after activation is equal to the value of G(x) (the above embodiment has tanh for scaling and translation). There are many other piecewise nonlinear activation functions similar to tanh. Their common characteristic is that at the origin, the slope is 1 and the value is 0. As the input increases, the slope decreases and gradually approaches 0, resulting in a limit to the output value. For digital circuits, this value is often obtained through table lookup interpolation. In analog computing circuits, the following embodiments of this invention propose using nonlinear incremental accumulation to obtain this output value. Step 5: If it belongs to a linear segment, the output value is equal to the original input, and the computing circuit copies the original input to the neuron output; Step 6: Parallel or cyclically repeat steps 1-5 to calculate the output of each neuron in the same depth layer; Step 7: Using the output of the previous layer as input, calculate layer by layer according to depth, repeating step 6, and finally obtain the target feature tensor / vector.

[0030] During backpropagation training, the calculation of the derivative function G'[x] of the activation function is similar: the operation is reversed, the derivative is calculated piecewise, and then the adjustment value is calculated, which is omitted.

[0031] Dynamic nonlinear computation: Specifically, in a computer system containing various computational circuits, data containing graphs, sounds, and meanings is read in or stored externally. After preliminary processing as needed, the data is organized into vectors of multiple specific dimensions. These vectors are then imported into a neural network. The computation of vector data within the neural network is performed layer by layer according to the network depth. In this layer-by-layer computation, each neuron independently calculates the multiplication and accumulation function xo = a * (w1 * x1 + w2 * x2 ... + wn * xn); where x1, x2, ... xn are the inputs to the activation functions of the neurons, and w1, w2, ... wn are the weights of the connections between the neurons. Then, the activation function G[xo] output of each xo is independently calculated, and 'a' is used to adjust or scale the range of values ​​for each weight, or to adjust the backpropagation parameters. The G[xo] output of each neuron in this layer is then used as the multiplication and accumulation input for the next layer. After layer-by-layer computation, the final output is a multidimensional tensor / vector that meets our needs.

[0032] In deep neural network computation systems, besides the fact that the G[xo] activation function itself is relatively computationally efficient (when c=1, v=0, u=1, the area near the origin is y=x), and the fact that it omits the output standardization / normalization step due to its output range, it is already superior to the current simplest activation function ReLU in terms of computational power alone; more importantly, after initial training, the model parameters quickly stabilize, and the adjustment amount becomes smaller. The input and output statistics of the piecewise function G(x) follow a sharp bell-shaped distribution. If G(x) is symmetric about the origin, then the axis of symmetry of the bell-shaped distribution is the y-axis at x=0.

[0033] In other words, most output values ​​are generated by the linear segment of the activation function G(x), with only a small portion being non-linear (tanh) and requiring computation. (The internal subfunction e^x of tanh performs exponential calculations.) Moreover, it is dynamic; whether different neurons perform non-linear calculations at different times is determined by the model's training parameters and the input values ​​of each layer.

[0034] The number of rounds of residual calculation and numerical standardization / normalization operations before and after the residuals can be omitted or reduced. Because the outputs of most layers in this network are linear functions, there is no issue of the network's lower-level parameters being unadjustable after nonlinear activation functions pass through deep networks. The proportion of linear neurons is higher than that of the residuals, thus saving computational resources on the residuals, and the backpropagation process will not be unable to adjust due to the multiplication of derivatives less than 1 in the nonlinear segments. Conversely, for shallow networks, the proportion of nonlinear activations is lower, resulting in better training effects and higher final network output accuracy (compared to a fixed proportion of nonlinear activations). The necessary proportion and number of nonlinear activation neurons are dynamically determined by the numerical values ​​of each neuron in the neural network.

[0035] Furthermore, in large models, the number of xn in the residual calculation xn+Hn(x) limits the dimensionality of the linked multidimensional tensors / vectors. Moreover, the fixed form of addition also reduces the combined expressive power of a single neuron channel, thus limiting the number of information channels. Sometimes, it's necessary to expand the dimensionality through matrix multiplication before performing residual calculations, such as hc and mhc residuals. These calculations are quite complex. Our approach eliminates the need for such dimensionality expansion calculations, or requires only a minimal number of residual calculation steps.

[0036] When interpolation is performed using lookup tables instead of general-purpose computation (CPU, GPU, etc.), computational speed / performance is generally independent of the shape of the nonlinear segment function curve and the calculation formula (accelerators like NPUs usually have built-in function lookup table interpolation circuits). However, the inference model is generated through training; training yields a more accurate model, resulting in higher inference accuracy. In terms of computational power, by reducing other residual and standardization / normalization steps, the overall computational power requirement is significantly reduced, at the cost of only a small proportion of nonlinear computation.

[0037] In summary, the computational process for extracting feature tensors / vectors from original image-phonetic data using neural networks based on this activation function is simpler, requiring fewer steps. This means the neural network model can be more concise and flexible, allowing for deeper neural network layers. The resulting feature tensors / vectors are of better quality and more accurate. It also offers greater capability with the same computing power and network size.

[0038] Dynamic Gating Implementation Example* Gating refers to multiplying the outputs of corresponding neurons in two layers before outputting to the next layer. The parameters v, k, j, c, and u in the G(x) function can be preset or trained and adjustable; they can be layer-shared parameters, neuron-specific parameters, or global parameters. As a gating mechanism, the v parameter can be set as a trainable parameter, allowing G(x) to shift up and down, transforming it from a symmetrical curve about the origin in the first and third quadrants to a sigmoid-like curve only in the first and second quadrants. In other words, training determines whether the activation function of a gated neuron is of type 0 / 1, type -1 and 1, or a mixture of both. We can design it to better suit our logic and habits, for example, with an absolute value of 1 for the limit and a slope of 1 for the middle segment, as follows: like Figure 14 F(x) = tanh(x), the full-quadrant piecewise activation function y = G(x) = Segment 1: c * u * F((-k + x) / u) + c * k + vx>k Section 2: c * x + vj <= x <= k Segment 3: c * u * F((-j + x) / u) + c * j + vx <j Where v=0, k=0.5, j=-k, c=1, u=0.5, the function is symmetric at the point [0,0], and its output value range is [-1,1]. like Figure 15 F(x) = tanh(x), the full-quadrant piecewise activation function y = G(x) = Segment 1: c * u * F((-k + x) / u) + c * k + vx>k Section 2: c * x + vj <= x <= k Segment 3: c * u * F((-j + x) / u) + c * j + vx <j Where v=0.5, k=0.25, j=-k, c=1, u=0.25, the function is symmetric at the point [0,0], and the output value range is [0,1]. These are two types of gating. Figure 14 It is a [-1,1] type gate control. Figure 15 It is a [0,1] type gating. Let a = 0.5*(1-v), k = a, u = a, and let v take values ​​in the range [0, 0.5], with the final value determined through training and unique to the neuron. That is, k, u, and j are all determined by a. Then v = 0 is... Figure 14 v=0.5 means Figure 15 When v is between 0 and 0.5, it is a mixed state because the values ​​of v, k, j, and u are all between the corresponding parameter values ​​of the functions in the two diagrams above.

[0039] If we let a = 0.5 * (1 - |v|), k = a, u = a, and let v take values ​​in the range [-0.5, 0.5], with the final value determined through training and unique to the neuron, then there are three variations: [-1, 1] gating, [0, 1] gating, and [-1, 0] gating. During training initialization, v can be randomly set to [-0.5, 0, 0.5].

[0040] The same-layer dynamic activation function* is similar to the dynamic gating implementation, but this implementation does not use the output for multiplication with the outputs of other neurons. In the output function, let a = 0.5 * (1 - |v|), k = a, u = a, and let v take values ​​in the range [-0.5, 0.5], with the final value determined through training and unique to the neuron. Therefore, there are three variations of the activation function in the same layer: [-1, 1] type, [0, 1] type, and [-1, 0] type. Similarly, during training initialization, v can be randomly set to [-0.5, 0, 0.5].

[0041] The above-mentioned dynamic activation function at the same level* is similar to the dynamic gating implementation*. The specific values ​​of the parameters or the specific values ​​in the formula are for reference and convention only, and can be changed according to actual needs.

[0042] Gradual hard transition example* Some neural network acceleration kernels do not have the function of smooth nonlinear activation, or general computing kernels omit the calculation of smooth nonlinear activation in order to improve computing power (low-performance CPUs usually need to look up tables for interpolation). However, smooth nonlinear activation is used during training, so backpropagation training is more stable and efficient, and the training success rate is higher. But during inference, due to the limitations of computing cores and computing power, it is necessary to change to a hard transition fully linear piecewise function.

[0043] The G(x) function can be easily modified into a hard inflection point. That is, the three parameters u, k, j of the G(x) function, such as those symmetric to the origin, can be set to be shared by the layers or globally. As the training process progresses and the system error decreases, the parameters can be gradually adjusted by the program settings, so that the function curve changes from a large radian to a small radian, and eventually gets closer and closer to a hard inflection point, and finally stabilizes at a certain small radian.

[0044] like Figure 16 F(x) = tanh(x), the full-quadrant piecewise activation function y = G(x) = Segment 1: c * u * F((-k + x) / u) + c * k + vx>k Section 2: c * x + vj <= x <= k Segment 3: c * u * F((-j + x) / u) + c * j + vx <j Where v=0, k=1, j=-k, c=1, u=1, the function is symmetric at the point [0,0], and the output value range is [-2,2]. like Figure 17 F(x) = tanh(x), the full-quadrant piecewise activation function y = G(x) = Segment 1: c * u * F((-k + x) / u) + c * k + vx>k Section 2: c * x + vj <= x <= k Segment 3: c * u * F((-j + x) / u) + c * j + vx <j Where v=0, k=0.5, j=-k, c=1, u=0.5, the function is symmetric at the point [0,0], and its output value range is [-2,2]. Letting k + u = b ensures that the upper limit of the function G(x) remains equal to b. Figure 16 and Figure 17In the curve, b is 2.k = buv. Gradually adjusting the value of u to near 0, G(x) gradually transforms into a function approaching a hard transition. That is, the limit and slope remain unchanged, but the curvature of the nonlinear segment is reduced. The value of u can be linked to changes in training error; as the training error decreases, the value of u is reduced until a preset value is reached. The curvature of the activation function gradually decreases to the preset value, thus the trained network parameters are closer to the hard transition activation function used for inference.

[0045] ---------- To explain the hardware circuitry, the following excerpt is from the priority document concerning the capacitive computing circuitry, specifically the content directly related to the activation function.

[0046] The network link input current modulation control module and the weighted current modulation control module jointly control the charging and discharging current of the capacitor module. In this embodiment, the network link input current modulation control module is a circuit module that dynamically controls the current and is used to generate the modulation current. (Known forms of modulation current include, but are not limited to: pulse width modulation (PWM), pulse density modulation (PDM), pulse position modulation (PPM), single pulse level, etc.; the modulation current is converted into charging and discharging current by the control signal through the switching circuit of the current channel, and its implementation includes direct clock generation, multi-channel signal mixing, reference signal trimming or delay insertion, etc. The current waveform also includes, but is not limited to, centrally symmetrical or edge-aligned, rectangular wave, triangular wave or sine wave, etc. The modulation current in other embodiments is the same.) The modulation current plus the switching of the internal resistance of the weighted current modulation control module (5)(6)(17)(18) realizes the control of the current. In this embodiment, the weighted current modulation control module is a small uniform resistor. In other embodiments, it can be a complex resistor network or an equivalent circuit. In this embodiment, the effect of the current flowing into the capacitor controlled by the network link input current modulation control module and the weight current modulation control module is equivalent to the link parameter w * input value x in the neural network. The network link input current modulation control module contains a fast access unit (the form of the fast access unit includes, but is not limited to: register, register file, flip-flop, latch, various SRAM / DRAM cells, array, and various new storage technology access units. Other embodiments are the same), which is set by the control information channel (1) (the control information channel and related fast access units also include functions such as enable, capacitor link port high impedance control, positive and negative setting of link parameter w and input value x, etc.), and generates current modulation (PWM, etc.) by comparing with the timing information of the timing information channel (2). Here, a part (bits) of the current modulation fast access unit represents the network input value x, and a part represents the link parameter w (which works together with the weight current modulation control module). The driving voltage channel (3) is linked to a dynamically selectable capacitor charging and discharging driving voltage source shared by the entire neural network layer. For example, the driving voltage a = 0 volts during discharge and the driving voltage a = 2 volts during charging. The specific voltage value depends on the actual needs; this is just for ease of description. The above control information channels, timing information channels, and even resistance / current channels are not limited to a single physical power-on link. The specific number of physical power-on links is determined according to the complexity of the data and actual needs.

[0047] (As is well known, a timer has multiple channels, and each channel has a fast access unit. Comparing the current with the timer's fast access unit generates multiple corresponding current modulations, which is a basic function of a microcontroller. In addition to being driven by a crystal oscillator clock, a timer can also be implemented using a delay chain (such as an inverter).) In this embodiment, the current is basically controlled by current modulation and a standard resistor network. In addition to expressing x, current modulation is also used to express w. Sometimes, in order to save chip area, some embodiments (specific value resistor embodiments) use different specific resistor values ​​to express the link parameter w, or even use a specifically generated stable voltage source to replace current modulation.

[0048] The capacitor modules of the last layer neurons are divided into positive and negative capacitor modules. In this embodiment, the left capacitor module represents a positive value (7), and the right capacitor module represents a negative value (22). The network link input current modulation control module will select to charge or discharge the positive or negative capacitor module according to the sign bit of the fast access unit set by the internal link parameter w and the sign bit of the fast access unit set by the input value x. That is, it determines which capacitor module to charge or discharge according to the sign of (w*x).

[0049] The capacitor control modules (8) and (16) correspond to their respective capacitor modules. During the charging and discharging process of the capacitor modules, the capacitor modules are pre-charged and discharged through the initial voltage source channel (9). In the charging mode, the capacitor modules are pre-charged to the initial voltage (in this embodiment, the initial voltage b = 1 volt); in the discharging mode, the capacitor modules are pre-charged to the initial voltage (in this embodiment, the initial voltage b = 2 volts). In the current embodiment, the initial voltage source can be set to an initial voltage b = 1 volt or b = 2 volts through the layer parameters of the neural network (1 or 2 volts is just an example reference, and the specific voltage value should be set according to the actual situation). If the neural network is more complex and there are different activation functions in one layer, or other more complex requirements, then multiple initial voltage source channels (9) with different voltages are needed. After the pre-charging and discharging steps, the input current modulation control module and the weight current modulation control module are linked through the network, and the neural network is charged and discharged, which is the actual working process of the neural network. After the neural network charging and discharging steps, the capacitor discharge detection begins. The capacitor module discharges to GND through the standard resistor R inside the capacitor control module until the detection voltage c = 1 volt (the specific voltage value is set according to the actual situation; in this embodiment, it is 1 volt). The capacitor control module contains a voltage comparator. When the voltage of the discharged capacitor module is lower than the detection voltage c, the output port (15) of the capacitor control module flips, displaying a level representing the comparison result (the level is set as needed; in this embodiment, it is set to high). All operations of the capacitor control module are manipulated by its own capacitor control channel (23). (Voltage comparison is a basic function of the chip, and its circuitry is common knowledge.) The neuron output module (14) logically merges the output signals of the capacitor control modules representing positive and negative values ​​(here, the symbol P / N represents the output of the positive and negative capacitor control modules, i.e., the positive module output is P and the negative module output is N). The layer output (13) of the neuron output module (14) represents the information of the duration difference of the discharge. In this embodiment, its information is logically equal to the level duration of P xor N. The layer output (12) of the neuron output module (14) represents the comparison of the magnitude of the positive and negative capacitor voltages or the comparison of the magnitude of the positive and negative capacitor discharge durations. In this embodiment, it is logically equal to the level of P output by the positive module before P xor N flips last. (It is known that logical operations such as xor are basic functions of digital circuits.) Note that in this embodiment, what needs to be output is the duration difference information, not just the duration difference level itself. The duration difference information has multiple forms of expression. Although this embodiment outputs the duration difference level for processing by the microcontroller and other modules, some embodiments (such as the current modulation layer output embodiment) require converting the duration difference information into a current modulation signal for output. That is, the neuron output module (14) has a time-to-digital converter circuit for detecting the duration difference. The layer output control signal channel (11) controls the time-to-digital converter circuit of the neuron output module to capture the Pxor N and obtain its duration through the timing information input through the timing information channel (10) and store it in the internal fast access unit. Finally, based on the value of the internal fast access unit and the timing information input through the timing information channel (10), the current modulation signal is output. (It is known that detecting the square wave pulse width is a basic function of the microcontroller. At low frequencies, a counter can be used to count clock pulses, and at high frequencies, a delay chain signal / multi-phase multi-line signal is used.) The following are the steps of the working method of the analog-digital hybrid neural network circuit in this embodiment (some steps can be performed in parallel according to actual needs, and the specific voltage values ​​are set according to actual needs): 1. Initialization settings: Set the fast access unit values ​​of the network link input current modulation control module, including enable, high impedance control of capacitor link port, positive and negative settings of link parameter w and input value x, etc.

[0050] 2. Select appropriate resistance values ​​or resistor networks as needed.

[0051] 3. Generate an appropriate current modulation signal.

[0052] 4. Capacitor pre-charge and discharge: Use the initial voltage source channel (9) to pre-charge and discharge the capacitor module (7)(22) to a specific initial voltage (charging mode b=1 volt, discharging mode b=2 volts). Pre-charge and discharge can be achieved by directly connecting to the initial voltage source or by detection through the voltage comparator (of the capacitor control module).

[0053] 5. Perform charging and discharging. Based on the positive and negative values ​​of w and x set internally, select to perform charging and discharging operations on the positive capacitor module (7) and the negative capacitor module (22).

[0054] 6. The charging and discharging current of the capacitor module (7)(22) is jointly controlled by the network-linked input current modulation control module (21)(19) and the weighted current modulation control module (4)(20)(5)(6)(17)(18).

[0055] 7. After the charging and discharging process, during detection, the charging and discharging continues until the detection voltage is reached. (Regarding energy efficiency, the detection process can be further optimized by setting multiple reference comparison voltages and selecting different resistance values ​​for the detection discharge resistor based on the voltage range of the positive and negative capacitors during detection, thereby accelerating the detection discharge speed. Because the result is a difference and the positive and negative values ​​switch synchronously, it does not affect the final result.) 8. Output detection result information / signal.

[0056] 9. Layer output processing: The neuron output module (14) combines the output signals (P and N) of the positive and negative capacitance control module and uses the time difference information to represent the final output. The time difference information and sign are the level and signal pulse, or the signed digital value after further time-to-digital conversion.

[0057] 10. In some embodiments, it may be necessary to convert this duration difference into a current-modulated output or other form of electrical information, which involves capturing the level duration of P xor N and outputting a current-modulated signal based on this information.

[0058] As is well known, the conversion from time pulse to digital quantity described in step 9 can be achieved through a time-to-digital converter (TDC). Specific implementations include, but are not limited to, counters, interpolators, inverter delay chains, vernier signal structures, time amplifier circuits, etc. In the case of a multi-layer neural network, if the input current modulation uses a single-pulse level, then the time pulse can be directly used as the input to the next layer of the neural network, or the entire network can use single pulses. Note that mathematically, it can be verified that if the PWM duty cycle variable in the derivation below is replaced with the pulse width of a single pulse, i.e., the pulse duration, the resulting formula, even in the counterintuitive case where the single pulses are not aligned in time, still holds true. That is, the output time difference is proportional to the product of the total pulse width and conductance of each input (the time integral of the total conductance), which still holds true due to the exponential multiplication effect. The capacitor discharge process is an exponential multiplication, with the discharge voltage V(t) = V0 * e^(-t / (RC)). e^a * e^b = e^(a+b), which is independent of the order, length, or even overlap of the single pulses corresponding to a or b. (That is, it only depends on the time integral of the total positive / negative conductance; the derivation is omitted). However, appropriate pulse repetition in PWM can average out various uncontrollable factors such as interference and inconsistencies, although it also increases power consumption. Single pulses, PWM, capacitor charge transfer, or other more complex time- and conductance-based similar forms are essentially all about the ratio of conduction time and current amplitude between various links and detection channels. The final result can be obtained by applying similar formulas to achieve the proportional relationship.

[0059] The following is a computational description of the working principle of the aforementioned hardware, and an explanation of how to generate the connection parameters w and input value x of the hardware. The calculations below are based on ideal components, neglecting leakage current, voltage and temperature variations, and other interference factors. Therefore, the calculations are approximate results. In actual neural network training, if a test statistical characteristic table of the hardware circuit is needed, the software layer will use a lookup table and interpolation method to obtain the actual input / weight / output values / derivatives of the circuit. The following calculations use the simplest PWM current modulation form as an example, but its essence is to control the current ratio of each current channel through switching; therefore, other forms of modulation are equivalent. Furthermore, in the following calculations, the pulse width, resistance, and capacitance do not require precise values; what is needed are precise and stable ratios.

[0060] The value x represents the duty cycle of the input portion in current modulation (PWM pulse width modulation; in single-pulse modulation, it's the pulse duration width, i.e., the ratio to the minimum pulse duration width). The duty cycle affects the voltage difference between the equivalent input voltage source and the capacitor voltage. Current input x = Xn = pulseXn The value w is the duty cycle of the portion representing the equivalent resistance in current modulation (pulseRn; in single-pulse modulation, this can be omitted and set to 1; if this weighted item exists in single-pulse modulation, it represents the scaling factor of the pulse duration or current intensity), divided by the corresponding resistance Rn and capacitance cap. Current weighted current limiting w = Wn = pulseRn / Rn / cap SSP(...), SumSelectPositive means selecting all values ​​greater than or equal to 0, summing them, and taking the absolute value.

[0061] SSN(...), SumSelectNegative means selecting all values ​​less than 0, summing them, and taking the absolute value.

[0062] V[t] represents the time function of the capacitor voltage, V'[t] is its derivative, and e is the natural constant. 1.1. In discharge mode, the initial voltage is 'a', and the discharge drive voltage is GND voltage 0. After the neural network discharges, it enters the detection process to continue discharging, with the stop / detection voltage c=b=1. The discharge resistor used in the detection process is R. t is a time variable, which is a standard time length that can be set and controlled.

[0063] Solve the difference equations respectively (b <= V[t] <= a). Dsolve[{V'[t]==SSP(...,Xn*(-V[t])*Wn), V[0]==a}, {V[t]}, t] DSolve[{V'[t]==SSN(...,Xn*(-V[t])*Wn), V[0]==a}, {V[t]}, t] Solution results The positive capacitance V[t] is given by gp1 = a * e^(-t*SSP(...,Xn*Wn)). The negative capacitance V[t] is given by gn1 = a * e^(-t*SSN(...,Xn*Wn)). During detection, the positive capacitor discharge time function tp1 = cap*R*ln[gp1 / b]=cap*R*(ln[gp1]-ln[b]) During detection, the discharge time function of the negative capacitor is tn1 = cap*R*ln[gn1 / b]=cap*R*(ln[gn1]-ln[b]). The difference between the two, tp1-tn1, is calculated according to the logarithmic rule. ln[a * e^(-t*SSP(...,Xn*Wn))]=ln[a]+ln[e^(-t*SSP(...,Xn*Wn)] = ln[a]-t*SSP(...,Xn*Wn) ln[a * e^(-t*SSN(...,Xn*Wn))]=ln[a]+ln[e^(-t*SSN(...,Xn*Wn)] = ln[a]-t*SSN(...,Xn*Wn) tp1 = cap*R*(ln[a]-t*SSP(...,Xn*Wn)-ln[b]) tn1 = cap*R*(ln[a]-t*SSN(...,Xn*Wn)-ln[b]) tp1-tn1 = -cap*R*(X1*W1+X2*W2+...+Xn*Wn)*t As can be seen from tp1 and tn1, its activation function is a linear function. Substituting this into Wn = pulseRn / Rn / cap... tp1-tn1 = -R*(X1*pulseR1 / R1+X2*pulseR2 / R2+...+Xn*pulseRn / Rn)*t Here, R and R1...Rn are all proportional. The value of pulseR1*R / R1 is the value of the weighting term. If all resistors are standard resistors, that is, R1...Rn equals R, then... tp1-tn1 = -(X1*pulseR1+X2*pulseR2+...+Xn*pulseRn)*t If we flip the positive and negative signals of the neuron's output module, then the result of tp1-tn1 is... R*(X1*pulseR1 / R1+X2*pulseR2 / R2+...+Xn*pulseRn / Rn)*t (X1*pulseR1+X2*pulseR2+...+Xn*pulseRn)*t, In simple terms, if the modulated current is generated using PWM, the total duty cycle is duty_n = pulseXn * pulseRn. The input value Xn = t * pulseXn, representing the duty cycle of the input portion. The weight value Qn = pulseRn * R / Rn, where pulseRn is the duty cycle representing the weight portion of the neuron's connection, Rn is the resistance of the weight term, and R is the standard resistor for detecting discharge. In essence, the calculation is achieved by adjusting the ratio of the standard resistor to the weight term resistance, the duty cycle of the weight term and the input term, and the total discharge time to make them equivalent to the input and weight values ​​in the software.

[0064] Additionally, there is a time t. If t is not the base time, assuming t=8, the final discharge detection time or input value should be subtracted from the time t. Since t=8 is a multiple of 2, the discharge detection time or input value can be directly shifted by 3 bits.

[0065] If there are deviations in the production process, resulting in one capacitor having a larger capacitance than the other (cap1 and cap2), then the above formula becomes... tp1 = cap1*R*(ln[a]-t*SSP(...,Xn*Wn)-ln[b]) tn1 = cap2*R*(ln[a]-t*SSN(...,Xn*Wn)-ln[b]) tp1-tn1 =R*(cap1-cap2)*ln(a / b)-R*(X1*pulseR1 / R1+X2*pulseR2 / R2+...+Xn*pulseRn / Rn)*t The calculation error caused by the capacitance difference: err_cap = R * (cap1 - cap2) * ln(a / b) Because the capacitance difference is an independent term, its error value can be easily obtained through calculations with a 0% duty cycle. For digital circuits (with digital-to-analog conversion), this error can be subtracted from the final calculation result. For analog circuits (without analog-to-analog conversion), the error can be corrected by adjusting the neuron bias term connections or by setting additional bias term connections.

[0066] The standard resistor R and other resistors R1, R2...Rn may have manufacturing errors. However, we only use the resistance ratio R / Rn here, which is relatively accurate. In addition, manufacturing errors, operating temperature, voltage, line inductance, frequency, leakage current, and switching delay will all affect the calculation results. Various errors can be adjusted using appropriate fine-tuning circuits or by inserting a delay into the PWM. These adjustment circuits and how to reduce manufacturing errors are very specific design tasks, which will not be discussed in detail here. A relatively simple method for adjusting the weights is to obtain the average deviation of the weights for each link through multiple calculations based on specific parameters (such as 0), and then correct the final result by modifying the ideal weight values ​​(before using the weights).

[0067] 1.2. If the detection process begins, and the device recharges from gn1 or gp1 to a=2, the stop / detection voltage is b=a. The discharge resistor used in the detection process is R. The charging drive voltage is c=a+1, then... The positive capacitor discharge time function during detection is tp3 = cap*R*ln[(c-gp1) / (ca)] ; tp3 = cap*R*ln[3 - 2 * e^(-t*SSP(...,Xn*Wn))] The discharge time function of the negative capacitor during detection is tn3 = cap*R*ln[(c-gn1) / (ca)] ; tn3 = cap*R*ln[3 - 2 * e^(-t*SSN(...,Xn*Wn))] It is evident that tp3 and tn3 are some kind of nonlinear functions. 2.1. Now returning to the charging mode in the simplest embodiment, the initial voltage is b=1, the charging drive voltage is a=2, and after the neural network charging is complete, it enters the detection process to begin discharging to ground, with the stop / detection voltage c=b=1. The discharge resistor used in the detection process is Rt, which is a time variable. Mathematical software solves the difference equations (b <= V[t] <= a). DSolve[{V'[t]==SSP(...,Xn*(aV[t])*Wn), V[0]==b}, {V[t]}, t] DSolve[{V'[t]==SSN(...,Xn*(aV[t])*Wn), V[0]==b}, {V[t]}, t] Solution results For a positive capacitor, V[t] is given by gp2 = a+(ba) * e^(-t*SSP(...,Xn*Wn)); for a negative capacitor, V[t] is given by gn2 = a+(ba) * e^(-t*SSN(...,Xn*Wn)). The discharge time function for the positive capacitor during detection is tp2 = cap*R*ln[gp2 / b]; the discharge time function for the negative capacitor during detection is tn2 = cap*R*ln[gn2 / b]. The difference between the two, tp2-tn2, can be substituted into b=1 and a=2. tp2 = cap*R*(ln[2-e^(-t*SSP(...,Xn*Wn))]); tn2 = cap*R*(ln[2-e^(-t*SSN(...,Xn*Wn))]) tp2-tn2 = cap*R*(ln[2-e^(-t*SSP(...,Xn*Wn))] - ln[2-e^(-t*SSN(...,Xn*Wn))]) As can be seen from tp2 and tn2, their activation function is the difference between two nonlinear functions tp2 and tn2. The curves of tp2 and tn2 in the first quadrant are similar in shape but have better training performance than the cap*R*ln(2)*tanh function. tanh is a mature neural network activation function with good performance. The maximum values ​​of tp2 and tn2 are cap*R*ln(2).

[0068] In other words, when nonlinear activation is required in this embodiment, the activation function used is to group the Xn*Wn groups according to their positive and negative signs, sum the sums of each group, take the absolute value (val), and then perform a nonlinear transformation of cap*R*ln(2-e^(-t*val)), and then directly take the value or take the difference between the two groups after the nonlinear transformation.

[0069] 2.2. If the detection process begins, the capacitor is recharged from gn2 or gp2 to a=2, and the stop / detection voltage a is reached. The charging resistor used during the detection charging process is R. The charging drive voltage is c=a+k, k=1. Then, the positive capacitor charging time function during detection is tp4 = cap*R*ln[(c-gp2) / (ca)] = cap*R*ln[a+k - (a+(ba) * e^(-t*SSP(...,Xn*Wn)))] = cap*R*ln[1 + (ba) / k * e^(-t*SSP(...,Xn*Wn))] = cap*R*ln[1 + e^(-t*SSP(...,Xn*Wn))] The charging time function of the negative capacitor during detection is tn4 = cap*R*ln[(c-gn2) / (ca)] = cap*R*ln[ a+k - (a+(ba) * e^(-t*SSN(...,Xn*Wn)))]= cap*R*ln[ 1 + (ba) / k * e^(-t*SSN(...,Xn*Wn))] = cap*R*ln[ 1 + e^(-t*SSN(...,Xn*Wn))] As can be seen from tp4 and tn4, their activation function is the difference between the same two nonlinear functions tp4 and tn4. 3.0. Example of Positive and Negative Capacitor Detection and Comparison* Similarly, the capacitor is first pre-charged to a predetermined voltage 'a', then the network discharges for forward propagation of the neural network, and then the detection process is executed. The voltage comparator directly compares the voltages of the positive and negative capacitors. The capacitor with the larger voltage discharges to ground through a standard resistor R during the detection process until the voltages of the two capacitors are equal. The discharge time is the time difference. Its sign represents the result of comparing the magnitudes of the positive and negative capacitors at the start of detection. Its unsigned value is consistent with |tp1-tn1|, and is also a linear function. However, because the voltage comparator has an offset voltage, and various leakage currents may be inconsistent, this will affect the accuracy and efficiency of the calculation results.

[0070] 3.1. Example of Dual-Capacitor Detection and Comparison* (Each capacitor has positive and negative terminals, i.e., 4 capacitors) Here, we supplement another example of detection and discharge comparison (dual-capacitor detection and comparison), where an identical capacitor is added to the capacitor module. This capacitor does not participate in the charging and discharging of the neural network; it only plays a role during detection (the network discharges, and detection also discharges). Before detection, this capacitor is pre-charged to 'a', and then it discharges to ground (0 voltage) through 'R'. The stop / detection voltage is set to the voltages of the capacitors in the capacitor module that participate in the charging and discharging of the neural network, i.e., gp and gn, i.e., the voltages of two capacitors with the same sign are compared using analog voltage comparison. During detection, the positive capacitor discharge time function tp0 = cap*R*ln[a / gp] = cap*R*(t*SSP(...,Xn*Wn)) During detection, the discharge time function of the negative capacitor is tn0 = cap*R*ln[a / gn] = cap*R*(t*SSN(...,Xn*Wn)). diff0 = tp0 - tn0 = cap*R*(X1*W1+X2*W2+...+Xn*Wn)*t; From tp0 and tn0, it can be seen that its activation function is also a linear function. 3.2 In the dual-capacitor detection process (where the network is charging and detection is also charging), the capacitor is pre-charged to b=1V before detection, and then charged through R with a driving voltage of a=2V (the specific voltage value depends on the actual setting; this is only for reference). The stop / detection voltage is set to the voltages of the capacitors participating in the neural network charging and discharging in the capacitor module, i.e., gp2 and gn2, that is, the voltages of two capacitors with the same sign are compared using analog voltage comparison. During testing, the positive capacitor charging time function is tp5 = cap*R*ln[(a-gp2) / (ab)]. = cap*R*ln[(a-(a+(ba) * e^(-t*SSP(...,Xn*Wn)))) / (ab)] =cap*R*ln[e^(-t*SSP(...,Xn*Wn))] The charging time function of the negative capacitor during detection is tn5 = cap*R*ln[(a-gn2) / (ab)]. = cap*R*ln[(a-(a+(ba) * e^(-t*SSN(...,Xn*Wn)))) / (ab)] = cap*R*ln[e^(-t*SSN(...,Xn*Wn))] tp5-tn5 = -cap*R*(X1*W1+X2*W2+...+Xn*Wn)*t As can be seen from TP5 and TN5, their activation functions are also linear functions. The optimal method for training model parameters is to calculate them using the statistical characteristics of the circuit through table lookup and interpolation. After obtaining the final model parameters through backpropagation training, these parameters are then input into the circuit of this embodiment to achieve neural network output (for inference). For parameters exceeding the range in the software, they can be divided into multiple parameters and multiple inputs, or the hardware circuit can be improved to allow the weighted current modulation control module to select more different resistance values.

[0071] In the computational part, as summarized above, tp5-tn5, tn1-tp1, tp0-tn0, etc., are ideally equivalent to the summation of product terms commonly used in software neural networks (the constant term is simply setting the input of one of the product terms to 1). tp4-tn4, tp3-tn3, tp2-tn2, etc., can also be used in neural networks, but there has been no attempt or publicly available information from the software and academic communities.

[0072] Looking at the derivation process of tp5-tn5, tn1-tp1, tp0-tn0, for example, tn1, tp1, and tn1-tp1, the calculation results are independent of the capacitance value of the charging and discharging capacitors. This greatly facilitates the design and production of chips or circuits, because the capacitance value is difficult to determine precisely. Our invention's calculation results do not depend on the capacitance value, which brings a huge advantage to mass production applications. In addition, because tn1-tp1 calculates the difference, the errors in tn1 and tp1 caused by the leakage current of the positive and negative capacitor control module and the capacitor module can be mutually canceled to a certain extent (of course, the leakage current should be minimized in the design). The calculation errors caused by the resistance value deviation and the errors caused by the current modulation time deviation are relatively small, and their accuracy is high. The capacitor and resistance deviations (only related to the ratio) can be corrected by obtaining consistency parameters during testing and written into the internal memory of the neuron output module. The neuron output module automatically adjusts the XOR time difference when it is working.

[0073] As can be seen from tp5-tn5 and tp0-tn0, although the dual capacitors use twice the amount of capacitors, they can still perform the calculation of accumulating positive and negative product terms whether charging or discharging.

[0074] This embodiment implements the essential multiply-accumulate, linear activation, and nonlinear activation functions required by neural networks. Other activation functions are not strictly necessary for neural networks and can be subsequently handled by computing chips / GPUs / microcontrollers, or see more embodiments below. The multi-channel multiply-accumulate parallel circuit described above reduces the number of CMOS transistors used by at least an order of magnitude compared to a single floating-point multiplier in existing chips.

[0075] In some embodiments, grouping operations are not desired (in embodiments with single-group nonlinear activation*). Instead, it is preferable to use the traditional multiply-add form w1x1 + w2x2... followed by nonlinear transformation activation. This is consistent with the implementation of traditional neural networks. While there are hardware workarounds in this invention, they will increase system latency. Details are as follows: The layer network uses a discharge mode. Each neuron output module has an additional equivalent conversion capacitor, pre-charged to c=1 volt. During the effective level of the P xor N output (as explained above, this duration is linear), the conversion capacitor is charged with a driving voltage a=2 through a resistor R (the resistor value here is set to the same value as the detection discharge resistor). Then, similar to the detection discharge process above, the charged conversion capacitor is discharged, also through the equivalent resistor R to ground. The discharge stops when the conversion capacitor voltage is less than or equal to the detection voltage b=c=1 volt. Similarly, the duration of the conversion capacitor detection discharge is a nonlinear transformation function with a shape similar to cap*R*ln(2)*tanh first quadrant curve. That is, f(y) = cap * R * (ln[2 - e^(-y)]), (y = tp0 - tn0 or y = tp1 - tn1). In effect, it's equivalent to adding a non-linear layer with one-to-one neuron connections on top of a linear activation layer.

[0076] Some embodiments may require implementing a nonlinear activation function (ADC detection type embodiment*) in the capacitor control module (8)(16) by using a high-speed ADC to read the capacitor voltage into a fast access unit, where a=2, b=1. The positive capacitance V[t] is given by: gp2 = a + (ba) * e^(-t*SSP(...,Xn*Wn)) = 2 - e^(-t*SSP(...,Xn*Wn)). The negative capacitance V[t] is given by: gn2 = a+(ba) * e^(-t*SSN(...,Xn*Wn)) = 2-e^(-t*SSN(...,Xn*Wn)). Using op-amps to perform subtraction (gp2-1, gn2-1), and then importing the result into an ADC, the voltage result from the ADC is stored in a fast access unit (the value in the fast access unit). Similarly, a nonlinear function with a zero-crossing extreme of 1 and a tanh-like shape can be obtained in the first quadrant. That is 1-e^(-y), y=t*(SSP or SSN)(...,Xn*Wn) ---------- Examples of capacitive analog calculations* The preceding references primarily describe a specific implementation scheme for a neural network-specific computational circuit based on parallel equivalent resistance as neuron links, equivalent time pulses as current modulation, and capacitor charging / discharging as the core. Its linear and nonlinear activation calculation functions are clearly defined and are piecewise functions. G(x) is a single-quadrant, bi-segmented circuit passing through the origin [0,0], with the origin as the starting point. Its activation-related circuitry has two levels of numerical conversion: linear segment numerical conversion and nonlinear segment numerical conversion. In this circuit, the unsigned analog quantity, used as the multiplication and accumulation value input to the neuron link, is compared with the reference analog quantity corresponding to the intersection point of the piecewise function. Based on this comparison, the nonlinear segment numerical conversion circuit is enabled; that is, when the input value exceeds the linear segment, the nonlinear segment numerical conversion circuit is activated (e.g., outputting corresponding linear or nonlinear equivalent time pulses depending on the charging / discharging mode mentioned above). After the computational circuit generates and outputs the linear segment analog quantity for numerical conversion, it continues to generate and output the nonlinear segment analog quantity based on the comparison result. Multiple such neurons and multi-layered such neuron computing circuits, due to pure analog computing, have performance and energy efficiency that are orders of magnitude better than existing pure transistor switching computing circuits.

[0077] The circuit function described above, mathematically, involves dividing the analog input x to the piecewise function G(x) into two analog inputs x0 corresponding to the intersection point of the first quadrant of the piecewise function: a linear analog input x1 and a nonlinear analog input x2, i.e., x = x1 + x2. When x > x0, the linear analog input x1 = x0, and the nonlinear analog input x2 = x - x0. When 0 <= x <= x0, the linear analog input x1 = x, and the nonlinear analog input x2 = 0. The two output analog inputs, linear and nonlinear, are calculated using different analog circuits based on the linear subfunction y1 = c * x1 and the nonlinear subfunction y2 = c * F(x2), respectively. The output of G(x) is the sum of two analog inputs with equivalent pulse times, i.e., G(x) = y1 + y2, output in two rounds or two channels.

[0078] Furthermore, in digital multi-core computing systems, the output values ​​of nonlinear functions are often obtained through table lookups and interpolation, especially in streamlined general-purpose computing cores. The calculation of G(x) functions with the same u value can share a lookup interpolation data table (this table is very small), meaning different G(x) activation functions share the same nonlinear segment curve shape and the same memory data. The value of G(x) is the function value obtained from the nonlinear f(x) lookup table, plus the function value at the endpoint of the linear segment, i.e., the linear offset.

[0079] In analog capacitive computing circuits, various activation functions can be generated based on the charging and discharging modes and related voltage settings. These functions include a tanh-like function part in the first quadrant. Such numerical transformation characteristic functions with a tanh-like form are common in physical devices and physical laws.

[0080] like Figure 8 F = tanh (815), F = ln[2-e^(-x)] (816), F = ln[3-2e^(-x)] / 2 (817). The characteristics of the first quadrant of these three functions are: firstly, they all have a limit value; as x increases, the derivative tends to 0, and the function output value tends to the limit; secondly, the slope (derivative is 1) and value (value is 0) at the origin can smoothly connect with the shape of y=x at the origin. Therefore, ln[q+1-q*e^(-x)] / q can also be used to construct piecewise functions similar to those in the previous examples. Specifically: F(x) = ln[q+1-q*e^(-x)] / q (non-linear) When q=1, F(x) = ln[2-e^(-x)]. In this embodiment, the F(x) function is a calculation function extracted from the operation and working rules of the hardware circuit, which is completely consistent with the physical calculation. However, this is not a function symmetric at the origin. Therefore, we only use the first quadrant of F(x). By modifying the circuit to set the corresponding signs of the input and output, it is easy to convert a function that is only in the first quadrant into a symmetric function.

[0081] like Figure 9 The single-quadrant piecewise activation function y = G(x) = Segment 1: c * u * F((-k + x) / u) + c * k + vx>k Section 2: c * x + vx >= 0 Where v=1, k=1, c=1, u=1, the function 821 is only in the first quadrant, and the output value range is [1, 2+ln(2)]. In comparison, the piecewise functions 820 and 821 have the same curve shape but different positions.

[0082] like Figure 10 The single-quadrant piecewise activation function y = G(x) = Segment 1: c * u * F((-k + x) / u) + c * k + vx>k Section 2: c * x + vx >= 0 Where v=0, k=1, c=2, u=1, this function only operates in the first quadrant, and its output value range is [0, 2+2*ln(2)]. like Figure 11 The single-quadrant piecewise activation function y = G(x) = Segment 1: c * u * F((-k + x) / u) + c * k + vx>k Section 2: c * x + vx >= 0 Where v=0, k=1, c=2, u=1 / ln(2), this function is only in the first quadrant, the output value range is [0,2], and the limit value of 823 is 2; Segment 1 824 splicing the part of y=x that passes through the origin 825.

[0083] As can be seen, this piecewise function is also flexible, and its curve is similar to that of the tanh-type implementation. For an origin-symmetric activation function, if used for hardware circuit calculations, simply inverting the output value sign according to the input sign can easily transform the curve in the first quadrant into a multi-quadrant curve symmetric about the origin. Shifting the multi-quadrant curve vertically is a simple operation for both analog and digital hardware calculation circuits. For example, in the third quadrant, the characteristic function calculated by converting the inverted activation value is F(x) = -ln[q+1-q*e^(x)] / q.

[0084] Because this function is the activation function for capacitive computing (q can be set by voltage), the trained model can be directly used in analog computing circuits. The linear and nonlinear parts of the activation functions are completely identical, bridging the gap between digital and analog model parameters. Therefore, the corresponding activation functions in capacitive computing circuits are remarkably consistent with those in analog, numerical, and general-purpose computing devices. Capacitors and resistors made of specific materials are the best-matched components in integrated circuits, and are also the most stable, basic, and simplest. By applying analog capacitive circuits to inference calculations, the performance and energy efficiency of neural network computations can be improved by several orders of magnitude.

[0085] There are many functions similar to tanh and ln[q+1-q*e^(-x)] / q, such as variants of the sigmoid function, 1 / (1+e^(-x)), q*(1 / (1+e^(-x))-0.5), and 4 / (1+e^(-x))-2. Refer to the piecewise function design method of this invention for specific details, and set them as needed.

[0086] F(x) = q*(1 / (1+e^(-x))-0.5) (Nonlinear) When q=4, F(x) = 4 / (1+e^(-x))-2 like Figure 12 The single-quadrant piecewise activation function y = G(x) = Segment 1: c * u * F((-k + x) / u) + c * k + vx>k Section 2: c * x + vx >= 0 Where v=0, k=1, c=1, u=1, this function only operates in the first quadrant, and the output value range is [0,3]. In summary, the above embodiments have sufficiently illustrated the implementation of the present invention. The present invention has various potential or hybrid embodiments and is not limited to the content described in the partial text.

Claims

1. An artificial neural network data processing system, characterized in that: a computing circuit performs computation of a multi-layer neural network, an input comprises a multi-dimensional vector / tensor containing graph / audio / semantic information, and a feature vector / tensor as a result is finally output; wherein, at least one layer of neurons of the neural network has numerical conversion which is used as output of the neurons, based on multiply-accumulate values of respective links of a previous layer, and determined according to an activation function; said activation function is a piecewise function; in said numerical conversion, the multiply-accumulate value is compared with a reference input value at a piecewise junction of said piecewise function, and a numerical conversion calculation circuit, a calculation subroutine or function calculation parameters corresponding to a corresponding function segment are selected according to a comparison result; said piecewise function G(x) has a linear segment c * x + v in the middle, and at least one end of two ends has a non-linear segment; the non-linear segment comprises a sub-function F(x), if there is an upper non-linear segment, the region where x>k is c * u * F((-k + x) / u) + c * k + v, if there is a lower non-linear segment, the region where x<j is c * u * F((-j + x) / u) + c * j + v; said sub-function F(x) passes through the origin [0,0]; a first-order derivative value F'(0) of said sub-function F(x) at the origin is 1; said sub-function F(x) increases monotonically in the first quadrant with decreasing increments; said sub-function F(x) has an upper limit in the first quadrant; c, v, u, k, j are configurable fixed values, or parameters modifiable during training or learnable.

2. The neural network data processing system according to claim 1, characterized in that: said sub-function F(x) is F(x)=tanh(x).

3. The neural network data processing system according to claim 1, characterized in that: said sub-function F(x) is F(x)=q*(1 / (1+e^(-x))-0.5); q=4.

4. The neural network data processing system according to claim 1, characterized in that: said sub-function F(x) for the upper non-linear segment is F(x)=ln[q+1-q*e^(-x)] / q; said sub-function F(x) for the lower non-linear segment is F(x)=-ln[q+1-q*e^(x)] / q; q is a configurable parameter.

5. The computing circuit system according to claim 4, characterized in that: said piecewise function G(x) is single-quadrant, has two segments, passes through the origin [0,0] with the origin as a starting point; there is a special-purpose neural network computing circuit, activation-related parts of which have linear segment numerical conversion and non-linear segment numerical conversion; the output of the activation-related parts of said special-purpose neural network computing circuit is analog quantity; said computing circuit can generate symmetric numerical conversion in other quadrants corresponding to the first quadrant by changing the signs of input and output; the conversion function corresponding to said non-linear segment numerical conversion is said sub-function F(x); The unsigned analog quantity of the multiply-accumulated value used as the input of the neuron link is compared with the reference input analog quantity corresponding to the intersection point of the piecewise function. Based on this result, the nonlinear segment numerical conversion correlation circuit is enabled, that is, the nonlinear segment numerical conversion correlation circuit is activated when the input value exceeds the linear segment. The aforementioned computing circuit generates the linear analog quantity for the linear segment numerical conversion; The nonlinear segment numerical conversion related circuit is activated to generate the nonlinear analog quantity of the nonlinear segment numerical conversion. The linear and nonlinear analog quantities are output independently on different lines or in different time periods.

6. The computing circuit system according to claim 1, characterized in that... Training methods that involve gradual hard transitions: Step 1: During training, set k, j, and u in G(x) as layer-shared parameters, and let k+u equal the upper limit value minus v, and j+u equal the lower limit value minus v; Step 2: Train the neural network for a certain number of rounds, and then adjust u and the corresponding k and j; Step 3: Repeat step 2 until training is complete; Step 4: Change the activation function G(x) to a hard transition, i.e., a fully linear piecewise function; Step 5: In the inference process, the trained model and parameters are used to change the activation function G(x) to a hard transition, i.e., a fully linear piecewise function.

7. The computing circuit system according to claim 1, characterized in that, There are methods for calculating G(x): Step 1: The analog quantity x, which is the input of the piecewise function G(x), is divided into two analog quantities according to the corresponding input analog quantity x0 at the intersection point of the first quadrant of the piecewise function, namely, the linear analog quantity part x1 and the nonlinear analog quantity part x2, i.e., x = x1 + x2. When x>x0, the linear analog component is x1=x0, and the nonlinear analog component is x2=x-x0; When 0 <= x <= x0, the linear analog component is x1 = x, and the nonlinear analog component is 0; Step 2: The analog quantities of the linear and nonlinear outputs are calculated separately using analog circuits based on the linear subfunction y1=c*x1 and the nonlinear subfunction y2=c*F(x2). Step 3, the output of G(x) is the sum of the two analog outputs, that is, G(x) = y1 + y2.

8. The computing circuit system according to claim 1, characterized in that... There are methods for training dynamic activation functions in the same layer: Step 1: Construct a neural network, which includes dynamic layers; Step 2: In the dynamic layer, v, k, j, u in G(x) are set as neuron-specific and trainable parameters, and k, j, u are set to establish a linear linkage relationship. Step 3: During the training of the relevant layers of the neural network, adjust v, k, j, u so that the range of values ​​of the activation function output of the gated layer neurons is differentiated into positive and negative activation, or no activation, or negative activation, or intermediate states.

9. The computing circuit system according to claim 8, characterized in that... There are methods for dynamic gating: The linear linkage relationship described herein is attributed to a single trainable parameter, meaning that all other parameters depend on this single trainable parameter. The dynamic layer is a dynamic gated layer. The output of a neuron in the dynamic gated layer is multiplied by the output of a neuron in another layer before being input into the next layer.

10. The computing circuit system according to claim 1, characterized in that, There are methods for constructing physical neural networks: Step 1: Based on the analog quantity conversion calculation function of the physical calculation circuit of the inference terminal that meets the requirements of the F(x) subfunction, construct the G(x) activation function; Step 2: On the training end, construct a multi-layer neural network model based on G(x) and complete the training. Step 3: In the physical computing circuit at the inference end, use the parameters of the multilayer neural network model constructed based on G(x).