Systems and methods for human activity recognition

Analog neuromorphic computing hardware addresses power and computational limitations by converting neural networks into efficient, low-power edge solutions for diverse applications, including human activity recognition with personalized capabilities.

JP7770589B2Active Publication Date: 2025-11-14POLYN TECHNOLOGY LIMITED
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024552681
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-04
Publication Date
2025-11-14
Estimated Expiration
2042-03-04

AI Technical Summary

Technical Problem

Existing hardware implementations of neural networks face challenges in power consumption, computational limitations, and impracticality of memristor-based crossbars, especially in edge environments, which are also computationally intensive and lack personalization in human activity recognition.

Method used

Analog neuromorphic computing hardware that models trained neural networks, offering improved performance per watt and parallelism, suitable for edge environments, and allows for mass production with minimal retraining, using techniques to convert neural network topologies into equivalent analog circuits.

Benefits of technology

Reduces power consumption by over 40% and enables recognition of up to 50 different human activities, providing personalized solutions with reduced computational burden and enabling applications like drone navigation and autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007770589000090
    Figure 0007770589000090
  • Figure 0007770589000091
    Figure 0007770589000091
  • Figure 0007770589000092
    Figure 0007770589000092
Patent Text Reader

Abstract

Systems, methods and devices for human activity recognition are provided. An exemplary device includes an integrated circuit for human activity recognition. The integrated circuit includes an analog network of analog components configured to implement a trained neural network model (e.g., an autoencoder) that is trained to generate a plurality of descriptors for a plurality of predefined human activities based on a plurality of features extracted from a plurality of electrical signals from one or more sensors. The device also includes one or more digital components configured to classify the human activity as one of a plurality of predefined human activities (e.g., using a classifier such as a K-nearest neighbor) according to the plurality of descriptors generated by the integrated circuit. In some implementations, the device further includes one or more sensors configured to collect a plurality of electrical signals during the human activity.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of, and claims priority to, U.S. Patent Application Publication No. 17 / 189,109, entitled "Analog Hardware Realization of Neural Networks," filed March 1, 2021, which is a continuation of PCT Application No. PCT / RU2020 / 000306, entitled "Analog Hardware Realization of Neural Networks," filed June 25, 2020, each of which is incorporated herein by reference in its entirety. U.S. Patent Application Publication No. 17 / 189,109 is also a continuation-in-part of PCT Application No. PCT / EP2020 / 067800, entitled "Analog Hardware Realization of Neural Networks," filed June 25, 2020, which is incorporated herein by reference in its entirety.

[0002] Technical Field FIELD OF THE INVENTION

[0002] The disclosed implementations relate generally to neural networks, and more particularly to systems and methods for human activity recognition using analog neuromorphic computing hardware. [Background technology]

[0003] background

[0003] Traditional hardware has not been able to keep up with innovations in neural networks and the growing popularity of machine learning-based applications. As advances in digital microprocessors have plateaued, the complexity of neural networks continues to outpace the computational power of CPUs and GPUs. Neuromorphic processors based on spiking neural networks, such as Loihi and True North, have limited applications. In GPU-like architectures, the power and speed of such architectures are limited by the data transmission rate. Data transmission can consume up to 80% of chip power, significantly impacting the speed of computation. Edge applications require low power consumption, but there are currently no known high-performance hardware implementations that consume less than 50 milliwatts of power.

[0004]

[0004] Memristor-based architectures using crossbar technology remain impractical for fabricating recurrent and feedforward neural networks. For example, memristor-based crossbars have many drawbacks, including high latency and current leakage during operation, making them impractical. Fabricating memristor-based crossbars also poses reliability issues, especially when neural networks have both negative and positive weights. For large neural networks with many neurons, high dimensions make it impossible to use memristor-based crossbars for simultaneous propagation of different signals, which complicates signal summation when neurons are represented by operational amplifiers. Furthermore, memristor-based analog integrated circuits have several limitations, such as a small number of resistance states, first-cycle problems when forming memristors, complex channel formation when training memristors, unpredictable dependence on memristor dimensions, slow operation of memristors, and drift in resistance states.

[0005] Additionally, the training process required for neural networks presents unique challenges for hardware implementations of neural networks. A trained neural network is used for a specific inference task, such as classification. Once a neural network is trained, a hardware equivalent is fabricated. If the neural network is retrained, the hardware fabrication process is repeated, driving up costs. While some reconfigurable hardware solutions exist, such hardware cannot be easily mass-produced and costs significantly more (e.g., five times more) than non-reconfigurable hardware. Furthermore, edge environments, such as smart home applications, do not require reprogrammability as such. For example, 85% of all neural network applications do not require retraining during operation, making on-chip learning less useful. Furthermore, edge applications involve noisy environments that can make reprogrammable hardware unreliable.

[0006]

[0006] Human activity recognition is an application that classifies human activities using information obtained from various types of sensors. Human activity recognition is one of the more common use cases for fitness devices such as bracelets and smart watches. Traditionally, human activity recognition in fitness bracelets or smart watches is implemented using software algorithms. Such algorithms are very computationally intensive and carry a large portion (e.g., more than 70%) of the computational load of the device's microcontroller unit. Furthermore, traditional solutions for human activity recognition can only distinguish between a few basic classes of human activities (e.g., standing, walking, or high-intensity movement), and these solutions are not personalized. Summary of the Invention [Problem to be solved by the invention]

[0007] overview

[0007] Therefore, there is a need for methods, circuits, and / or interfaces that address at least some of the above-identified shortcomings. Analog circuits that model trained neural networks and are fabricated according to the techniques described herein can offer the advantage of improved performance per watt, may be useful for implementing hardware solutions in edge environments, and can address diverse applications such as drone navigation and autonomous vehicles. The cost advantages offered by the proposed fabrication methods and / or analog network architectures become even more pronounced as neural networks become larger. Analog hardware implementation aspects of neural networks also offer improved parallelism and neuromorphism. Furthermore, neuromorphic analog components are less sensitive to noise and temperature changes compared to their digital counterparts. [Means for solving the problem]

[0008] Chips fabricated according to the techniques described herein offer orders of magnitude improvements in size, power, and performance over conventional systems, making them ideal for edge environments, including retraining purposes. Such analog neuromorphic chips can be used to implement edge computing applications or in Internet of Things (IoT) environments. Initial processing (e.g., creating descriptors for image recognition), which can consume over 80-90% of power due to analog hardware, can be moved onto the chip, thereby reducing energy consumption and network load and opening new markets for applications.

[0009] Various edge applications can benefit from the use of such analog hardware. For example, video processing can involve directly connecting to CMOS sensors without a digital interface using the techniques described herein. Other various video processing applications include road sign recognition in automobiles, camera-based true depth and / or simultaneous localization and mapping in robots, room access control without server connectivity, and always-on solutions in security and healthcare. Such chips can be used for data processing and low-level data fusion from radar and lidar. Such techniques can be used to implement battery management functions in large battery packs, voice / speech processing without data center connectivity, voice recognition in mobile devices, wake-up spoken commands for IoT sensors, translators that translate one language into another, and configurable process control using large sensor arrays and / or hundreds of sensors in IoT applications with low signal strength. In particular, such techniques can be used to implement human activity recognition. Replacing traditional software algorithms with analog neural network implementations can help offload computational burden from a device's microcontroller, such that the device's overall power consumption is significantly reduced (e.g., power consumption reduced by more than 40%). The techniques described herein can also be used to distinguish between numerous human activities, which can be used to develop personalized solutions for different users or classes of users. For example, neural network determination or descriptor generation of human activities allows for the recognition of over 50 different human activities. The neural network enables activity personalization, and descriptors can be generated as unique digital fingerprints for specific users.

[0010]

[0010] After standard software-based neural network simulation / training, neuromorphic analog chips can be mass-produced according to some implementations. Client neural networks can be easily ported with customized chip design and manufacturing, regardless of the neural network structure. Furthermore, according to some implementations, a library is provided that allows for ready-to-implement on-chip solutions (network emulators). Such solutions require only training, changing one lithography mask, and subsequent mass production of the chip. For example, only a portion of the lithography mask needs to be changed during chip manufacturing.

[0011] The techniques described herein can be used to design and / or fabricate analog neuromorphic integrated circuits that are mathematically equivalent to a trained neural network (either a feedforward neural network or a recurrent neural network). According to some implementations, the process begins with a trained neural network, which is first converted into a converted network consisting of standard elements. The operation of the converted network is simulated using software with known models representing the standard elements. Software simulation is used to determine individual resistance values ​​for each of the resistors in the converted network. Lithography masks are laid out based on the placement of the standard elements in the converted network. Each of the standard elements is laid out in the mask using existing circuit libraries corresponding to the standard elements to simplify and speed the process. In some implementations, resistors are laid out in one or more masks separate from masks containing other elements (e.g., operational amplifiers) in the converted network. In this way, when the neural network is retrained, only the masks containing resistors or other types of fixed resistive elements representing the new weights in the retrained neural network need to be regenerated, thereby simplifying and speeding the process. The lithography mask is then sent to a fab for fabricating the analog neuromorphic integrated circuit.

[0012]

[0012] In one aspect, a method for hardware realization of a neural network is provided according to some implementation aspects. The method includes obtaining a neural network topology and weights of a trained neural network. The method also includes converting the neural network topology into an equivalent analog network of analog components. The method also includes calculating a weight matrix of the equivalent analog network based on the weights of the trained neural network. Each element of the weight matrix represents a respective connection between analog components of the equivalent analog network. The method also includes generating a schematic model for implementing the equivalent analog network based on the weight matrix, including selecting component values ​​for the analog components.

[0013]

[0013] In some implementations, generating the schematic model includes generating a resistance matrix of the weight matrix, where each element of the resistance matrix corresponds to a respective weight in the weight matrix and represents a resistance value.

[0014]

[0014] In some implementations, the method further includes obtaining new weights for the trained neural network, calculating a new weight matrix for an equivalent analog network based on the new weights, and generating a new resistance matrix for the new weight matrix.

[0015] In some implementations, the neural network topology includes one or more layers of neurons, each layer of neurons calculating a respective output based on a respective mathematical function, and converting the neural network topology into an equivalent analog network of analog components includes, for each of the one or more layers of neurons, (i) identifying one or more function blocks for the respective layer based on the respective mathematical function, each function block having a respective schematic implementation having a block output that matches the output of the respective mathematical function, and (ii) generating a respective multilayer network of analog neurons based on arranging the one or more function blocks, each analog neuron implementing a respective function of the one or more function blocks, and each analog neuron in a first layer of the multilayer network being connected to one or more analog neurons in a second layer of the multilayer network.

[0016] In some implementations, the one or more function blocks include one or more basis function blocks selected from the group consisting of: (i) a block output;

number

number

[0017] In some implementations, identifying the one or more function blocks includes selecting the one or more function blocks based on a type of the respective layer.

[0018] In some implementations, the neural network topology includes one or more layers of neurons, each layer of neurons calculating a respective output based on a respective mathematical function, and converting the neural network topology into an equivalent analog network of analog components includes: (i) decomposing a first layer of the neural network topology into a plurality of sublayers, the decomposition including decomposing a mathematical function corresponding to the first layer to obtain one or more intermediate mathematical functions, each sublayer implementing an intermediate mathematical function; and (ii) generating, for each sublayer of the first layer of the neural network topology, a respective multi-layer analog subnetwork of analog neurons based on (a) selecting one or more sub-function blocks for the respective sublayer based on the respective intermediate mathematical function, and (b) arranging the one or more sub-function blocks, each analog neuron implementing a respective function of the one or more sub-function blocks, and each analog neuron of the first layer of the multi-layer analog subnetwork being connected to one or more analog neurons of a second layer of the multi-layer analog subnetwork.

[0019]

[0019] In some implementations, the mathematical function corresponding to the first layer includes one or more weights, and decomposing the mathematical function includes adjusting the one or more weights so that combining one or more intermediate functions results in the mathematical function.

[0020]

[0020] In some implementations, the method further includes (i) generating an equivalent digital network of digital components for one or more output layers of the neural network topology, and (ii) connecting the outputs of one or more layers of the equivalent analog network to the equivalent digital network of digital components.

[0021]

[0021] In some implementations, the analog components include multiple operational amplifiers and multiple resistors, each operational amplifier representing an analog neuron of an equivalent analog network, and each resistor representing a connection between two analog neurons.

[0022]

[0022] In some implementations, selecting component values ​​for the analog components includes performing a gradient descent method to identify possible resistance values ​​for a plurality of resistors.

[0023]

[0023] In some implementations, the neural network topology includes one or more GRU or LSTM neurons, and converting the neural network topology includes generating one or more signal delay blocks for each recurrent connection of the one or more GRU neurons or LSTM neurons.

[0024] In some implementations, one or more signal delay blocks are activated at a frequency that corresponds to a predetermined input signal frequency for the neural network topology.

[0025]

[0025] In some implementations, the neural network topology includes one or more layers of neurons that perform an unconstrained activation function, and transforming the neural network topology includes applying one or more transformations selected from the group consisting of: (i) replacing the unconstrained activation function with a constrained activation; and (ii) adjusting the connections or weights of the equivalent analog network so that, for one or more given inputs, the difference in output between the trained neural network and the equivalent analog network is minimized.

[0026] In some implementations, the method further includes generating, based on the resistance matrix, one or more lithography masks for fabricating a circuit that implements an equivalent analog network of analog components.

[0027]

[0027] In some implementations, the method further includes (i) obtaining new weights for the trained neural network, (ii) calculating a new weight matrix for an equivalent analog network based on the new weights, (iii) generating a new resistance matrix for the new weight matrix, and (iv) generating a new lithography mask for fabricating a circuit implementing the equivalent analog network of analog components based on the new resistance matrix.

[0028]

[0028] In some implementations, the trained neural network is trained using software simulation to generate the weights.

[0029]

[0029] In another aspect, a method for hardware realization of a neural network is provided according to some implementation aspects. The method includes obtaining a neural network topology and weights of a trained neural network. The method also includes calculating one or more coupling constraints based on analog integrated circuit (IC) design constraints. The method also includes converting the neural network topology to an equivalent loosely coupled network of analog components that satisfies the one or more coupling constraints. The method also includes calculating a weight matrix of the equivalent loosely coupled network based on the weights of the trained neural network. Each element of the weight matrix represents a respective connection between the analog components of the equivalent loosely coupled network.

[0030] In some implementations, converting a neural network topology into an equivalent loosely coupled network of analog components involves determining the number of possible input couplings N, subject to one or more coupling constraints. i and output coupling degree N o This includes deriving

[0031] In some implementations, the neural network topology includes at least one densely connected layer with K inputs and L outputs and a weight matrix U. In such a case, transforming the at least one densely connected layer may be performed by transforming the at least one densely connected layer to a layer with a degree of input connectivity of N i and the output coupling is not more than N o K inputs, L outputs, and

number

[0032] In some implementations, the neural network topology includes at least one densely connected layer with K inputs and L outputs and a weight matrix U. In such cases, transforming the at least one densely connected layer includes transforming the K inputs, L outputs, and

number

[0033] In some implementations, the neural network topology has K inputs and L outputs, P i The maximum input coupling, P o It includes a single sparsely connected layer with a maximum output connection of N and a weight matrix of U, where missing connections are represented by zeros. In such a case, transforming a single sparsely connected layer into a layer with input connections of N i and the output coupling is not more than N o There are K inputs, L outputs, and each layer m has a corresponding weight matrix U m represented by

number

[0034] In some implementations, the neural network topology includes a convolutional layer with K inputs and L outputs. In such cases, converting the neural network topology into an equivalent loosely connected network of analog components involves converting the convolutional layer into an equivalent loosely connected network of analog components with K inputs, L outputs, P i The maximum input coupling and P o This involves decomposing the layer into a single loosely connected layer with a maximum output coupling of P i ≦N iand P o ≦N o。

[0035]

[0035] In some implementations, a weight matrix is ​​used to generate a rough model for implementing an equivalent loosely coupled network.

[0036] In some implementations, the neural network topology includes a recurrent neural layer. In such cases, converting the neural network topology into an equivalent loosely coupled network of analog components includes converting the recurrent neural layer into one or more tightly or loosely coupled layers with signal delay connections.

[0037] In some implementations, the neural network topology includes a recurrent neural layer. In such cases, converting the neural network topology to an equivalent loosely connected network of analog components includes decomposing the recurrent neural layer into several layers, at least one of which is equivalent to a densely or loosely connected layer with K inputs and L outputs and a weight matrix U, and missing connections are represented by zeros.

[0038] In some implementations, the neural network topology has K inputs, a weight vector U∈R K and a single-layer perceptron with computational neurons with activation function F. In such cases, converting a neural network topology to an equivalent loosely connected network of analog components involves (i) deriving the connectivity N of the equivalent loosely connected network subject to one or more coupling constraints, and (ii) calculating the connectivity N of the equivalent loosely connected network according to Eq.

number

number

[0039] In some implementations, the neural network topology includes a single-layer perceptron with K inputs, L computational neurons, and a weight matrix V including a row of weights for each of the L computational neurons. In such a case, converting the neural network topology to an equivalent loosely coupled network of analog components includes: (i) deriving the connectivity N of the equivalent loosely coupled network according to one or more coupling constraints; (ii) calculating the equation

number

number

[0040] In some implementations, the neural network topology includes a multi-layer perceptron with K inputs and S layers, where each layer i of the S layers includes computational neurons L iAnd, L i a corresponding weight matrix V containing a row of weights for each computational neuron i In such cases, converting a neural network topology into an equivalent loosely coupled network of analog components involves (i) deriving the degree of connectivity N of the equivalent loosely coupled network subject to one or more coupling constraints, (ii) converting a multilayer perceptron into a network with Q=Σ i=1,s (L i ) single-layer perceptron networks, each single-layer perceptron network including a respective one of the Q computational neurons. Decomposing the multilayer perceptron includes duplicating one or more of the K inputs shared by the Q computational neurons; and (iii) for each single-layer perceptron network of the Q single-layer perceptron networks, (a) Eq.

number

number

number

[0041] In some implementations, the neural network topology includes a convolutional neural network (CNN) with K inputs and S layers, where each layer i of the S layers includes computational neurons L i And, L i The corresponding weight matrix V contains a row of the weights of each computational neuron. i In such cases, converting a neural network topology into an equivalent loosely coupled network of analog components involves (i) deriving the degree of connectivity N of the equivalent loosely coupled network subject to one or more coupling constraints, (ii) converting the CNN into a network with Q=Σ i=1,S (L i) single-layer perceptron networks, each single-layer perceptron network including a respective one of the Q computational neurons. Decomposing the CNN includes duplicating one or more of the K inputs shared by the Q computational neurons; and (iii) for each single-layer perceptron network of the Q single-layer perceptron networks, (a)

number

number

number

[0042] In some implementations, the neural network topology includes a layer L with K inputs and K neurons. p , layer L with L neurons n and weight matrix W∈R L×K where R is the set of real numbers and L p Each neuron in layer L n connected to each neuron in layer L n Each neuron in layer L performs an activation function F, which n The output of is the expression Y for the input x. o =F(Wx). In such cases, transforming a neural network topology into an equivalent loosely connected network of analog components is a trapezoidal transformation, which (i) computes the number of possible input connections N, subject to one or more connection constraints. I >1 and possible output coupling N O > 1 and (ii) K·L <L·N I +K·N O a layer LA with K analog neurons that perform the identity activation function according to the determination that p , which runs the identity activation function

number

[0043] In some implementations, performing the trapezoidal transform is such that K·L≧L·N· I +K·N O According to this determination, (i) layer L p Divide K' L ≧ L N I +K'·N O A sublayer L with K' neurons is p1 and a sublayer L with (K-K') neurons p2and (ii) a sublayer L with K′ neurons. p1 (iii) performing the constructing and generating steps for a sublayer L having K-K neurons; p2 and recursively performing the dividing, constructing and generating steps for .

[0044] In some implementations, the neural network topology includes a multi-layer perceptron network. In such cases, the method further includes iteratively performing a trapezoidal transformation for each pair of successive layers of the multi-layer perceptron network and calculating a weight matrix of an equivalent loosely coupled network.

[0045] In some implementations, the neural network topology includes a recurrent neural network (RNN) that includes (i) a linear combination of two fully connected layers, (ii) element-wise summation, and (iii) a nonlinear function calculation. In such cases, the method further includes performing a trapezoidal transformation on the (i) two fully connected layers and (ii) the nonlinear function calculation, and calculating a weight matrix of an equivalent loosely connected network.

[0046] In some implementations, the neural network topology includes a long short-term memory (LSTM) network or a gated recurrent unit (GRU) network that includes (i) a linear combination of multiple fully connected layers, (ii) element-wise summation, (iii) a Hadamard product, and (iv) multiple nonlinear function calculations. In such cases, the method further includes performing a trapezoidal transform for the (i) multiple fully connected layers and (ii) multiple nonlinear function calculations and calculating a weight matrix of an equivalent loosely connected network.

[0047] In some implementations, the neural network topology includes a convolutional neural network (CNN) including (i) a plurality of partially connected layers and (ii) one or more fully connected layers. In such cases, the method further includes (i) converting the plurality of partially connected layers into equivalent fully connected layers by inserting missing connections with zero weights, and (ii) for each pair of successive layers of the equivalent fully connected layer and the one or more fully connected layers, iteratively performing a trapezoidal transformation and calculating a weight matrix of an equivalent loosely connected network.

[0048] In some implementations, the neural network topology has K input, L output neurons and a weight matrix U∈R L×K where R is a set of real numbers, and each output neuron implements an activation function F. In such cases, converting a neural network topology into an equivalent loosely connected network of analog components involves performing an approximation transformation that includes: (i) determining the number of possible input connections N, subject to one or more connection constraints; I >1 and possible output coupling N O > 1, (ii) the set

number

number

number

number

number

[0049] In some implementations, the neural network topology has K inputs, S layers, and L i=1、S A multilayer perceptron with computational neurons and the weight matrix of the ith layer

number

[0050]

[0050] In another aspect, a method for hardware realization of a neural network is provided according to some implementation aspects. The method includes obtaining a neural network topology and weights of a trained neural network. The method also includes converting the neural network topology into an equivalent analog network of analog components including a plurality of operational amplifiers and a plurality of resistors. Each operational amplifier represents an analog neuron of the equivalent analog network, and each resistor represents a connection between two analog neurons. The method also includes calculating a weight matrix of the equivalent analog network based on the weights of the trained neural network. Each element of the weight matrix represents a respective connection. The method also includes generating a resistance matrix of the weight matrix. Each element of the resistance matrix corresponds to a respective weight of the weight matrix and represents a resistance value.

[0051] In some implementations, generating the resistance matrix of the weight matrix includes (i) determining a predetermined range of possible resistance values ​​{R 最小 ,R 最大} and obtain the initial base resistance value R within a given range. ベース and (ii) selecting a set of limited-length resistance values ​​within a predetermined range, the set being a set of {R i , R j} for all combinations of the range [-Rベース ,R ベース Possible weights in

number

number

[0052] In some implementations, the predetermined range of possible resistance values ​​includes resistances according to nominal series E24 in the range 100 KΩ to 1 MΩ.

[0053] In some implementations, R + and R - is chosen independently for each layer of the equivalent analog network.

[0054] In some implementations, R + and R - is chosen independently for each analog neuron in the equivalent analog network.

[0055] In some implementations, the first one or more weights and the first one or more inputs of the weight matrix represent one or more connections to a first operational amplifier of an equivalent analog network. In such cases, the method further includes, before generating the resistance matrix, (i) modifying the first one or more weights by a first value, and (ii) configuring the first operational amplifier to multiply a linear combination of the first one or more weights and the first one or more inputs by the first value before performing the activation function.

[0056]

[0056] In some implementations, the method includes (i) obtaining a predetermined range of weights and (ii) updating a weight matrix according to the predetermined range of weights so that an equivalent analog network produces an output similar to a neural network trained on the same input.

[0057]

[0057] In some implementations, the trained neural network is trained such that each layer of the neural network topology has quantized weights.

[0058]

[0058] In some implementations, the method further includes retraining the trained neural network to reduce its sensitivity to errors in weights or resistor values ​​that would cause an equivalent analog network to produce a different output compared to the trained neural network.

[0059]

[0059] In some implementations, the method further includes retraining the trained neural network to minimize weights in any layer that are greater than the average absolute weight of that layer by a threshold greater than a predetermined threshold.

[0060]

[0060] In another aspect, a method for hardware realization of a neural network is provided according to some implementation aspects. The method includes obtaining a neural network topology and weights of a trained neural network. The method also includes converting the neural network topology into an equivalent analog network of analog components including a plurality of operational amplifiers and a plurality of resistors. Each operational amplifier represents an analog neuron of the equivalent analog network, and each resistor represents a connection between two analog neurons. The method also includes calculating a weight matrix of the equivalent analog network based on the weights of the trained neural network. Each element of the weight matrix represents a respective connection. The method also includes generating a resistance matrix of the weight matrix. Each element of the resistance matrix corresponds to a respective weight in the weight matrix. The method also includes pruning the equivalent analog network to reduce the number of the plurality of operational amplifiers or the plurality of resistors based on the resistance matrix to obtain an optimized analog network of analog components.

[0061]

[0061] In some implementations, pruning the equivalent analog network includes replacing resistors corresponding to one or more elements of the resistance matrix having a resistance value below a predetermined minimum threshold resistance value with a conductor.

[0062]

[0062] In some implementations, pruning the equivalent analog network includes removing one or more connections of the equivalent analog network that correspond to one or more elements of the resistance matrix that exceed a predetermined maximum threshold resistance value.

[0063]

[0063] In some implementations, pruning the equivalent analog network includes removing one or more connections of the equivalent analog network that correspond to one or more elements of the weight matrix that are approximately zero.

[0064]

[0064] In some implementations, pruning the equivalent analog network further includes removing one or more analog neurons of the equivalent analog network that do not have any input connections.

[0065]

[0065] In some implementations, pruning the equivalent analog network includes (i) ranking analog neurons of the equivalent analog network based on detecting use of the analog neurons when performing calculations on one or more data sets, (ii) selecting one or more analog neurons of the equivalent analog network based on the ranking, and (iii) removing one or more analog neurons from the equivalent analog network.

[0066]

[0066] In some implementations, detecting the use of analog neurons includes (i) establishing a model of an equivalent analog network using modeling software, and (ii) measuring the propagation of analog signals by generating calculations of one or more data sets using the model.

[0067]

[0067] In some implementations, detecting the use of analog neurons includes (i) establishing a model of an equivalent analog network using modeling software, and (ii) measuring the output signal of the model by generating calculations of one or more data sets using the model.

[0068]

[0068] In some implementations, detecting the use of analog neurons includes (i) establishing a model of an equivalent analog network using modeling software, and (ii) measuring the power consumed by the analog neurons by generating calculations of one or more data sets using the model.

[0069]

[0069] In some implementations, the method further includes recalculating a weight matrix of the equivalent analog network after pruning the equivalent analog network and before generating one or more lithography masks for fabricating a circuit that implements the equivalent analog network, and updating a resistance matrix based on the recalculated weight matrix.

[0070]

[0070] In some implementations, the method further includes, for each analog neuron of the equivalent analog network, (i) calculating a respective bias value for each analog neuron based on the weights of the trained neural network while calculating the weight matrix; (ii) removing each analog neuron from the equivalent analog network pursuant to a determination that the respective bias value is above a predetermined maximum bias threshold; and (iii) replacing each analog neuron with a linear junction in the equivalent analog network pursuant to a determination that the respective bias value is below a predetermined minimum bias threshold.

[0071]

[0071] In some implementations, the method further includes reducing the number of neurons of the equivalent analog network by increasing the number of connections from one or more analog neurons of the equivalent analog network before generating the weight matrix.

[0072]

[0072] In some implementations, the method further includes, before converting the neural network topology, pruning the trained neural network using a neural network pruning technique so that an equivalent analog network contains less than a predetermined number of analog components, and updating the neural network topology and weights of the trained neural network.

[0073]

[0073] In some implementations, pruning is performed iteratively, taking into account the accuracy or level of match of the outputs between the trained neural network and the equivalent analog network.

[0074]

[0074] In some implementations, the method further includes performing network knowledge extraction before converting the neural network topology to an equivalent analog network.

[0075] In another aspect, an integrated circuit according to some implementations is provided. The integrated circuit includes an analog network of analog components fabricated by a method including: (i) obtaining a neural network topology and weights of a trained neural network; (ii) converting the neural network topology into an equivalent analog network of analog components including a plurality of operational amplifiers and a plurality of resistors, each operational amplifier representing a respective analog neuron and each resistor representing a respective connection between a respective first analog neuron and a respective second analog neuron; (iii) calculating a weight matrix of the equivalent analog network based on the weights of the trained neural network, each element of the weight matrix representing a respective connection; (iv) generating a resistance matrix of the weight matrix, each element of the resistance matrix corresponding to a respective weight in the weight matrix; (v) generating one or more lithography masks for fabricating a circuit implementing the equivalent analog network of analog components based on the resistance matrix; and (vi) fabricating the circuit based on the one or more lithography masks using a lithography process.

[0076]

[0076] In some implementations, the integrated circuit further includes one or more digital-to-analog converters configured to generate an analog input of an equivalent analog network of analog components based on one or more digital.

[0077]

[0077] In some implementations, the integrated circuit further includes an analog signal sampling module configured to process one-dimensional or two-dimensional analog input at a sampling frequency based on the number of inferences of the integrated circuit.

[0078]

[0078] In some implementations, the integrated circuit further includes a voltage conversion module for scaling down or up the analog signal to fit the operating range of the multiple operational amplifiers.

[0079]

[0079] In some implementations, the integrated circuit further includes a tact signal processing module configured to process one or more frames acquired from the CCD camera.

[0080] In some implementations, the trained neural network is a long short-term memory (LSTM) network. In such cases, the integrated circuit further includes one or more clock modules for synchronizing signal timing and enabling time series processing.

[0081] In some implementations, the integrated circuit further includes one or more analog-to-digital converters configured to generate a digital signal based on the output of the equivalent analog network of analog components.

[0082]

[0082] In some implementations, the integrated circuit further includes one or more signal processing modules configured to process one-dimensional or two-dimensional analog signals obtained from the edge application.

[0083] In some implementations, the trained neural network is trained using a training data set including signals from an array of gas sensors for different gas mixtures for selective sensing of different gases in a gas mixture containing a predetermined amount of the gas to be detected. In such cases, the neural network topology is a one-dimensional deep convolutional neural network (1D-DCNN) designed to detect three binary gas components based on measurements by 16 gas sensors, and includes a 1D convolution block for each of the 16 sensors, three shared or common 1D convolution blocks, and three dense layers. In such cases, an equivalent analog network includes (i) up to 100 input / output connections per analog neuron, (ii) delay blocks that can introduce delays of any number of time steps, (iii) a signal limit of 5, (iv) 15 layers, (v) approximately 100,000 analog neurons, and (vi) approximately 4,900,000 connections.

[0084] In some implementations, a trained neural network is trained using a training dataset containing thermal aging time-series data of different MOSFETs to predict the remaining useful life (RUL) of MOSFET devices. In such a case, the neural network topology includes four LSTM layers with 64 neurons in each layer, followed by two dense layers with 64 neurons and one neuron, respectively. In such a case, the equivalent analog network includes (i) a maximum of 100 input / output connections per analog neuron, (ii) a signal limit of 5, (iii) 18 layers, (iv) 3,000-3,200 analog neurons, and (v) 123,000-124,000 connections.

[0085] In some implementations, a trained neural network has been trained using a training dataset containing time-series data including discharge and temperature data during continuous use of different commercially available Li-ion batteries to monitor the state of health (SOH) and state of charge (SOC) of lithium-ion batteries for use in a battery management system (BMS). In such cases, the neural network topology includes an input layer, two LSTM layers with 64 neurons each, followed by an output dense layer with two neurons for generating SOC and SOH values. In such cases, the equivalent analog network includes (i) a maximum of 100 input / output connections per analog neuron, (ii) a signal limit of 5, (iii) nine layers, (iv) 1,200-1,300 analog neurons, and (v) 51,000-52,000 connections.

[0086] In some implementations, a trained neural network has been trained using a training dataset containing time-series data including discharge and temperature data during continuous use of different commercially available Li-ion batteries to monitor the state of health (SOH) of lithium-ion batteries for use in a battery management system (BMS). In such a case, the neural network topology includes an input layer with 18 neurons, a simple recurrent layer with 100 neurons, and a dense layer with 1 neuron. In such a case, the equivalent analog network includes (i) a maximum of 100 input / output connections per analog neuron, (ii) a signal limit of 5, (iii) four layers, (iv) 200-300 analog neurons, and (v) 2,200-2,400 connections.

[0087] In some implementations, the trained neural network is trained using a training dataset including spoken commands to identify spoken commands. In such cases, the neural network topology is a deep-separable convolutional neural network (DS-CNN) layer with one neuron. In such cases, the equivalent analog network includes (i) a maximum of 100 input / output connections per analog neuron, (ii) a signal limit of 5, (iii) 13 layers, (iv) approximately 72,000 analog neurons, and (v) approximately 2.6 million connections.

[0088] In some implementations, the trained neural network is trained using a training dataset including photoplethysmography (PPG) data, accelerometer data, temperature data, and galvanic skin response signal data of different individuals performing various physical activities over a predetermined period of time, as well as baseline heart rates obtained from an ECG sensor for determining pulse rates during physical exercise based on the PPG sensor data and the three-axis accelerometer data. In such cases, the neural network topology includes two Conv1D layers, each with 16 filters and 20 kernels, two LSTM layers, each with 16 neurons, and two dense layers, each with 16 neurons and 1 neuron, that perform time-series convolution. In such a case, the equivalent analog network would include (i) delay blocks to generate any number of time steps, (ii) up to 100 input / output connections per analog neuron, (iii) a signal limit of 5, (iv) 16 layers, (v) 700-800 analog neurons, and (vi) 12,000-12,500 connections.

[0089] In some implementations, the trained neural network is trained to classify different objects based on pulsed Doppler radar signals, in which case the neural network topology includes a multi-scale LSTM neural network.

[0090] In some implementations, the trained neural network is trained to perform human activity type recognition based on inertial sensor data. In such cases, the neural network topology includes three channel-wise convolutional networks, each with a convolutional layer of 12 filters and a kernel dimension of 64, followed by a maximum pruning layer and two common dense layers of 1024 neurons and N neurons, respectively, where N is the number of classes. In such cases, the equivalent analog network includes (i) delay blocks to generate any number of time steps, (ii) up to 100 input / output connections per analog neuron, (iii) an output layer of 10 analog neurons, (iv) a signal limit of 5, (v) 10 layers, (vi) 1,200-1,300 analog neurons, and (vi) 20,000-21,000 connections.

[0091]

[0091] In some implementations, the trained neural network is further trained to detect abnormal patterns of human activity based on accelerometer data that is merged with heart rate data using a convolution operation.

[0092]

[0092] In another aspect, a method for generating a library for hardware realization of a neural network is provided. The method includes obtaining a plurality of neural network topologies, each corresponding to a respective neural network. The method also includes converting each neural network topology into a respective equivalent analog network of analog components. The method also includes generating a plurality of lithography masks for fabricating a plurality of circuits, each circuit implementing a respective equivalent analog network of analog components.

[0093] In some implementations, the method further includes obtaining a new neural network topology and weights for the trained neural network. The method also includes selecting one or more lithography masks from the plurality of lithography masks based on comparing the new neural network topology with the plurality of neural network topologies. The method also includes calculating a weight matrix for a new equivalent analog network based on the weights. The method also includes generating a resistance matrix for the weight matrix. The method also includes generating a new lithography mask for fabricating a circuit implementing the new equivalent analog network based on the resistance matrix and the one or more lithography masks.

[0094]

[0094] In some implementations, the new neural network topology includes multiple sub-network topologies, and selecting one or more lithography masks is further based on comparing each sub-network topology with each network topology of the multiple network topologies.

[0095] In some implementations, one or more subnetwork topologies of the plurality of subnetwork topologies fail to compare with any network topology of the plurality of network topologies. In such cases, the method further includes (i) converting each subnetwork topology of the one or more subnetwork topologies into a respective equivalent analog subnetwork of analog components, and (ii) generating one or more lithography masks for fabricating one or more circuits, each circuit of the one or more circuits implementing a respective equivalent analog subnetwork of analog components.

[0096]

[0096] In some implementations, converting each network topology into a respective equivalent analog network includes (i) decomposing each network topology into multiple subnetwork topologies, (ii) converting each subnetwork topology into a respective equivalent analog subnetwork of analog components, and (iii) configuring each equivalent analog subnetwork to obtain a respective equivalent analog network.

[0097]

[0097] In some implementations, decomposing the respective network topologies includes identifying one or more layers of the respective network topologies as multiple sub-network topologies.

[0098]

[0098] In some implementations, each circuit is obtained by (i) generating a schematic diagram of an equivalent analog network for each of the analog components, and (ii) generating a respective circuit layout design based on the schematic diagram.

[0099] In some implementations, the method further includes combining one or more circuit layout designs before generating a plurality of lithography masks for fabricating the plurality of circuits.

[0100]

[0100] In another aspect, a method for optimizing energy efficiency of an analog neuromorphic circuit according to some implementations is provided. The method includes obtaining an integrated circuit implementing an analog network of analog components including a plurality of operational amplifiers and a plurality of resistors. The analog network represents a trained neural network, each operational amplifier representing a respective analog neuron, and each resistor representing a respective connection between a respective first analog neuron and a respective second analog neuron. The method also includes generating inferences using the integrated circuit for a plurality of test inputs including simultaneously transferring signals from one layer of the analog network to a subsequent layer. While generating inferences using the integrated circuit, the method also includes: (i) determining whether levels of signal outputs of the plurality of operational amplifiers are balanced; and (ii) pursuant to a determination that levels of signal outputs are balanced, (a) determining analog neurons of an active set of the analog network that influence signal formation for signal propagation; and (b) powering off one or more analog neurons of the analog network that are different from the analog neurons of the active set for a predetermined period of time.

[0101]

[0101] In some implementations, determining the active set of analog neurons is based on calculating the delay of signal propagation through the analog network.

[0102]

[0102] In some implementations, determining the active set of analog neurons is based on detecting the propagation of signals through the analog network.

[0103]

[0103] In some implementations, the trained neural network is a feedforward neural network, the analog neurons of the active set belong to an active layer of the analog network, and turning off the power supply includes turning off the power supply of one or more layers prior to the active layer of the analog network.

[0104]

[0104] In some implementations, the predetermined period is calculated based on simulating the propagation of a signal through an analog network, taking into account signal delays.

[0105] In some implementations, the trained neural network is a recurrent neural network (RNN) and the analog network further includes one or more analog components other than the plurality of operational amplifiers and the plurality of resistors. In such cases, the method further includes, pursuant to determining that the levels of the signal outputs are balanced, turning off power to the one or more analog components for a predetermined period of time.

[0106]

[0106] In some implementations, the method further includes turning on the power supply of one or more analog neurons of the analog network after a predetermined period of time.

[0107]

[0107] In some implementations, determining whether the signal output levels of multiple operational amplifiers are balanced is based on detecting whether one or more operational amplifiers of the analog network are outputting above a predetermined threshold signal level.

[0108]

[0108] In some implementations, the method further includes repeating turning off for a predetermined period and turning on for a predetermined period the active set of analog neurons while generating inferences.

[0109]

[0109] In some implementations, the method further includes: (i) for each inference cycle, in accordance with a determination that the level of the signal output is balanced, (a) determining, during a first time interval, analog neurons of a first layer of the analog network that influence signal formation for signal propagation; (b) turning off power supplies to a first one or more analog neurons of the analog network before the first layer for a predetermined period of time; and (ii) turning off power supplies to a second one or more analog neurons of the analog network, including analog neurons of the first layer and the first one or more analog neurons, for a predetermined period of time during a second time interval following the first time interval.

[0110]

[0110] In some implementations, the one or more analog neurons are analog neurons of a first one or more layers of an analog network, and the active set of analog neurons are analog neurons of a second layer of the analog network, which second layer of the analog network is different from the first one or more layers.

[0111] In another aspect, a method for recognizing human activities is provided. The method includes tracking a user's activity using one or more sensors, including acquiring a plurality of electrical signals from the one or more sensors. The method also includes forming a feature vector by extracting a plurality of features from the plurality of electrical signals. The features correspond to inputs of a neural network model trained to generate a plurality of descriptors for a plurality of predefined human activities. The method also includes applying an analog neuro-computing hardware device to the feature vector to generate an embedding vector that specifies the descriptors. The analog neuro-computing hardware device implements the trained neural network model. The method also includes applying a trained machine learning classifier to the embedding vector to classify the user's activity as one of the predefined human activities.

[0112]

[0112] In some implementations, the trained neural network model is an autoencoder that includes an encoder and a decoder.

[0113] In some implementations, the trained machine learning classifier is a KNN (K Nearest Neighbor) classifier. In some implementations, the number of neighbors of the KNN classifier is equal to 5. In some implementations, the trained machine learning classifier is trained separately for each of the predefined human activities using binary classification.

[0114]

[0114] In some implementations, the one or more sensors include one or more of an IMU, a camera, a microphone, and a biofeedback device.

[0115]

[0115] In some implementations, the method further includes smoothing the output of the trained machine learning classifier to obtain basic classes of activity.

[0116]

[0116] In some implementations, the analog neurocomputing hardware device is fabricated by steps including obtaining a neural network topology and weights of a trained neural network model; converting the neural network topology into an equivalent analog network of analog components; calculating a weight matrix of the equivalent analog network based on the weights of the trained neural network model, where each element of the weight matrix represents one or more connections between analog components of the equivalent analog network; generating a schematic model for implementing the equivalent analog network based on the weight matrix, including selecting component values ​​for the analog components; and fabricating an integrated circuit according to the schematic model using a lithography process.

[0117]

[0117] In some implementations, generating the general model includes generating a resistance matrix of the weight matrix, where each element of the resistance matrix corresponds to a respective weight in the weight matrix and represents a resistance value.

[0118]

[0118] In some implementations, the trained machine learning classifier is implemented using one or more digital components, and the trained machine learning classifier can be retrained for a new user.

[0119] In another aspect, a method for recognizing human activities is provided. The method includes acquiring a sequence of electrical signals from one or more sensors that track a user's activity. The method also includes forming a plurality of feature vectors by extracting features from the sequence of electrical signals. The features correspond to inputs of a neural network model trained to generate a plurality of descriptors for a plurality of predefined human activities. The method also includes applying an analog neuro-computing hardware device to the plurality of feature vectors to generate a plurality of embedding vectors, each of which specifies a corresponding descriptor. The method also includes using the plurality of embedding vectors to classify the user's activity as one of the predefined human activities.

[0120] In some implementations, the method further includes receiving, from the user, a set of descriptors describing a particular physical activity, and classifying the user's activity as one of the particular physical activities using the set of descriptors and the plurality of embedding vectors. In some implementations, the method further includes generating statistics of the user's personal daily routine based on the classifying the user's activity as one of the particular physical activities.

[0121]

[0121] In some implementations, the method further includes storing, for a user, a plurality of embedding vectors as describing a particular activity, and using the plurality of embedding vectors to classify subsequent activities of the user as the particular activity.

[0122]

[0122] In some implementations, the method further includes receiving a set of descriptors describing a particular activity from a trainer different from the user, and providing feedback to the user if the activity matches the particular activity based on the plurality of embedding vectors and the set of descriptors.

[0123]

[0123] In another aspect, a human activity recognition device is provided. The device includes an integrated circuit for human activity recognition. The integrated circuit includes an analog network of analog components configured to implement a trained neural network model, the trained neural network model being trained to generate a plurality of descriptors for a plurality of predefined human activities based on a plurality of features extracted from a plurality of electrical signals from one or more sensors. The device also includes one or more digital components configured to classify the human activity as one of the plurality of predefined human activities according to the plurality of descriptors generated by the integrated circuit.

[0124]

[0124] In some implementations, the human activity recognition device further includes one or more sensors configured to collect a plurality of electrical signals during the human activity.

[0125]

[0125] In some implementations, the trained neural network model is an autoencoder that includes an encoder and a decoder.

[0126]

[0126] In some implementations, one or more digital components implement a trained machine learning classifier that is a KNN (K Nearest Neighbor) classifier that can be retrained.

[0127]

[0127] In some implementations, the number of neighbors in the KNN classifier is equal to five.

[0128]

[0128] In some implementations, the trained machine learning classifier is trained separately for each of a plurality of predefined human activities using binary classification.

[0129]

[0129] In some implementations, the one or more digital components are further configured to smooth the output of the trained machine learning classifier to obtain fundamental classes of activity.

[0130]

[0130] In some implementations, the one or more sensors include one or more of an IMU, a camera, a microphone, and a biofeedback device.

[0131] In some implementations, the integrated circuit is fabricated by steps including: obtaining a neural network topology and weights of a trained neural network model; converting the neural network topology into an equivalent analog network of analog components; calculating a weight matrix of the equivalent analog network based on the weights of the trained neural network model, where each element of the weight matrix represents one or more connections between analog components of the equivalent analog network; generating a schematic model for implementing the equivalent analog network based on the weight matrix, including selecting component values ​​for the analog components; and fabricating the integrated circuit according to the schematic model using a lithography process. In some implementations, generating the schematic model includes generating a resistance matrix of the weight matrix. Each element of the resistance matrix corresponds to a respective weight in the weight matrix and represents a resistance value.

[0132] In some implementations, the computer system includes one or more processors, a memory, and a display. One or more programs include instructions for performing any of the methods described herein.

[0133] In some implementations, a non-transitory computer-readable storage medium stores one or more programs configured for execution by a computer system having one or more processors, a memory, and a display, the one or more programs including instructions for performing any of the methods described herein.

[0134]

[0134] Thus, methods, systems and devices are disclosed that can be used for the hardware implementation of trained neural networks.

[0135] BRIEF DESCRIPTION OF THE DRAWINGS

[0135] To better understand the above-described systems, methods and graphical user interfaces, as well as additional systems, methods and graphical user interfaces that provide data visualization analysis and data preparation, please refer to the following description of implementation aspects in conjunction with the following drawings, in which like reference numbers refer to corresponding parts throughout the drawings: [Brief explanation of the drawings]

[0136] [Figure 1A]

[0136] A block diagram of a system for hardware realization of a trained neural network using analog components, according to some implementation aspects. [Figure 1B]

[0136] FIG. 1B is a block diagram of an alternative representation of the system of FIG. 1A for hardware realization of a trained neural network using analog components, according to some implementations. [Figure 1C]

[0136] FIG. 1B is a block diagram of another representation of the system of FIG. 1A for hardware realization of a trained neural network using analog components, according to some implementations. [Figure 2A]

[0137] FIG. 1 is a system diagram of a computing device according to some implementations. [Figure 2B]

[0137] Optional modules of a computing device are illustrated according to some implementations. [Figure 3A]

[0138] 1 illustrates an exemplary process for generating a schematic model of an analog network corresponding to a trained neural network, according to some implementations. [Figure 3B]

[0138] An exemplary manual prototyping process used to generate a target chip model, according to some implementations, is shown. [Figure 4A]

[0139] 1 illustrates an example of a neural network converted into a mathematically equivalent analog network, according to some implementations. [Figure 4B]

[0139] An example of a neural network converted into a mathematically equivalent analog network according to some implementations is shown. [Figure 4C]

[0139] An example of a neural network converted into a mathematically equivalent analog network according to some implementations is shown. [Figure 5]

[0140] 1 illustrates an example mathematical model of a neuron, according to some implementations. [Figure 6A]

[0141] 1 illustrates an exemplary process for an analog hardware realization of a neural network for computing the XOR of input values, according to some implementations. [Figure 6B]

[0141] Illustrates an exemplary process for an analog hardware realization of a neural network for computing the XOR of input values, according to some implementations. [Figure 6C]

[0141] Illustrates an exemplary process for an analog hardware realization of a neural network for computing the XOR of input values, according to some implementations. [Figure 7]

[0142] 1 illustrates an exemplary perceptron, according to some implementations. [Figure 8]

[0143] 1 illustrates an exemplary pyramid-neural network, according to some implementations. [Figure 9]

[0144] 1 illustrates an exemplary pyramidal single neural network, according to some implementations. [Figure 10]

[0145] 1 illustrates an example of a transformed neural network, according to some implementations. [Figure 11A]

[0146] 1 illustrates the application of the T-transform algorithm to a single-layer neural network, according to some implementations. [Figure 11B]

[0146] We illustrate the application of the T-transform algorithm to a single-layer neural network, according to some implementations. [Figure 11C]

[0146] We illustrate the application of the T-transform algorithm to a single-layer neural network, according to some implementations. [Figure 12]

[0147] 1 illustrates an exemplary recurrent neural network (RNN), according to some implementations. [Figure 13A]

[0148] FIG. 1 is a block diagram of an LSTM neuron, according to some implementations. [Figure 13B]

[0149] 1 illustrates a delay block according to some implementations. [Figure 13C]

[0150] 1 is a neuron schema of an LSTM neuron, according to some implementations. [Figure 14A]

[0151] FIG. 1 is a block diagram of a GRU neuron, according to some implementations. [Figure 14B]

[0152] 1 is a neuron schema of a GRU neuron, according to some implementations. [Figure 15A]

[0153] 1 is a neuron schema of a variant of a single Conv1D filter, according to some implementations. [Figure 15B]

[0153] Neuron schema of a variant of a single Conv1D filter, according to some implementations. [Figure 16]

[0154] 1 illustrates an example architecture of a transformed neural network, according to some implementations. [Figure 17A]

[0155] 1 provides an example chart illustrating the dependency between output error and classification error or weight error, according to some implementations. [Figure 17B]

[0155] An exemplary chart illustrating the dependency between output error and classification error or weight error according to some implementations is provided. [Figure 17C]

[0155] An exemplary chart illustrating the dependency between output error and classification error or weight error according to some implementations is provided. [Figure 18]

[0156] 1 provides an exemplary scheme of a neuron model used for resistor quantization, according to some implementations. [Figure 19A]

[0157] 1 illustrates a schematic diagram of an operational amplifier fabricated on CMOS, according to some implementations. [Figure 19B]

[0157] A table of descriptions of the exemplary circuit shown in Figure 19A is shown, according to some implementations. [Figure 20A]

[0158] 1 illustrates a schematic diagram of an LSTM block, according to some implementations. [Figure 20B]

[0158] A schematic diagram of an LSTM block according to some implementations is shown. [Figure 20C]

[0158] A schematic diagram of an LSTM block according to some implementations is shown. [Figure 20D]

[0158] A schematic diagram of an LSTM block according to some implementations is shown. [Figure 20E]

[0158] A schematic diagram of an LSTM block according to some implementations is shown. [Figure 20F]

[0158] A table of descriptions of the example circuits shown in Figures 20A-20D is provided, according to some implementations. [Figure 21A]

[0159] 1 illustrates a schematic diagram of a multiplier block according to some implementations. [Figure 21B]

[0159] A schematic diagram of a multiplier block according to some implementations is shown. [Figure 21C]

[0159] A schematic diagram of a multiplier block according to some implementations is shown. [Figure 21D]

[0159] A schematic diagram of a multiplier block according to some implementations is shown. [Figure 21E]

[0159] A schematic diagram of a multiplier block according to some implementations is shown. [Figure 21F]

[0159] A schematic diagram of a multiplier block according to some implementations is shown. [Figure 21G]

[0159] A schematic diagram of a multiplier block according to some implementations is shown. [Figure 21H]

[0159] A schematic diagram of a multiplier block according to some implementations is shown. [Figure 21I]

[0159] A schematic diagram of a multiplier block according to some implementations is shown. [Figure 21J]

[0159] Figure 21A-21I shows a table of descriptions of the schematic diagrams shown in Figures 21A-21I, according to some implementations. [Figure 22A]

[0160] 1 illustrates a schematic diagram of a sigmoid neuron, according to some implementations. [Figure 22B]

[0160] Figure 22B shows a table of descriptions for the schematic diagram shown in Figure 22A, according to some implementations. [Figure 23A]

[0161] 1 illustrates a schematic diagram of a hyperbolic tangent function block according to some implementations. [Figure 23B]

[0161] A table of descriptions for the schematic diagram shown in Figure 23A is provided, according to some implementations. [Figure 24A]

[0162] 1 illustrates a schematic diagram of a single neuron CMOS operational amplifier, according to some implementations. [Figure 24B]

[0162] A schematic diagram of a single neuron CMOS operational amplifier is shown, according to some implementations. [Figure 24C]

[0162] A schematic diagram of a single neuron CMOS operational amplifier is shown, according to some implementations. [Figure 24D]

[0162] Figure 24 shows a table of descriptions of the schematic diagrams shown in Figures 24A-24C, according to some implementations. [Figure 25A]

[0163] 1 illustrates a schematic diagram of a variant of a single neuron CMOS operational amplifier, according to some implementations. [Figure 25B]

[0163] A schematic diagram of a variant of a single neuron CMOS operational amplifier is shown, according to some implementations. [Figure 25C]

[0163] A schematic diagram of a variant of a single neuron CMOS operational amplifier is shown, according to some implementations. [Figure 25D]

[0163] A schematic diagram of a variant of a single neuron CMOS operational amplifier is shown, according to some implementations. [Figure 25E]

[0163] Figure 25A-25D shows a table of descriptions of the schematic diagrams shown in Figures 25A-25D, according to some implementations. [Figure 26A]

[0164] 1 illustrates an exemplary weight distribution histogram, according to some implementations. [Figure 26B]

[0164] An exemplary weight distribution histogram is shown, according to some implementations. [Figure 26C]

[0164] An exemplary weight distribution histogram is shown, according to some implementations. [Figure 26D]

[0164] An exemplary weight distribution histogram is shown, according to some implementations. [Figure 26E]

[0164] An exemplary weight distribution histogram is shown, according to some implementations. [Figure 26F]

[0164] An exemplary weight distribution histogram is shown, according to some implementations. [Figure 26G]

[0164] An exemplary weight distribution histogram is shown, according to some implementations. [Figure 26H]

[0164] An exemplary weight distribution histogram is shown, according to some implementations. [Figure 26I]

[0164] An exemplary weight distribution histogram is shown, according to some implementations. [Figure 26J]

[0164] An exemplary weight distribution histogram is shown, according to some implementations. [Figure 26K]

[0164] An exemplary weight distribution histogram is shown, according to some implementations. [Figure 27A]

[0165] 1 illustrates a flowchart of a method for hardware realization of a neural network, according to some implementations. [Figure 27B]

[0165] A flowchart of a method for hardware realization of a neural network according to some implementation aspects is shown. [Figure 27C]

[0165] A flowchart of a method for hardware realization of a neural network according to some implementation aspects is shown. [Figure 27D]

[0165] A flowchart of a method for hardware realization of a neural network according to some implementation aspects is shown. [Figure 27E]

[0165] A flowchart of a method for hardware realization of a neural network according to some implementation aspects is shown. [Figure 27F]

[0165] A flowchart of a method for hardware realization of a neural network according to some implementation aspects is shown. [Figure 27G]

[0165] A flowchart of a method for hardware realization of a neural network according to some implementation aspects is shown. [Figure 27H]

[0165] A flowchart of a method for hardware realization of a neural network according to some implementation aspects is shown. [Figure 27I]

[0165] A flowchart of a method for hardware realization of a neural network according to some implementation aspects is shown. [Figure 27J]

[0165] A flowchart of a method for hardware realization of a neural network according to some implementation aspects is shown. [Figure 28A]

[0166] 1 illustrates a flowchart of a method for hardware realization of a neural network according to hardware design constraints, according to some implementations. [Figure 28B]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28C]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28D]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28E]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28F]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28G]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28H]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28I]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28J]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28K]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28L]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28M]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28N]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28O]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28P]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28Q]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28R]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 28S]

[0166] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 29A]

[0167] 1 illustrates a flowchart of a method for hardware realization of a neural network according to hardware design constraints, according to some implementations. [Figure 29B]

[0167] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 29C]

[0167] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 29D]

[0167] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 29E]

[0167] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 29F]

[0167] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 30A]

[0168] 1 illustrates a flowchart of a method for hardware realization of a neural network according to hardware design constraints, according to some implementations. [Figure 30B]

[0168] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 30C]

[0168] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 30D]

[0168] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 30E]

[0168] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 30F]

[0168] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 30G]

[0168] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 30H]

[0168] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 30I]

[0168] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 30J]

[0168] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 30K]

[0168] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 30L]

[0168] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 30M]

[0168] A flowchart of a method for hardware realization of a neural network according to hardware design constraints is shown, according to some implementations. [Figure 31A]

[0169] 1 illustrates a flowchart of a method for fabricating an integrated circuit including an analog network of analog components, according to some implementations. [Figure 31B]

[0169] A flowchart of a method for fabricating an integrated circuit including an analog network of analog components is shown, according to some implementations. [Figure 31C]

[0169] A flowchart of a method for fabricating an integrated circuit including an analog network of analog components is shown, according to some implementations. [Figure 31D]

[0169] A flowchart of a method for fabricating an integrated circuit including an analog network of analog components is shown, according to some implementations. [Figure 31E]

[0169] A flowchart of a method for fabricating an integrated circuit including an analog network of analog components is shown, according to some implementations. [Figure 31F]

[0169] A flowchart of a method for fabricating an integrated circuit including an analog network of analog components is shown, according to some implementations. [Figure 31G]

[0169] A flowchart of a method for fabricating an integrated circuit including an analog network of analog components is shown, according to some implementations. [Figure 31H]

[0169] A flowchart of a method for fabricating an integrated circuit including an analog network of analog components is shown, according to some implementations. [Figure 31I]

[0169] A flowchart of a method for fabricating an integrated circuit including an analog network of analog components is shown, according to some implementations. [Figure 31J]

[0169] A flowchart of a method for fabricating an integrated circuit including an analog network of analog components is shown, according to some implementations. [Figure 31K]

[0169] A flowchart of a method for fabricating an integrated circuit including an analog network of analog components is shown, according to some implementations. [Figure 31L]

[0169] A flowchart of a method for fabricating an integrated circuit including an analog network of analog components is shown, according to some implementations. [Figure 31M]

[0169] A flowchart of a method for fabricating an integrated circuit including an analog network of analog components is shown, according to some implementations. [Figure 31N]

[0169] A flowchart of a method for fabricating an integrated circuit including an analog network of analog components is shown, according to some implementations. [Figure 31O]

[0169] A flowchart of a method for fabricating an integrated circuit including an analog network of analog components is shown, according to some implementations. [Figure 31P]

[0169] A flowchart of a method for fabricating an integrated circuit including an analog network of analog components is shown, according to some implementations. [Figure 31Q]

[0169] A flowchart of a method for fabricating an integrated circuit including an analog network of analog components is shown, according to some implementations. [Figure 32A]

[0170] 1 illustrates a flowchart of a method for generating a library for hardware realization of a neural network, according to some implementations. [Figure 32B]

[0170] A flowchart of a method for generating a library for hardware realization of a neural network according to some implementations is shown. [Figure 32C]

[0170] A flowchart of a method for generating a library for hardware realization of a neural network according to some implementations is shown. [Figure 32D]

[0170] A flowchart of a method for generating a library for hardware realization of a neural network according to some implementations is shown. [Figure 32E]

[0170] A flowchart of a method for generating a library for hardware realization of a neural network according to some implementations is shown. [Figure 33A]

[0171] 1 illustrates a flowchart of a method for optimizing the energy efficiency of an analog neuromorphic circuit (modeling a trained neural network), according to some implementations. [Figure 33B]

[0171] A flowchart of a method for optimizing the energy efficiency of an analog neuromorphic circuit (modeling a trained neural network) according to some implementations is shown. [Figure 33C]

[0171] A flowchart of a method for optimizing the energy efficiency of an analog neuromorphic circuit (modeling a trained neural network) according to some implementations is shown. [Figure 33D]

[0171] A flowchart of a method for optimizing the energy efficiency of an analog neuromorphic circuit (modeling a trained neural network) according to some implementations is shown. [Figure 33E]

[0171] A flowchart of a method for optimizing the energy efficiency of an analog neuromorphic circuit (modeling a trained neural network) according to some implementations is shown. [Figure 33F]

[0171] A flowchart of a method for optimizing the energy efficiency of an analog neuromorphic circuit (modeling a trained neural network) according to some implementations is shown. [Figure 33G]

[0171] A flowchart of a method for optimizing the energy efficiency of an analog neuromorphic circuit (modeling a trained neural network) according to some implementations is shown. [Figure 33H]

[0171] A flowchart of a method for optimizing the energy efficiency of an analog neuromorphic circuit (modeling a trained neural network) according to some implementations is shown. [Figure 33I]

[0171] A flowchart of a method for optimizing the energy efficiency of an analog neuromorphic circuit (modeling a trained neural network) according to some implementations is shown. [Figure 33J]

[0171] A flowchart of a method for optimizing the energy efficiency of an analog neuromorphic circuit (modeling a trained neural network) according to some implementations is shown. [Figure 33K]

[0171] A flowchart of a method for optimizing the energy efficiency of an analog neuromorphic circuit (modeling a trained neural network) according to some implementations is shown. [Figure 34]

[0172] 1 shows a table describing the MobileNet v1 architecture, according to some implementations. [Figure 35]

[0173] 1 illustrates an exemplary 1D convolutional neural network used for speech intelligibility, according to some implementations. [Figure 36]

[0174] 36 illustrates an example T-transform of the fully connected layer of the neural network shown in FIG. 35, according to some implementations. [Figure 37]

[0175] 36 illustrates an exemplary T-transform of the 1D convolution of the neural network shown in FIG. 35 according to some implementations. [Figure 38]

[0176] 36 illustrates an exemplary T-transform of the max pruning operator of the neural network shown in FIG. 35, according to some implementations. [Figure 39]

[0177] 1 illustrates an example of splitting a dataset into a training dataset and a validation dataset, according to some implementations. [Figure 40]

[0178] 1 shows two graph plots of activity labels according to some implementations. [Figure 41A]

[0179] 1 illustrates an exemplary graph plot for splitting an experiment into training or testing test data, according to some implementations. [Figure 41B]

[0179] An exemplary graph plot for splitting an experiment into training or testing test data according to some implementations is shown. [Figure 42A]

[0180] 1 illustrates a schematic diagram of an exemplary encoder, according to some implementations. [Figure 42B]

[0180] A schematic diagram of an exemplary decoder according to some implementations is shown. [Figure 42C]

[0180] A schematic diagram of an exemplary classifier is shown, according to some implementation aspects. [Figure 43A]

[0181] 10 shows graphs of results according to some implementations. [Figure 43B]

[0181] Graphs of results from several implementations are shown. [Figure 43C]

[0181] Graphs of results from several implementations are shown. [Figure 43D]

[0181] Graphs of results from several implementations are shown. [Figure 43E]

[0181] Graphs of results from several implementations are shown. [Figure 43F]

[0181] Graphs of results from several implementations are shown. [Figure 44A]

[0182] 10 shows graphs of results according to some implementations. [Figure 44B]

[0182] Graphs of results from several implementations are shown. [Figure 44C]

[0182] Graphs of results from several implementations are shown. [Figure 44D]

[0182] Graphs of results from several implementations are shown. [Figure 45A]

[0183] 1 shows a table illustrating predictions of a base class activity classifier model, according to some implementations. [Figure 45B]

[0183] A table showing predictions of the base class activity classifier model according to some implementations is shown. [Figure 46A]

[0184] 10 shows graphs of results according to some implementations. [Figure 46B]

[0184] Graphs of results from several implementations are shown. [Figure 46C]

[0184] Graphs of results from several implementations are shown. [Figure 46D]

[0184] Graphs of results from several implementations are shown. [Figure 46E]

[0184] Graphs of results from several implementations are shown. [Figure 46F]

[0184] Graphs of results from several implementations are shown. [Figure 47A]

[0185] 10 shows graphs of results according to some implementations. [Figure 47B]

[0185] Graphs of results from several implementations are shown. [Figure 47C]

[0185] Graphs of results from several implementations are shown. [Figure 47D]

[0185] Graphs of results from several implementations are shown. [Figure 47E]

[0185] Graphs of results from several implementations are shown. [Figure 47F]

[0185] Graphs of results from several implementations are shown. [Figure 49A]

[0186] 10 shows graphs of results according to some implementations. [Figure 49B]

[0186] Graphs of results from several implementations are shown. [Figure 49C]

[0186] Graphs of results from several implementations are shown. [Figure 49D]

[0186] Graphs of results from several implementations are shown. [Figure 48]

[0187] 47E and 47F show visualizations for t-distributed stochastic neighbor embedding (t-SNE) with exercise sets to embed different approaches of the CrossFit_sit-up exercise in the experiment shown in FIGS. 47E and 47F, according to some implementations. [Figure 50]

[0188] 1 illustrates an exemplary human activity recognition device, according to some implementations. [Figure 51]

[0189] 1 illustrates a flowchart of an exemplary method for recognizing human activity, according to some implementations. [Figure 52]

[0190] 1 illustrates a flowchart of another example method for recognizing human activity, according to some implementations. DETAILED DESCRIPTION OF THE INVENTION

[0137]

[0191] Reference will now be made to implementations, examples of which are illustrated in the accompanying figures. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, it will be apparent to one skilled in the art that the present invention may be practiced without these specific details.

[0138] Implementation Description

[0192] 1A is a block diagram of a system 100 for hardware implementation of a trained neural network using analog components, according to some implementations. The system includes converting (126) a trained neural network 102 to an analog neural network 104. In some implementations, analog integrated circuit constraints 184 constrain (146) the conversion (126) to produce the analog neural network 104. The system then derives (calculates or generates) weights 106 for the analog neural network 104 through a process sometimes referred to as weight quantization (128). In some implementations, the analog neural network includes multiple analog neurons, each represented by an analog component such as an operational amplifier, and each analog neuron is connected to another analog neuron via a connection. In some implementations, the connection is represented using a resistor that reduces current flow between two analog neurons. In some implementations, the system converts (148) the weights 106 to the resistance value 112 of the connection. The system then generates (130) one or more schematic models 108 for implementing the analog neural network 104 based on the weights 106. In some implementations, the system optimizes the resistor values ​​112 (or weights 106) to form an optimized analog neural network 114 that is further used to generate (150) the schematic model 108. In some implementations, the system generates (132) a lithography mask for connections 110 and / or generates (136) a lithography mask for analog neurons 120. In some implementations, the system fabricates (134 and / or 138) an analog integrated circuit 118 that implements the analog neural network 104. In some implementations, the system generates (152) a library of lithography masks 116 based on the lithography mask for connections 110 and / or the lithography mask for analog neurons 120. In some implementations, the system uses 154 a library 116 of lithography masks to fabricate an analog integrated circuit 118 .In some implementations, when the trained neural network 142 is retrained (142), the system regenerates (or recalculates) (144) the resistor values ​​112 (and / or weights 106), the schematic model 108, and / or the lithography masks 110 for the connections. In some implementations, the system reuses the lithography masks 120 for the analog neurons 120. That is, in some implementations, only the lithography masks 110 for the weights 106 (or the resistor values ​​112 corresponding to the changed weights) and / or the connections are regenerated. Because only the lithography masks for the connections, weights, schematic model, and / or corresponding connections are regenerated, the process (or path) for fabricating an analog integrated circuit of the retrained neural network is substantially simplified, as indicated by dashed line 156, resulting in a faster time-to-market for hardware respins of the neural network compared to traditional approaches for hardware implementation of neural networks.

[0139]

[0193] 1B is a block diagram of an alternative representation of a system 100 for hardware realization of a trained neural network using analog components, according to some implementations. The system includes training the neural network in software (156), determining connection weights, generating an electronic circuit equivalent to the neural network (158), calculating resistor values ​​corresponding to each connection weight (160), and subsequently generating a lithography mask having the resistor values ​​(162).

[0140]

[0194] 1C is a block diagram of another representation of a system 100 for hardware implementation of a trained neural network using analog components, according to some implementations. The system, according to some implementations, is distributed as a software development kit (SDK) 180. A user develops and trains a neural network (164) and inputs the trained neural net 166 into the SDK 180. The SDK estimates (168) the complexity of the trained neural net 166. If the complexity of the trained neural net can be reduced (e.g., some connections and / or neurons can be removed, some layers can be reduced, or the density of neurons can be changed), the SDK 180 prunes (178) the trained neural net and retrains (182) the neural net to obtain an updated trained neural net 166. Once the complexity of the trained neural net is reduced, the SDK 180 converts the trained neural net 166 into a sparse network of analog components (e.g., a pyramidal or trapezoidal network) (170). The SDK 180 also generates a circuit model 172 of the analog network. In some implementations, the SDK uses software simulation to estimate (176) the deviation of the output produced by the circuit model 172 compared to the trained neural network for the same input. If the estimation error exceeds a threshold error (e.g., a value set by the user), the SDK 180 prompts the user to reconfigure, redevelop, and / or retrain the neural network. In some implementations, not shown, the SDK automatically reconfigures the trained neural net 166 to reduce the estimation error. This process is repeated multiple times until the error is reduced below the threshold error. In FIG. 1C, the dashed line from block 176 ("Estimate Circuit-Induced Errors") to block 164 ("Develop and Train Neural Network") indicates a feedback loop.For example, if the pruned network does not exhibit the desired accuracy, some implementations prune the network differently until the accuracy exceeds a predetermined threshold for the given application (e.g., 98% accuracy). In some implementations, this process involves recalculating the weights, since pruning involves retraining the entire network.

[0141]

[0195] In some implementations, the components of the system 100 described above are implemented as computing modules within one or more computing devices or server systems. FIG. 2A is a system diagram of a computing device 200 according to some implementations. As used herein, the term “computing device” includes both personal devices 102 and servers. The computing device 200 typically includes one or more processing units / cores (CPUs) 202 for executing modules, programs, and / or instructions stored in memory 214 that perform processing operations, one or more network or other communication interfaces 204, and one or more communication buses 212 for interconnecting the memory 214 and these components. The communication bus 212 may include circuitry for interconnecting and controlling communications between the system components. The computing device 200 may include a user interface 206 including a display device 208 and one or more input devices or mechanisms 210. In some implementations, the input device / mechanism 210 includes a keyboard, and in some implementations, the input device / mechanism includes a “soft” keyboard that optionally appears on the display device 208, allowing a user to “press keys” that appear on the display 208. In some implementations, the display 208 and the input device / mechanism 210 include a touchscreen display (also called a touch-sensitive display). In some implementations, the memory 214 includes high-speed random-access memory such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some implementations, the memory 214 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. In some implementations, the memory 214 includes one or more storage devices located remotely from the CPU 202. The memory 214, or alternatively, the non-volatile memory devices within the memory 214, include a computer-readable storage medium.In some implementations, memory 214 or the computer-readable storage medium of memory 214 stores the following programs, modules and data structures, or a subset thereof: • an operating system 216, which includes procedures for handling various basic system services and for performing hardware-dependent tasks; one or more communication network interfaces 204 (wired or wireless) and a communications module 218 used to connect the computing device 200 to other computers and devices via one or more communications networks, such as the Internet, other wide area networks, local area networks, metropolitan area networks, etc.; A trained neural network 220 including weights 222 and neural network topology 224. Examples of input neural networks are described below with reference to Figures 4A-4C, 12, 13A, and 14A, according to some implementations: a neural network transformation module 226 including a transformed analog neural network 228, a mathematical formulation 230, basis function blocks 232, an analog model 234 (sometimes referred to as a neuron model), and / or analog integrated circuit (IC) design constraints 236. Exemplary operations of the neural network transformation module 226 are described below with reference to the flowcharts shown in at least Figures 5, 6A-6C, 7, 8, 9, 10, and 11A-11C, and Figures 27A-27J and 28A-28S; and / or A weight matrix calculation (sometimes referred to as weight quantization) module 238, including transformed network weights 272, and optionally including a resistance calculation module 240, resistance values ​​242. Exemplary operations of the weight matrix calculation module 238 and / or weight quantization, according to some implementations, are described with reference to at least Figures 17A-17C, 18, and 29A-29F.

[0142]

[0196] Some implementations include one or more optional modules 244, as shown in Figure 2B. Some implementations include an analog neural network optimization module 246. Examples of analog neural network optimization, according to some implementations, are described below with reference to Figures 30A-30M.

[0143]

[0197] Some implementations include a lithography mask generation module 248 that further includes lithography masks 250 for analog components other than resistors (or connections) (e.g., operational amplifiers, multipliers, delay blocks, etc.). In some implementations, the lithography masks are generated based on a chip design layout according to the chip design using a Cadence, Synopsys, or Mentor Graphics software package. Some implementations use a design kit from a silicon wafer fabrication factory (sometimes called a fab). The lithography mask is intended for use in the specific fab that provides the design kit (e.g., the TSMC 65 nm design kit). The generated lithography mask file is used to manufacture the chip in the fab. In some implementations, a Cadence, Mentor Graphics, or Synopsys software package-based chip design is semi-automatically generated from a SPICE or fast SPICE (Mentor Graphics) software package. In some implementations, a user with chip design skills drives the conversion from a SPICE or fast SPICE circuit to a Cadence, Mentor Graphics, or Synopsys chip design. Some implementations combine cadence design blocks in a single neuronal unit and establish appropriate interconnections between the blocks.

[0144]

[0198] Some implementations include a library generation module 254 that further includes a library of lithography masks 256. An example of library generation, according to some implementations, is described below with reference to Figures 32A-32E.

[0145]

[0199] Some implementations include an integrated circuit (IC) fabrication module 258, which further includes an analog-to-digital converter (ADC), digital-to-analog converter (DAC), or other similar interface 260 and / or fabricated ICs or models 262. Exemplary integrated circuits and / or associated modules, according to some implementations, are described below with reference to Figures 31A-31Q.

[0146]

[0200] Some implementations include an energy efficiency optimization module 264, which further includes an inference module 266, a signal monitoring module 268, and / or a power optimization module 270. Examples of energy efficiency optimization, according to some implementations, are described below with reference to Figures 33A-33K.

[0147]

[0201] Each of the above-identified executable modules, applications, or sets of procedures may be stored in one or more of the aforementioned memory devices and correspond to sets of instructions for performing the functions described above. The above-identified modules or programs (i.e., sets of instructions) need not be implemented as separate software programs, procedures, or modules; thus, various subsets of these modules may be combined or otherwise rearranged in various implementations. In some implementations, memory 214 stores a subset of the above-identified modules and data structures. Additionally, in some implementations, memory 214 stores additional modules or data structures not described above.

[0148]

[0202] 2A illustrates a computing device 200, which is intended more as a functional description of various features that may be present than as a structural schematic of the implementations described herein. In practice, items shown separately may be combined and some items may be separated, as will be recognized by those skilled in the art.

[0149] Exemplary Process for Generating a Schematic Model of an Analog Network

[0203] 3A illustrates an exemplary process 300 for generating a schematic model of an analog network corresponding to a trained neural network, according to some implementations. As shown in FIG. 3A, a trained neural network 302 (e.g., a mobile net) is transformed (322) into a target or equivalent analog network 304 (using a process sometimes referred to as a T-transformation). The target neural network (sometimes referred to as a T-network) 304 is exported (324) to SPICE (as a SPICE model 306) using a single neuron model (SNM), which is then exported (326) from SPICE to a full-on-chip design using cadence and a cadence model 308. The cadence model 308 is cross-validated (328) against the initial neural network for one or more validation inputs.

[0150]

[0204] In the above and below descriptions, a mathematical neuron is a mathematical function that receives one or more weighted inputs and produces a scalar output. In some implementations, a mathematical neuron can have memory (e.g., long short-term memory (LSTM), recurrent neurons). A trivial neuron is an "ideal" mathematical neuron,

number

number

number

[0151]

[0205] FIG. 3B illustrates an exemplary manual prototyping process used to generate a target chip model 320 based on an SNM model on Cadence 314, according to some implementations. While the following description uses Cadence, it should be noted that, according to some implementations, alternative tools from Mentor Graphics Design or Synopsys (e.g., Synopsys Design Kit) may be used instead of Cadence tools. This process includes selecting SNM constraints, including inbound and outbound constraints and signal constraints, selecting analog components (e.g., resistors including specific resistor array technologies) for connections between neurons, and developing a Cadence SNM model 314. A prototype SNM model 316 (e.g., a PCB prototype) is developed (330) based on the SNM model on Cadence 314. The prototype SNM model 316 is compared to the SPICE model for equivalence. In some implementations, when the neural network meets the equivalence requirements, the neural network is selected for on-chip prototyping. Because the size of the neural network is small, the T-transform can be manually verified for equivalence. An on-chip SNM model 318 is then generated (332) based on the SNM model prototype 316. The on-chip SNM model is optimized as much as possible according to some implementations. In some implementations, after finalizing the SNM, the on-chip density of the SNM model is calculated before generating (334) a target chip model 320 based on the on-chip SNM model 318. During the prototyping process, a practitioner may iteratively select a neural network task or application and a specific neural network (e.g., having a neural network on the order of 0.1 to 1.1 million neurons), perform the T-transform, establish a cadence neural network model, and design an interface and / or target chip model.

[0152] Example Input Neural Network

[0206] 4A, 4B, and 4C illustrate examples of trained neural networks (e.g., neural network 220) input into system 100 and converted into a mathematically equivalent analog network, according to some implementations. FIG. 4A illustrates an exemplary neural network (sometimes referred to as an artificial neural network) composed of artificial neurons that receive inputs, combine the inputs using activation functions, and generate one or more outputs. Inputs include data such as images, sensor data, and documents. Typically, each neural network performs a specific task, such as object recognition. The network includes connections between neurons, each connection providing the output of one neuron as input to another neuron. After training, each connection is assigned a corresponding weight. As shown in FIG. 4A, neurons are typically organized into multiple layers, with each layer of neurons connected only to the layer immediately preceding and following it. The input layer of neuron 402 receives external inputs (e.g., inputs X1, X2, ..., X n) input layer 402 is followed by one or more hidden layers of neurons (e.g., layers 404 and 406), which are followed by an output layer 408 that generates an output 410. Various types of connection patterns connect the neurons in successive layers, such as a fully connected pattern that connects every neuron in one layer to every neuron in the next layer, or a pruned pattern that connects the output of a group of neurons in one layer to a single neuron in the next layer. In contrast to the neural network shown in FIG. 4A, sometimes referred to as a feedforward network, the neural network shown in FIG. 4B includes one or more connections from neurons in one layer to either other neurons in the same layer or neurons in a preceding layer. The example shown in FIG. 4B is an example of a recurrent neural network and includes two input neurons 412 (accepting input X1) and 414 (accepting input X2) in the input layer, followed by two hidden layers. The first hidden layer includes neurons from the input layer and fully connected neurons 416 and 418, and neurons 420, 422, and 424 in the second hidden layer. The output of second hidden layer neuron 420 is connected to first hidden layer neuron 416, providing a feedback loop. The hidden layer containing neurons 420, 422, and 424 is input to output layer neuron 426, which produces output y.

[0153]

[0207] FIG. 4C illustrates an example of a convolutional neural network (CNN) according to some implementations. In contrast to the neural networks illustrated in FIGS. 4A and 4B, the example illustrated in FIG. 4C includes a different type of neural network layer, including a first stage of layers for feature learning and a second stage of layers for classification tasks such as object recognition. The feature learning stage includes a convolutional and rectified linear unit (ReLU) layer 430 followed by a pruning layer 432, which is followed by another convolutional and ReLU layer 434, which is followed by another pruning layer 436. The first layer 430 extracts features from an input 428 (e.g., an input image or portion thereof) and performs a convolution operation and one or more nonlinear operations (e.g., ReLU, tanh, or sigmoid) on the input. Pruning layers, such as layer 432, reduce the number of parameters when the input is large. The output of pruning layer 436 is smoothed by layer 438 and input to a fully connected neural network having one or more layers (e.g., layers 440 and 442). The output of the fully connected neural network is input to a softmax layer 444, which classifies the output of layer 442 of the fully connected network to produce one of many different outputs 446 (e.g., the object class or type of the input image 428).

[0154]

[0208] Some implementations store the layout or organization of the input neural network, including the number of neurons in each layer, the total number of neurons, the operation or activation function of each neuron, and / or the connections between neurons, in memory 214 as neural network topology 224.

[0155]

[0209] 5 shows an example of a mathematical model 500 for a neuron, according to some implementations. The mathematical model involves multiplying an input incoming signal 502 by a synaptic weight 504 and summing it with a summation unit 506. The result of the summation unit 506 is input to a nonlinear transformation unit 508 to generate an output signal 510, according to some implementations.

[0156]

[0210] 6A-6C illustrate an exemplary process for analog hardware implementation of a neural network for computing the XOR of input values ​​(classification of the XOR result) according to some implementations. FIG. 6A shows a table 600 of possible input values ​​X1 and X2 along the x-axis and y-axis, respectively. Expected result values ​​are indicated by hollow circles (representing a value of 1) and filled or dark circles (representing a value of 0). This is a typical XOR problem for two input signals and two classes. The expected result is 1 only if either value X1 or X2 is 1, but not both; otherwise, it is 0. The training set consists of four possible input signal combinations (binary values ​​of X1 and X2 inputs). FIG. 6B shows a ReLU-based neural network 602 for solving the XOR classification of FIG. 6A according to some implementations. The neurons do not use bias values ​​but use ReLU activation. Inputs 604 and 606 (corresponding to X1 and X2, respectively) are input to a first ReLU neuron 608-2. Inputs 604 and 606 are also input to a second ReLU neuron 608-4. The result of the two ReLU neurons 608-2 and 608-4 is input to a third neuron 608-6, which performs a linear summation of the input values ​​to generate an output value 510 (Out value). The neural network 602 has weights −1 and 1 (for input values ​​X1 and X2, respectively) for the ReLU neuron 608-2, weights −1 and 1 (for input values ​​X1 and X2, respectively) for the ReLU neuron 608-4, and weights 1 and 1 (for the outputs of the ReLU neurons 608-2 and 608-4, respectively). In some implementations, the weights of the trained neural network are stored in memory 214 as weights 222.

[0157]

[0211] 6C shows an exemplary analog network equivalent of network 602, according to some implementations. Analog equivalents 614 and 616 of X1 and X2 inputs 604 and 606 are fed to first-layer analog neurons N1 618 and N2 620. Neurons N1 and N2 are tightly coupled to second-layer neurons N3 and N4. The second-layer neurons (i.e., neuron N3 622 and neuron N4 624) are connected to output neuron N5 626, which generates output Out (equivalent to output 610 of network 602). Neurons N1, N2, N3, N4, and N5 have ReLU (max=1) activation functions.

[0158]

[0212] Some implementations use Keras learning, which converges in approximately 1000 iterations and yields the connection weights. In some implementations, the weights are stored in memory 214 as part of weights 222. In the example below, the data format is "neuron[first link weight, second link weight, bias]". ●N1[-0.9824321,0.976517,-0.00204677], ●N2[1.0066702,-1.0101418,-0.00045485], ●N3[1.0357606,1.0072469,-0.00483723], ●N4[-0.07376373,-0.7682612,0.0], and ●N5[1.0029935,-1.1994369,-0.00147767].

[0159]

[0213] Next, to calculate resistor values ​​for connections between neurons, some implementations calculate a resistor range. Some implementations set a resistor nominal value (R+, R−) of 1 MΩ, a possible resistor range of 100 KΩ to 1 MΩ, and a nominal series E24. Some implementations calculate the w1, w2, and wbias resistor values ​​for each connection as follows: For each weight value wi (e.g., weight 222), some implementations evaluate all possible (Ri−, Ri+) resistor pair options within the chosen nominal series and select the one with the smallest error value

number

[0160] [Table 1]

[0161] Exemplary Advantages of Transformed Neural Networks

[0214] Before describing examples of transformation, some advantages of transformed neural networks over traditional architectures should be noted. As described herein, an input trained neural network is transformed into an analog network of pyramidal or trapezoidal shape. Some advantages of pyramidal or trapezoidal structures over crossbar structures include lower latency, simultaneous analog signal propagation, manufacturability using standard integrated circuit (IC) design elements including resistors and operational amplifiers, high computational parallelism, high accuracy (e.g., accuracy increases with the number of layers compared to traditional methods), tolerance to errors in each weight and / or each connection (e.g., pyramids balance errors), low RC (low resistance-capacitance delay associated with signal propagation through the network), and / or the ability to manipulate the bias and function of each neuron in each layer of the transformed network. Pyramids are also excellent computational blocks in their own right, being multilevel perceptrons capable of modeling any neural network with one output. According to some implementations, networks with several outputs are implemented using different pyramidal or trapezoidal geometries. A pyramid can be thought of as a multilayer perceptron with one output and several layers (e.g., N layers), where each neuron has n inputs and one output. Similarly, a trapezoid is a multilayer perceptron with each neuron having n inputs and m outputs. Each trapezoid is, in some implementations, a pyramid-like network with each neuron having n inputs and m outputs, where n and m are limited by IC analog chip design limitations.

[0162]

[0215] Some implementations perform a reversible transformation of any trained neural network into a pyramidal or trapezoidal subsystem. Therefore, pyramidal and trapezoidal networks can be used as universal building blocks for transforming any neural network. An advantage of pyramidal- or trapezoid-based neural networks is the possibility of implementing any neural network using standard IC analog elements (e.g., operational amplifiers, resistors, and signal delay lines in the case of recurrent neurons) using standard lithography techniques. It is also possible to restrict the weights of the transformed network to a certain interval. That is, in some implementations, the reversible transformation is performed with weights restricted to some predefined range. Another advantage of using pyramidal or trapezoidal networks is the simultaneous propagation of analog signals, which increases the parallelism of signal processing or computation speed and provides lower latency. Furthermore, many modern neural networks are loosely coupled networks that perform much better when transformed into pyramidal networks than when transformed into crossbar networks (e.g., they are more compact, have lower RC values, and have no leakage current), and pyramidal and trapezoidal networks are relatively more compact than crossbar-based memristor networks.

[0163]

[0216] Furthermore, analog neuromorphic trapezoidal chips possess several properties not typical of analog devices. For example, the signal-to-noise ratio does not increase as the number of cascaded analog chips increases, external noise is suppressed, and temperature effects are significantly reduced. These properties make trapezoidal analog neuromorphic chips similar to digital circuits. For example, in some implementations, individual neurons use operational amplifiers to adjust the signal level, operating at frequencies between 20,000 and 100,000 Hz and unaffected by noise or signals with frequencies higher than the operating range. Trapezoidal analog neuromorphic chips also perform output signal filtering due to the unique nature of how operational amplifiers function. Such trapezoidal analog neuromorphic chips suppress close-in phase noise. Furthermore, this noise is significantly reduced due to the low-ohm outputs of the operational amplifiers. Due to the signal level adjustment at each operational amplifier output and the amplifier's synchronous action, temperature-induced parameter drift does not affect the signal at the final output. The trapezoid-like neuromorphic circuit is tolerant to errors and noise in the input signal and to deviations in the resistor values ​​corresponding to the weight values ​​of the neural network.Due to the very nature of the analog neuromorphic trapezoid-like circuit based on operational amplifiers, the trapezoid-like analog neuromorphic network is also tolerant to any kind of systematic error, such as an error in resistor value setting, if such error is the same for all resistors.

[0164] Example reversible transform (T-transform) of a trained neural network

[0217] In some implementations, the example transformations described herein are performed by a neural network transformation module 226 that transforms the trained neural network 220 based on the mathematical formulation 230, the basis function blocks 232, the analog component models 234, and / or the analog design constraints 236 to obtain a transformed neural network 228.

[0165]

[0218] Figure 7 shows an example perceptron 700 according to some implementations. The perceptron includes K = 8 inputs and eight neurons 702-2, ..., 702-16 in the input layer that receive the eight inputs. The output layer has four neurons 704-2, ..., 704-8 that correspond to L = 4 outputs. The input layer neurons are fully connected to the output layer neurons, making 8 x 4 = 32 connections. Let the connection weights be represented by a weight matrix WP (elements WP i,j corresponds to the weight of the connection between the ith neuron in the input layer and the jth neuron in the output layer). Furthermore, let each neuron implement an activation function F.

[0166]

[0219] Figure 8 illustrates an exemplary pyramidal neural network (P-NN) 800, a type of target neural network (T-NN or TNN) equivalent to the perceptron shown in Figure 7, according to some implementations. To perform this conversion of the perceptron (Figure 7) to the PN-NN architecture (Figure 8), assume that the number of inputs for the T-NN is limited to Ni = 4 and the number of outputs is limited to No = 2. The T-NN includes an input layer LTI of neurons 802-2, ..., 802-34, which is a concatenation of two copies of the input layer of neurons 802-2, ..., 802-16, for a total of 2 x 8 = 16 input neurons. The set of neurons 804, including neurons 802-20, ..., 802-34, is a copy of neurons 802-2, ..., 802-18, and the inputs are replicated. For example, the input to neuron 802-2 is also input to neuron 802-20, and input 20 to neuron 802-4 is also input to neuron 802-22, and so on. Figure 8 also includes a hidden layer LTH1 of linear neurons 806-02, ..., 806-16 (2 × 16 ÷ 4 = 8 neurons). Each group of Ni neurons from the input layer LTI is fully connected to two neurons from the LTH1 layer. Figure 8 also includes an output layer LTO with 2 × 8 ÷ 4 = 4 neurons 808-02, ..., 808-08, each implementing an activation function F. Each neuron in layer LTO is connected to a different neuron from a different group in layer LTH1. The network shown in Figure 8 includes 40 connections. Some implementations perform the weight matrix calculation for the P-NN of Figure 8 as follows: The weights of the hidden layer LTH1 (WTH1) are calculated from the weight matrix WP, and the weights corresponding to the output layer LTO (WTO) form a sparse matrix with elements equal to 1.

[0167]

[0220] Figure 9 illustrates a Pyramid Single Neural Network (PSNN) 900 corresponding to the output neurons of Figure 8 according to some implementations. The PSNN includes a layer (LPSI) of input neurons 902-02, ..., 902-16 (corresponding to the eight input neurons of network 700 in Figure 7). Hidden layer LPSH1 includes 8 / 4 = two linear neurons 904-02 and 904-04, with each group of Ni neurons from LTI connected to one neuron in the LPSH1 layer. Output layer LPSO consists of one neuron 906 with activation function F connected to both hidden layer neurons 904-02 and 904-04. To calculate the weight matrix of PSNN 900, some implementations calculate a vector WPSH1 equal to the first row of WP for the LPSH1 layer. For the LPSO layer, some implementations calculate a weight vector WPSO with two elements, each equal to 1. This process is repeated for the first, second, third, and fourth output neurons. A P-NN, such as the network shown in Figure 8, is an amalgamation of PSNNs (of four output neurons). The input layer of every PSNN is a separate copy of the input layer of P. In this example, the P-NN 800 includes an input layer with 8 x 4 = 32 inputs, a hidden layer with 2 x 4 = 8 neurons, and an output layer with 4 neurons.

[0168] An exemplary transformation using a target neuron with N inputs and one output

[0221] In some implementations, the example transformations described herein are performed by a neural network transformation module 226 that transforms the trained neural network 220 based on the mathematical formulation 230, the basis function blocks 232, the analog component models 234, and / or the analog design constraints 236 to obtain a transformed neural network 228.

[0169] Single-layer perceptron with one output

[0222] Let a single-layer perceptron SLP(K,1) contain K inputs and one output neuron with activation function F. Furthermore, U∈R K Let be the vector of weights for SLP(K,1). The following algorithm Neuron2TNN1 builds a T neural network from T neurons (called TN(N,1)) with N inputs and one output. Algorithm Neuron2TNN1 1. Build the input layer of the T-NN by including all inputs from SLP(K,1). 2. If K>N: aK input neurons, so that every group has no more than N inputs.

number

number

number

number

number

number

[0170]

[0223] where:

number

number

number

[0171]

[0224] Figure 10 shows an example of a constructed T-NN according to some implementations. All layers except the first perform an identity transformation of the inputs of these layers. The weight matrix of the constructed T-NN has the following form according to some implementations: Layer 1 (e.g., layer 1002):

number

number

[0172]

[0225] The output value of the T-NN is calculated using the following formula: y=F(W m W m-1 ...W 2 W 1 x)

[0173]

[0226] The output of the first layer is calculated as an output vector according to the following formula:

number

[0174]

[0227] The obtained vector is multiplied by the weight matrix of the second layer.

number

[0175]

[0228] Every subsequent layer outputs a vector whose components are equal to a linear combination of some subvector of x.

[0176]

[0229] Finally, the output of the T-NN is equal to:

number

[0177]

[0230] This is the same value calculated by SLP(K,1) for the same input vector x. Therefore, the output value of SLP(K,1) and the output value of the constructed T-NN are equal.

[0178] Single-layer perceptron with several outputs

[0231] Suppose we have a single-layer perceptron SLP(K,L) with K input and L output neurons, each performing an activation function F. Furthermore, U∈R L×K Let be the weight matrix of SLP(K,L). The following algorithm Layer2TNN1 constructs a T neural network from neurons TN(N,1). Algorithm Layer2TNN1 1. For every output neuron i=1,...,L aK inputs, 1 output neuron and a weight vector U ij , j=1,2,...,K, we use the algorithm Neuron2TNN1 as the SLP i (K,1). As a result, TNN i is constructed. 2. All TNNs i We construct a PTNN by composing the following into a single neural network: a.All TNNs i To concatenate the input vectors of SLP(K,L), the input of the PTNN has L groups of K inputs, each group being a copy of the input layer of SLP(K,L).

[0179]

[0232] The output of the PTNN is a pair of SLPs. i (K,1) and TNN i Since the outputs of are equal, they are equal to the output of SLP(K,L) for the same input vector.

[0180] Multilayer Perceptron

[0233] A multilayer perceptron (MLP) has K inputs, S layers, and L i It contains computational neurons in the i-th layer, and the MLP (K, S, L1, ... L S ) shall be expressed as

number

[0181]

[0234] Below is an exemplary algorithm for constructing T neural networks from neurons TN(N,1) according to some implementations. Algorithm MLP2TNN1 1. For every layer i=1,...,S aL i-1 inputs, L i output neurons and weight matrix U i SLP consisting of i (L i-1 ,L i ) and apply algorithm Layer2TNN1 to it, resulting in PTNN i Build. 2. All PTNNs i We build MTNN by stacking these into one neural network, and TNN i-1 The output of the TNN i is set as the input for

[0182]

[0235] Any pair SLP i (L i-1 ,L i ) and PTNN i Since the outputs of the MTNN are equal, the output of the MLP(K,S,L1,...LS ) output.

[0183] N I inputs and N O An example T-transformation for a target neuron with outputs

[0236] In some implementations, the example transformations described herein are performed by a neural network transformation module 226 that transforms the trained neural network 220 based on the mathematical formulation 230, the basis function blocks 232, the analog component models 234, and / or the analog design constraints 236 to obtain a transformed neural network 228.

[0184] Example transformation of a single-layer perceptron with several outputs

[0237] Let a single-layer perceptron SLP(K,L) contain K input and L output neurons, each of which implements an activation function F. Furthermore, U∈R L×K Let be the weight matrix of SLP(K,L). The following algorithm calculates the weights of neurons TN(N I ,N O ) to construct a T neural network. Algorithm Layer2TNNX 1. Construct a PTNN from an SLP(K,L) using algorithm Layer2TNN1 (see above). The PTNN has an input layer consisting of L groups of K inputs. 2. From L groups

number

[0185]

[0238] According to some implementations, the output of PTNNX is calculated by the same formula as for PTNN (described above), so the outputs are equal.

[0186]

[0239] 11A-11c illustrate a network of two output neurons and a TN (N I 11A illustrates an application 1100 of the above algorithm to a single-layer neural network (NN) having a layer 1104 (neurons 1 and 2). FIG. 11A shows an exemplary source or input NN according to some implementations. K inputs are input to two neurons 1 and 2 belonging to layer 1104. FIG. 11B illustrates a PTNN constructed after the first step of the algorithm according to some implementations. The PTNN consists of two parts that implement subnets corresponding to output neurons 1 and 2 of the NN shown in FIG. 11A. In FIG. 11B, input 1102 is replicated and input to two sets of input neurons 1106-2 and 1106-4. Each set of input neurons is connected to neurons in a subsequent layer having two sets of neurons 1108-2 and 1108-4, each set containing m neurons. The input layer is followed by identity transformation blocks 1110-2 and 1110-4, each containing one or more layers with an identity weight matrix. The output of identity transformation block 1110-2 is connected to output neuron 1112 (corresponding to output neuron 1 in FIG. 11A), and the output of identity transformation block 1110-4 is connected to output neuron 1114 (corresponding to output neuron 1 in FIG. 11A). Figure 11C shows the application of the final step of the algorithm, which involves replacing the two copies of the input vectors (1106-2 and 1106-4) with one vector 1116 (step 3) and re-establishing connectivity within the first layer 1118 by creating two output links from every input neuron, one link connecting to the subnet associated with output 1 and another link connecting to the subnet for output 2.

[0187] Example transformations of multilayer perceptrons

[0240] A multilayer perceptron (MLP) has K inputs, S layers, and L i It contains computational neurons in the i-th layer, and the MLP (K, S, L1,...L S ) shall be expressed as

number

[0188]

[0241] In some implementations, every pair of SLPs i (L i-1 ,L i ) and PTNNX i Since the outputs of MTNNX are equal, the output of MTNNX is the same as that of MLP(K,S,L1,...L S ) output.

[0189] Example transformations of recurrent neural networks

[0242] Recurrent neural networks (RNNs) contain backward connections that allow them to store information. Figure 12 shows an example RNN 1200, according to some implementations. The example shows an input X t 1206, executes activation function A, and returns the value h t 1202. The back arrow from block 1204 to itself indicates a backward connection, according to some implementations. An equivalent network would be one in which the activation block has input X t At time 0, the network receives input X t 1208 and executes activation function A 1204, generating the value h o At time 1, the network accepts input X1 1212 and the network's output at time 0, executes activation function A 1204, and outputs value h1 1210; at time 2, the network accepts input X2 1216 and the network's output at time 1, executes activation function A 1204, and outputs value h1 1218. This process continues, according to some implementations, until time t, at which point the network accepts input X2 1216 and the network's output at time 1, executes activation function A 1204, and outputs value h1 1218. t 1206 and the output of the network at time t-1, and executes activation function A 1204 to obtain the value h t Outputs 1202.

[0190]

[0243] Data processing in the RNN is performed according to the following equation: h t =f(W (hh) h t-1 +W (hx) x t )

[0191]

[0244] In the above equation, x t is the current input vector, and h t-1 is the previous input vector x t-1 This expression is the output of the RNN for several operations, namely, two fully connected layers W (hh) h t-1 and W(hx )x t It consists of computing a linear combination of (f), element-wise addition, and computing a nonlinear function (f). The first and third operations can be implemented by a trapezoid-based network (one fully connected layer is implemented by a pyramid-based network, which is a special case of a trapezoid network). The second operation is a general operation that can be implemented in networks of any structure.

[0192]

[0245] In some implementations, layers of an RNN that do not have recurrent connections are transformed using the Layer2TNNX algorithm described above. After the transformation is complete, recurrent links are added between the associated neurons. Some implementations use delay blocks, as described below with reference to FIG. 13B.

[0193] Example transformations of LSTM networks

[0246] A long short-term memory (LSTM) neural network is a special case of an RNN. The operation of an LSTM network is expressed by the following equation: ●f t =σ(W f [h t-1 ,x t ]+b f ), ●i t =σ(W i [h t-1 ,x t ]+b i ), ●D t =tanh(W D [h t-1 ,x t ]+b D ), ●C t =(f t ×C t-1 +i t ×D t ), ●o t =σ(W o [h t-1 ,x t ]+b o ), and ●h t =ot ×tanh(C t ).

[0194]

[0247] In the above formula, W f , W i , W D and W O is the trainable weight matrix, and b f , b i , b D and b O is the trainable bias, and x t is the current input vector, and h t-1 is the previous input vector x t-1 is the internal state of the LSTM calculated for o t is output for the current input vector, where the subscript t refers to time instance t and the subscript t-1 refers to time instance t-1.

[0195]

[0248] 13A is a block diagram of an LSTM neuron 1300, according to some implementations. A sigmoid (σ) block 1318 receives an input h t-1 1330 and x t Process 1332 and output f t The second sigmoid (σ) block 1320 generates the input h t-1 1330 and x t Process 1332 and output i t The hyperbolic tangent (tanh) block 1322 generates 1338. t-1 1330 and x t Process 1332 and output D t 1340. The third sigmoid (σ) block 1328 generates the input h t-1 1330 and x t Process 1332 and output O t 1342. The multiplier block 1304 generates f t 1336 and the output C of the summation block 1306 (from the preceding time instance) t-1 1302 to generate an output, which is output i t 1338 and D tMultiply by 1340 and output C t The output of the second multiplier block 1314 is summed by summing block 1306 to produce 1310. t 1310 is input to another tanh block 1312 to produce an output which is multiplied by a third multiplier block 1316 to produce an output O t multiplied by 1342 to produce the output h t Generates 1334.

[0196]

[0249] These expressions utilize several types of operations: (i) computation of linear combinations of several fully connected layers, (ii) element-wise addition, (iii) Hadamard product, and (iv) nonlinear function computation (e.g., sigmoid (σ) and hyperbolic tangent (tanh)). Some implementations implement operations (i) and (iv) by trapezoid-based networks (one fully connected layer is implemented by a pyramid-based network, which is a special case of a trapezoid network). Some implementations use networks of various structures for the general operations (ii) and (iii).

[0197]

[0250] In some implementations, layers of an LSTM layer that do not have recurrent connections are transformed using the Layer2TNNX algorithm described above. After the transformation is complete, in some implementations, recurrent links are added between the associated neurons.

[0198]

[0251] 13B illustrates a delay block according to some implementations. As mentioned above, some of the expressions in the LSTM operations depend on saving, restoring, and / or recalling outputs from previous time instances. For example, multiplier block 1304 multiplies the output of sum block 1306 (from the previous time instance) C t-1 13B shows two examples of delay blocks, according to some implementations. Example 1350 includes a left delay block 1354 that processes input x at time t. t Accepts 1352 and outputs x t-dtThe example on the right, 1360, shows that, according to some implementations, cascaded delay blocks 1364 and 1366 generate an output x t-2dt After a time delay of 2 units indicated by 1368, the input x t Outputs 1362.

[0199]

[0252] 13C is a neuron schema of an LSTM neuron, according to some implementations. The schema includes weighted summation nodes (sometimes called adder blocks) 1372, 1374, 1376, 1378, and 1396, multiplier blocks 1384, 1392, and 1394, and delay blocks 1380 and 1382. The input x t 1332 is connected to adder blocks 1372, 1374, 1376 and 1378. t-1 Output of h t-1 1330 is also input to adder blocks 1372, 1374, 1376 and 1378. Adder block 1372 produces an output f t 1336. Similarly, adder block 1374 generates an output i t 1338. Similarly, adder block 1376 generates an output D t 1340. Similarly, adder block 1378 generates an output O t 1342. The multiplier block 1392 generates an output that is input to a sigmoid block 1390 that generates an output i t 1338, f t 1336 and preceding time instance C t-1 The output of adder block 1396 from 1302 is used to generate the first output. Multiplier block 1394 multiplies the output i t 1338 and D t 1340 to generate a second output. An adder block 1396 sums the first and second outputs to generate an output C t Generates 1310. Output Ct 1310 is input to a hyperbolic tangent block 1398 to produce an output, which is the output of the sigmoid block 1390, O t 1342 are input to a multiplier block 1384 to produce an output h t 1334. Delay block 1382 is used to recall (e.g., save and restore) the output of adder block 1396 from a prior time instance. Similarly, delay block 1380 retrieves (e.g., saves and restores) the prior input x (e.g., from a prior time instance). t-1 13B. An example of a delay block is described above with reference to FIG. 13B, according to some implementations.

[0200] Example transformation of GRU network

[0253] A Gated Recurrent Unit (GRU) neural network is a special case of an RNN. The operation of an RNN can be expressed in the following formula: ●z t =σ(W z x t +U z h t-1 ), ●r t =σ(W r x t +U r h t-1 ), ●j t =tanh(Wx t +r t ×Uh t-1 ), ●h t =z t ×h t-1 +(1-z t )×j t ).

[0201]

[0254] In the above formula, x t is the current input vector, and h t-1 is the previous input vector x t-1 is the output calculated for

[0202]

[0255] 14A is a block diagram of a GRU neuron, according to some implementations. The sigmoid (σ) block 1418 receives the input h t-1 1402 and x t Process 1422 and output r t The second sigmoid (σ) block 1420 generates the input h t-1 1402 and x t Process 1422 and output z t 1428. The multiplier block 1412 produces the output r t Enter 1426 t-1 1402 to produce an output, which is the sum of the input x t 1422) is input to a hyperbolic tangent (tanh) block 1424, which outputs j t The second multiplier block 1414 produces output j t 1430 and output z t 1428 to generate the first output. t 1428 to generate an output, which is the sum of the output and the input h t-1 1402 to produce an output which is input to a third multiplier block 1404 which multiplies the first output (from multiplier block 1414) into an adder block 1406 to produce an output h t Generate 1408. Input h t-1 1402 is the output of the GRU neuron from the previous time interval output t-1.

[0203]

[0256] 14B is a neuron schema of a GRU neuron 1440, according to some implementations. The schema includes weighted summation nodes (sometimes called adder blocks) 1404, 1406, 1410, 1406, and 1434, multiplier blocks 1404, 1412, and 1414, and a delay block 1432. The input x t 1422 is connected to adder blocks 1404, 1410 and 1406. t-1 Output of h t-11402 is also input to adder blocks 1404 and 1406 and multiplier blocks 1404 and 1412. Adder block 1404 produces an output Z t 1428. Similarly, adder block 1406 generates an output that is input to sigmoid block 1420, which in turn generates an output r that is input to multiplier block 1412. t 1426. The output of multiplier block 1412 is input to adder block 1410, whose output is input to hyperbolic tangent block 1424, which generates output 1430. Output 1430 as well as the output of sigmoid block 1418 are input to multiplier block 1414. The output of sigmoid block 1418 is input to multiplier block 1404, which multiplies its output with an input from delay block 1432 to generate a first output. The multiplier block generates a second output. Adder block 1434 sums the first and second outputs to generate output h t 1408. A delay block 1432 is used to recall (e.g., save and restore) the output of adder block 1434 from a previous time instance. Examples of delay blocks are described above with reference to FIG. 13B, according to some implementations.

[0204]

[0257] Since the operation types used in the GRU are the same as those in the LSTM network (described above), the GRU is converted into a trapezoid-based network according to the principles described above for LSTM (e.g., using the Layer2TNNX algorithm) in some implementations.

[0205] Example transformations of convolutional neural networks

[0258] In general, a convolutional neural network (CNN) includes several basic operations such as convolution (a set of linear combinations of pieces of an image (or internal map) with a kernel), activation functions, and pruning (e.g., maximum, average, etc.). Every computational neuron in a CNN follows the general processing scheme of neurons in an MLP, i.e., a linear combination of several inputs with subsequent computation of an activation function. Therefore, a CNN is transformed using the MLP2TNNX algorithm described above for a multilayer perceptron, according to some implementations.

[0206]

[0259] Conv1D is a convolution performed over the time coordinate. Figures 15A and 15B are neuron schemas of variants of a single Conv1D filter, according to some implementations. In Figure 15A, the weighted sum node 1502 (sometimes called an adder block and labeled "+") has five inputs, and therefore corresponds to a 1D convolution with a kernel of 5. The inputs are x from time t t 1504, x from time t-1 t-1 1514 (obtained by inputting the input into delay block 1506), x from time t-2 t-2 1516 (obtained by inputting the output of delay block 1506 into another delay block 1508), x from time t-3 t-3 1518 (obtained by inputting the output of delay block 1508 into another delay block 1510) and x from time t-4 t-4 1520 (obtained by inputting the output of delay block 1510 into another delay block 1512). For large kernels, it may be beneficial to utilize delay blocks of different frequencies, so that some of the blocks produce larger delays. Some implementations replace several small delay blocks with one large delay block, as shown in FIG. 15B. In addition to the delay blocks of FIG. 15A, an example would be to add x from time t-3 t-3 1518 and the delay_3 block 1524 from time t-5 t-5Another delay block 1526 is used to generate delay_3 1522. Delay_3 1524 block is an example of multiple delay blocks, according to some implementations. While this operation does not reduce the total number of blocks, according to some implementations, it may reduce the total number of resulting operations performed on the input signal, reducing error accumulation.

[0207]

[0260] In some implementations, the convolutional layers are represented by trapezoidal neurons and the fully connected layers are represented by crossbars of resistors. Some implementations use the crossbars to calculate the resistance matrix of the crossbars.

[0208] Exemplary Approximation Algorithm for Single-Layer Perceptrons with Multiple Outputs

[0261] In some implementations, the example transformations described herein are performed by a neural network transformation module 226 that transforms the trained neural network 220 and / or the analog neural network optimization module 246 based on the mathematical formulation 230, the basis function blocks 232, the analog component models 234, and / or the analog design constraints 236 to obtain a transformed neural network 228.

[0209]

[0262] Let a single-layer perceptron SLP(K,L) contain K input and L output neurons, each of which implements an activation function F. Furthermore, U∈R L×K Let be the weight matrix of SLP(K,L). The following is an example of approximating the neuron TN(N I ,N O) is an example for constructing a T neural network from a matrix of 1000 neurons. The algorithm applies the Layer2TNN1 algorithm (described above) in the first stage to reduce the number of neurons and connections, followed by Layer2TNNX to process the reduced-size input. The output of the resulting neural network is calculated using the shared weights of the layers constructed by the Layer2TNN1 algorithm. The number of these layers is determined by the value p, a parameter of the algorithm. If p is equal to 0, only the Layer2TNNX algorithm is applied, and the transformation is equivalent. If p > 0, p layers share weights, and the transformation is approximate. AlgorithmLayer2TNNX_Approx 1. Parameter p is

number

number

number

number

number

number

[0210]

[0263] 16 shows an example architecture 1600 of the resulting neural net, according to some implementations. This example includes a PNN 1602 connected to a TNN 1606. The PNN 1602 includes a layer of K inputs, with N connected as inputs 1612 to the TNN 1606. p The TNN 1606 generates L outputs 1610 according to some implementations.

[0211] Approximation algorithms for multilayer perceptrons with several outputs.

[0264] A multilayer perceptron (MLP) has K inputs, S layers, and L iIt contains computational neurons in the i-th layer, and the MLP (K, S, L1, ... L S ) and further,

number

[0212] Exemplary Methods for Compressing Transformed Neural Networks

[0265] In some implementations, the example transformations described herein are performed by a neural network transformation module 226 that transforms the trained neural network 220 and / or the analog neural network optimization module 246 based on the mathematical formulation 230, the basis function blocks 232, the analog component models 234, and / or the analog design constraints 236 to obtain a transformed neural network 228.

[0213]

[0266] This section describes exemplary methods for compressing transformed neural networks according to some implementations. Some implementations compress analog pyramidal neural networks to minimize the number of operational amplifiers and resistors required to realize an analog network-on-chip. In some implementations, the method for compressing analog neural networks is pruning, similar to pruning in software neural networks. Nevertheless, compressing pyramidal analog networks, which can be realized in hardware as integrated analog chips, has some peculiarities. Because the number of elements, such as operational amplifiers and resistors, defines the weights of an analog-based neural network, it is crucial to minimize the number of operational amplifiers and resistors placed on-chip. This also helps minimize the chip's power consumption. Modern neural networks, such as convolutional neural networks, can be compressed 5 to 200 times without significant loss of network accuracy. In many cases, entire blocks of modern neural networks can be pruned without significant loss of accuracy. Converting a dense neural network into a loosely connected pyramidal, trapezoidal, or crossbar-like neural network presents an opportunity to prune the loosely connected pyramidal or trapezoidal analog network, which are then represented by operating amplifiers and resistors in an analog IC chip. In some implementations, such techniques are applied in addition to traditional neural network compression techniques. In some implementations, compression techniques are applied based on the specific architecture (e.g., pyramidal vs. trapezoidal vs. crossbar) of the input neural network and / or the converted neural network.

[0214]

[0267] For example, because the network is realized with analog elements such as operational amplifiers, some implementations determine the current flowing through the operational amplifier when presented with a standard training data set, thereby determining whether knots (operational amplifiers) are needed for the entire chip. Some implementations analyze a SPICE model of the chip and determine knots and connections that do not draw current and consume power. Some implementations determine the current flow through the analog IC network and therefore determine which knots and connections to prune. Furthermore, some implementations also remove connections if the connection weight is too high and / or replace resistors with direct connectors if the connection weight is too low. Some implementations prune knots if all connections leading to this knot have weights below a predetermined threshold (e.g., close to 0), delete connections where the operational amplifier always provides zero at the output if the amplifier provides a linear function without amplification, and / or change the operational amplifier to a linear junction.

[0215]

[0268] Some implementations apply compression techniques specific to pyramidal, trapezoidal, or crossbar-type neural networks. Some implementations generate pyramidal or trapezoidal networks with a larger input volume (than without compression), thus minimizing the number of layers in the pyramid or trapezoid. Some implementations generate more compact trapezoidal networks by maximizing the number of outputs for each neuron.

[0216] Exemplary Generation of Optimal Resistor Sets

[0269] In some implementations, the example calculations described herein are performed by a weight matrix calculation or weight quantization module 238 (e.g., using a resistance calculation module 240) that calculates the weights 272 and / or the corresponding resistance values ​​242 of the weights 272 for the connections of the transformed neural network.

[0217]

[0270] This section describes examples of generating optimal resistor sets for a trained neural network, according to some implementations. Exemplary methods are provided for converting connection weights to resistor nominal values ​​for implementing a neural network on a microchip with potentially smaller resistor nominal values ​​and potentially higher allowed resistor variations (sometimes referred to as an NN model).

[0218]

[0271] The test set "Test" contains approximately 10,000 values ​​of input vectors (x and y coordinates), with both coordinates varying in the range [0;1] in steps of 0.01. The network NN output for a given input X is given by Out = NN(X). Furthermore, the input value class is determined as follows: Class_nn(X) = NN(X) > 0.61?1:0.

[0219]

[0272] Below we compare the mathematical network model M with the rough network model S. The rough network model includes possible resistor variations in rv and processes a "test" set, each time producing a different vector of output values ​​S(test) = Out_s. The output error is defined by the following equation:

number

[0220]

[0273] The classification error is defined by the following formula:

number

[0221]

[0274] Some implementations set the desired classification error to 1% or less.

[0222] Exemplary Error Analysis

[0275] 17A shows an example chart 1700 illustrating the dependency between output error and classification error on an M-network, according to some implementations. In FIG. 17A, the x-axis corresponds to classification margin 1704, and the y-axis corresponds to total error 1702 (see above). The graph shows the total error (the difference between the output of model M and the actual data) for different classification margins of the output signal. For this example, the chart shows that the optimal classification margin 1706 is 0.610.

[0223]

[0276] If another network, O, produces output values ​​with a constant shift relative to the associated M output value, there will be a classification error between O and M. To keep the classification error below 1%, this shift should be in the range [-0.045, 0.040]. Thus, the possible output error of S is 45mV.

[0224]

[0277] Possible weight errors are determined by analyzing the dependency between weight / bias relative error and output error across the entire network. Charts 1710 and 1720, shown in Figures 17B and 17C, respectively, are obtained by averaging 20 randomly modified networks across a "test" set, according to some implementations. In these charts, the x-axis represents absolute weight error 1712, and the y-axis represents absolute output error 1714. As can be seen from the charts, an output error limit of 45 mV (y = 0.045) allows for a relative error value of 0.01 or an absolute error value (x value) of 0.01 for each weight. The maximum weight modulus (the maximum absolute value of a weight among all weights) of the neural network is 1.94.

[0225] Exemplary Process for Selecting a Resistor Set

[0278] Resistors configured with {R+,R-} pairs selected from this set have value functions that exceed the required weight range [-wlim;wlim] with some resistor error r_err. In some implementations, the value function of a resistor set is calculated as follows: • An array of possible weight options is calculated together with a weighted average error that depends on the resistor error. Weight options in the array are limited to the required weight range [-wlim;wlim], Values ​​that are worse than their neighbors in terms of weight error are removed, An array of distances between adjacent values ​​is calculated, The value function is the mean square or maximum component of the distance array.

[0226]

[0279] Some implementations iteratively search for an optimal resistor set by continuously adjusting each resistor value in the resistor set based on a learning rate value. In some implementations, the learning rate varies over time. In some implementations, the initial resistor set is selected as uniform (e.g., [1;1;...;1]), and the minimum and maximum resistor values ​​are selected to be within a two-order of magnitude range (e.g., [1;100] or [0.1;10]). Some implementations choose R+ = R−. In some implementations, the iterative process converges to a local minimum. In one case, the process yielded the following set: [0.17, 1.036, 0.238, 0.21, 0.362, 1.473, 0.858, 0.69, 5.138, 1.215, 2.083, 0.275]. This is a locally optimal resistor set of 12 resistors for a weight range [-2;2] with rmin=0.1 (minimum resistance), rmax=10 (maximum resistance), and r_err=0.001 (resistance estimation error). Some implementations do not use the entire available range [rmin;rmax] to find a good local optimum. Only a portion of the available range (e.g., in this case, [0.17;5.13]) is used. The resistor settings are relative, not absolute. In this case, a relative value range of 30 is sufficient for the resistor set.

[0227]

[0280] In one example, the following resistor set of length 20 is obtained for the above parameters: [0.300, 0.461, 0.519, 0.566, 0.648, 0.655, 0.689, 0.996, 1.006, 1.048, 1.186, 1.222, 1.261, 1.435, 1.488, 1.524, 1.584, 1.763, 1.896, 2.02]. In this example, the value 1.763 is also the R-=R+ value. This set is then used to generate the weights for the NN and the corresponding model S. A set of 20 resistors is more than needed, since the mean square output error of model S was 11 mV, considering the relative resistance error was close to zero. The maximum error across the set of input data was calculated to be 33 mV. In one example, an S, DAC, and ADC converter with 256 levels were analyzed as separate models, and the results showed a mean square output error of 14 mV and a maximum output error of 49 mV. An output error of 45 mV on the NN corresponds to a relative recognition error of 1%. An output error value of 45 mV also corresponds to an acceptable relative weight error of 0.01 or an absolute weight error of 0.01. The maximum weight modulus within the NN is 1.94. In this way, an optimal (or suboptimal) resistor set is determined using an iterative process based on the desired weight range [-wlim;wlim], resistor error (relative), and possible resistor range.

[0228]

[0281] Typically, a very wide resistor set is not very useful (e.g., 1 to 1 / 5 orders of magnitude is sufficient) unless different precision is required within different layers or portions of the weight spectrum. For example, if the weights are in the [0, 1] range, but most of the weights are in the [0, 0.001] range, better precision within that range is needed. In the example above, considering that the relative resistor error is close to zero, a set of 20 resistors is more than sufficient to quantize the NN network with a given precision. In one example, with the set of resistors [0.300, 0.461, 0.519, 0.566, 0.648, 0.655, 0.689, 0.996, 1.006, 1.048, 1.186, 1.222, 1.261, 1.435, 1.488, 1.524, 1.584, 1.763, 1.896, 2.02] (values ​​listed are relative), an average S output error of 11 mV was obtained.

[0229] Exemplary Process for Resistor Quantization

[0282] In some implementations, the example calculations described herein are performed by a weight matrix calculation or weight quantization module 238 (e.g., using a resistance calculation module 240) that calculates the weights 272 and / or the corresponding resistance values ​​242 of the weights 272 for the connections of the transformed neural network.

[0230]

[0283] This section describes an exemplary process for quantizing resistance values ​​corresponding to trained neural network weights, according to some implementations. The exemplary process substantially simplifies the process of manufacturing chips using analog hardware components to implement neural networks. As described above, some implementations use resistors to represent neural network weights and / or biases for operational amplifiers that represent analog neurons. The exemplary process described herein specifically reduces the complexity of lithographically fabricating a set of resistors for a chip. The procedure for quantizing the resistance values ​​requires only the selected value of the resistors for chip fabrication. In this manner, the exemplary process simplifies the overall process of chip fabrication and enables on-demand automatic resistor lithography mask fabrication.

[0231]

[0284] 18 provides an exemplary scheme of a neuron model 1800 used for resistor quantization according to some implementations. In some implementations, the circuit is based on an operational amplifier 1824 (e.g., an AD824 series precision amplifier) ​​that receives input signals from negative weight fixed resistors (R1- 1804, R2- 1806, Rb-bias 1816, Rn- 1818, and R- 1812) and positive weight fixed resistors (R1+ 1808, R2+ 1810, Rb+bias 1820, Rn+ 1822), and R+ 1814). The positive weight voltages are fed to the direct input of the operational amplifier 1824, and the negative weight voltages are fed to the inverting input of the operational amplifier 1824. The operational amplifier 1824 is used to enable a weighted sum of the weighted outputs from each resistor, where the negative weights are subtracted from the positive weights. The operational amplifier 1824 also amplifies the signal to an extent necessary for circuit operation. In some implementations, the operational amplifier 1824 also performs RELU conversion of the output signal at the output cascade of the operational amplifier 1824.

[0232]

[0285] The following formula determines the weight based on the resistor value: • The voltage at the output of a neuron is determined by the following equation:

number

number

[0233]

[0286] The following exemplary optimization procedure, according to some implementations, quantizes the value of each resistor to minimize the error in the neural network output. 1. Obtain the set of connection weights and biases {w1,...,wn,b}. 2. Obtain the minimum and maximum possible resistor values ​​{Rmin, Rmax}. These parameters are determined based on the technology used for fabrication. Some implementations use TaN or tellurium high-resistivity materials. In some implementations, the minimum resistor value is determined by the least squares fit that can be formed by lithography. The maximum value is determined by the allowable length for a resistor (e.g., a resistor made from TaN or tellurium) to fit into the desired area, which is determined by the area of ​​the operational amplifier section on the lithography mask. In some implementations, because the array of resistors is stacked (e.g., one in the BEOL and another in the FEOL), the area of ​​the resistor array is smaller than the area of ​​a single operational amplifier. 3. Assume each resistor has a relative tolerance value of r_err. 4. The goal is to select a set of resistance values ​​{R1,...,Rn} for a given length N within a defined [Rmin;Rmax] based on the {w1,...,wn,b} values. An exemplary search algorithm for finding a suboptimal {R1,...,Rn} set based on a specific optimality criterion is provided below. 5. Another algorithm is to select {Rn, Rp, Rni, Rpi} for the network when {R1..Rn} is determined.

[0234] Example {R1,...,Rn} Search Algorithm

[0287] Some implementations use an iterative approach for resistor set search. Some implementations select an initial (random or uniform) set {R1,...,Rn} within a defined range. Some implementations select one of the elements of the resistor set as R-=R+ value. Some implementations modify each resistor in the set by the current learning rate value until such modification produces a "better" set (according to the value function). This process is repeated for all resistors in the set with several different learning rate values ​​until no further improvement is possible.

[0235]

[0288] Some implementations define the value function for the resistor set as follows: • The possible weight options are calculated according to the formula (described above).

number

[0236]

[0289] The required weight range of the model [-wlim;wlim] is set to [-5;5], and other parameters include N=20, r_err=0.1%, rmin=100KΩ, rmax=5MΩ, where rmin and rmax are the minimum and maximum values ​​of the resistance, respectively.

[0237]

[0290] In one example, the following resistor set of length 20 was obtained for the above parameters: [0.300, 0.461, 0.519, 0.566, 0.648, 0.655, 0.689, 0.996, 1.006, 1.048, 1.186, 1.222, 1.261, 1.435, 1.488, 1.524, 1.584, 1.763, 1.896, 2.02] MΩ.

[0238] Exemplary {Rn,Rp,Rni,Rpi} Search Algorithm

[0291] Some implementations determine Rn and Rp using an iterative algorithm, such as the algorithm described above. Some implementations set Rp = Rn (the task of determining Rn and Rp is symmetric, and the two quantities typically converge to similar values). Then, the weights w i For each, some implementations select the resistor pair {Rni, Rpi} that minimizes the estimated weight error value.

number

[0239]

[0292] Some implementations then use the set {Rni;Rpi;Rn;Rp} values ​​to implement the neural network schematic. In one example, the schematic, according to some implementations, produced a mean square output error (sometimes referred to as the S mean square output error described above) of 11 mV and a maximum error of 33 mV over a set of 10,000 uniformly distributed input data samples. In one example, the S model was analyzed as a separate model along with a digital-to-analog converter (DAC) and an analog-to-digital converter (ADC) with 256 levels. The model, according to some implementations, produced a mean square output error of 14 mV and a maximum output error of 49 mV for the same data set. DACs and ADCs have levels because they convert analog values ​​to bit values ​​and vice versa. An 8-bit digital value equals 256 levels. For an 8-bit ADC, precision cannot be better than 1 / 256.

[0240]

[0293] Some implementations use Mathcad or any other similar software to calculate the resistance values ​​of the analog IC chip when the connection weights are known, based on Kirchhoff's circuit laws and the basic principles of operational amplifiers (described below with reference to FIG. 19A). In some implementations, operational amplifiers are used to both amplify the signal and transform it with an activation function (e.g., ReLU, sigmoid, tangent hyperbolic, or linear mathematical equation).

[0241]

[0294] Some implementations fabricate resistors in lithographic layers, where the resistors are formed as cylindrical holes in a SiO2 matrix, with the resistance value set by the diameter of the hole. Some implementations use amorphous TaN, CrN, TiN, or tellurium as the high-resistivity material to create high-density resistor arrays. Several ratios of Ta to N, Ti to N, and Cr to N provide high resistance to create ultra-high-density, high-resistivity element arrays. For example, for TaN, Ta5N6, and Ta3N5, the higher the N to Ta ratio, the higher the resistivity. Some implementations use Ti2N, TiN, CrN, or Cr5N and scale the ratios accordingly. TaN deposition is a standard procedure used in chip manufacturing and is available at all major foundries.

[0242] Exemplary Operational Amplifier

[0295] FIG. 19A shows a schematic diagram of an operational amplifier made in CMOS (CMOS op amp) 1900 according to some implementations. In FIG. 19A, In+ (positive input or pos) 1404, In− (negative input or neg) 1406, and Vdd− (positive supply voltage relative to GND) 1402 are contact inputs. Contact Vss− (negative supply voltage or GND) is indicated by label 1408. The circuit output is Out 1410 (contact output). The parameters of CMOS transistors are determined by the ratio of geometric dimensions, i.e., L (gate channel length) to W (gate channel width), examples of which are shown in the table shown in FIG. 19B (described below). The current mirror is made up of NMOS transistors M11 1944, M12 1946 and resistor R1 1921 (an exemplary resistance value of 12 kΩ) and provides the offset current for the differential pair (M1 1926 and M3 1930). The differential amplifier stage (differential pair) is made up of NMOS transistors M1 1926 and M3 1930. Transistors M1 and M3 are for amplification, while PMOS transistors M2 1928 and M4 1932 act as active current loads. From transistor M3, the signal is input to the gate of output PMOS transistor M7 1936. From transistor M1, the signal is input to the active load for PMOS transistor M5 (inverter) 1934 and NMOS transistor M6 1934. The current flowing through transistor M5 1934 sets the current for NMOS transistor M8 1938. Transistor M7 1936 is included in the common source scheme for the positive half-wave signal. Transistor M8 1938 is enabled by the common source circuit for the negative half-wave signal. To increase the overall load capacity of the operational amplifier, the outputs of M7 1936 and M8 1938 include an inverter with transistors M9 1940 and M10 1942. Capacitors C1 1912 and C2 1914 are for blocking.

[0243]

[0296] FIG. 19B shows a table 1948 of descriptions of the exemplary circuit shown in FIG. 19A , according to some implementations. The parameter values ​​are provided as examples, and various other configurations are possible. Transistors M1, M3, M6, M8, M10, M11, and M12 are N-channel MOSFET transistors with explicit substrate connections. The other transistors M2, M4, M5, M7, and M9 are P-channel MOSFET transistors with explicit substrate connections. The table shows exemplary shutter ratios of length (L, column 1) and width (W, column 2) given to each of the transistors (column 3).

[0244]

[0297] In some implementations, operational amplifiers such as the examples described above are used as fundamental elements of integrated circuits for hardware implementations of neural networks. In some implementations, the operational amplifiers are 40 square microns in size and fabricated according to the 45 nm node standard.

[0245]

[0298] In some implementations, activation functions such as ReLU, hyperbolic tangent, and sigmoid functions are represented by operational amplifiers with modified output cascades. For example, a RELU, sigmoid, or tangent function is realized as an output cascade of operational amplifiers (sometimes called op-amps) using corresponding well-known analog schematics according to some implementations.

[0246]

[0299] In the examples described above and below, in some implementations, operational amplifiers are replaced by inverters, current mirrors, two-quadrant or four-quadrant multipliers and / or other analog function blocks to enable weighted sum operations.

[0247] An example scheme of an LSTM block

[0300] 20A-20E show a schematic diagram of an LSTM neuron 20000 according to some implementations. The neuron's inputs are Vin1 20002 and Vin2 20004, which are values ​​in the range [-0.1, 0.1]. The LSTM neuron also inputs the resulting value of the neuron's calculation at time H(t-1) (previous value; see above for the explanation of LSTM neurons) 20006 and the neuron's state vector at time C(t-1) (previous value) 20008. The neuron LSTM's output (shown in FIG. 20B) includes the result of calculating the neuron at the current time H(t) 20118 and the neuron's state vector at the current time C(t) 20120. The scheme includes the following: 20A, a "neuron O" assembled with operational amplifiers U1 20094 and U2 20100. Resistors R_Wo1 20018, R_Wo2 20016, R_Wo3 20012, R_Wo4 20010, R_Uop1 20014, R_Uom1 20020, Rr 20068, and Rf2 20066 set the connection weights of a single "neuron O." "Neuron O" uses a sigmoid (module X1 20078, FIG. 20B) as a nonlinear function. "Neuron C" assembled with operational amplifiers U3 20098 (shown in FIG. 20C) and U4 20100 (shown in FIG. 20A). Resistors R_Wc1 20030, R_Wc2 20028, R_Wc3 20024, R_Wc4 20022, R_Ucp1 20026, R_Ucm1 20032, Rr 20122, and Rf2 20120 set the connection weights of "Neuron C". "Neuron C" uses the hyperbolic tangent (module X2 22080, FIG. 2B) as a nonlinear function. "Neuron I" is assembled with operational amplifiers U5 20102 and U6 20104, as shown in Figure 20C. Resistors R_Wi1 20042, R_Wi2 20040, R_Wi3 20036 and R_Wi4 20034, R_Uip1 20038, R_Uim1 20044, Rr 20124 and Rf2 20126 set the connection weights of "Neuron I". "Neuron I" uses a sigmoid (module X3 20082) as a nonlinear function. As shown in Figure 20D, "neuron f" resistors R_Wf1 20054, R_Wf2 20052, R_Wf3 20048, R_Wf4 20046, R_Ufp1 20050, R_Ufm1 20056, Rr 20128 and Rf2 20130 assembled with operational amplifiers U7 20106 and U8 20108 set the weights of the connections of "neuron f". "Neuron f" uses a sigmoid (module X4 20084) as a nonlinear function.

[0248]

[0301] The outputs of modules X2 20080 (FIG. 20B) and X3 20082 (FIG. 20C) are input to the X5 multiplier module 20086 (FIG. 20B). The outputs of modules X4 20084 (FIG. 20D) and the buffer to U9 20010 are input to the multiplier module X6 20088. The outputs of modules X5 20086 and X6 20088 are input to the adder (U10 20112). The divider 10 is assembled with resistors R1 20070, R2 20072 and R3 20074. A non-linear function of the hyperbolic tangent (module X7 20090, FIG. 20B) is obtained at the emission of the divisor signal. The output C(t) 20120 (the current state vector of the LSTM neuron) is obtained from a buffer inverter for the U11 20114 output signal. The outputs of modules X1 20078 and X7 20090 are input to a multiplier (module X8 20092) whose output is input to a divide-by-10 for U12 20116. The result of calculating the LSTM neuron at the current time H(t) 20118 is obtained from the output signal of U12 20116.

[0249]

[0302] 20E shows example values ​​of different configurable parameters (e.g., voltages) of the circuits shown in FIGS. 20A-20D, according to some implementations. According to some implementations, Vdd 20058 is set to +1.5V, Vss 20064 is set to −1.5V, Vdd1 20060 is set to +1.8V, Vss1 20062 is set to −1.0V, and GND 20118 is set to GND.

[0250]

[0303] FIG. 20F shows a table 20132 of descriptions of the exemplary circuits shown in FIGS. 20A-20D, according to some implementations. Parameter values ​​are provided as examples, and various other configurations are possible. Transistors U1-U12 are CMOS op-amps (described above with reference to FIGS. 19A and 19B). X1, X3, and X4 are modules that perform a sigmoid function. X2 and X7 are modules that perform a hyperbolic tangent function. X5 and X8 are modules that perform a multiplication function. Example resistor ratings include Rw=10 kΩ and Rr=1.25 kΩ. Other resistors are expressed relative to Rw. For example, Rf2=12×Rw, R_Wo4=5×Rw, R_Wo3=8×Rw, R_Uop1=2.6×Rw, R_Wo2=12×Rw, R_W1=w×Rw and R_Uom1=2.3×Rw, R_Wc4=4×Rw, R_Wc3=5.45×Rw, R_Ucp1=3×Rw, R_Wc2=12×Rw, R_Wc1=2.72×Rw, R_Ucm1 = 3.7 × Rw, R_Wi4 = 4.8 × Rw, W_Wi3 = 6 × Rw, W_Uip1 = 2 × Rw, R_Wi2 = 12 × Rw, R_Wi1 = 3 × Rw, R_Uim1 = 2.3 × Rw, R_Wf4 = 2.2 × Rw, R_Wf3 = 5 × Rw, R_Wfp = 4 × Rw, R_Wf2 = 2 × Rw, R_Wf1 = 5.7 × Rw and Rfm1 = 4.2 × Rw.

[0251] Example Scheme of a Multiplier Block

[0304] 21A-21I show schematic diagrams of a multiplier block 21000 according to some implementations. The neuron 21000 is based on the principle of a four-quadrant multiplier, constructed using operational amplifiers U1 21040 and U2 21042 (shown in FIG. 21B), U3 21044 (shown in FIG. 21H), and U4 21046 and U5 21048 (shown in FIG. 21I) and CMOS transistors M1 21052-M68 21182. The inputs of the multiplier include V_one 21020 21006 and V_two 21008 (shown in FIG. 21B), as well as a node Vdd (positive power supply voltage, e.g., +1.5 V with respect to GND) 21004 and a node Vss (negative power supply voltage, e.g., −1.5 V with respect to GND) 21002. In this scheme, additional power supply voltages are used: contact input Vdd1 (positive power supply voltage, e.g., +1.8 V relative to GND), contact Vss1 (negative power supply voltage, e.g., −1.0 V relative to GND). The result of the circuit calculation is output to mult_out (output pin) 21170 (shown in FIG. 21I).

[0252]

[0305] 21B, the input signal (V_one) from V_one 21006 is connected to an inverter with a gain of one at U1 21040, the output of which forms a signal negA 21006 of equal amplitude but opposite sign to the signal V_one. Similarly, the signal (V_two) from input V_two 21008 is connected to an inverter with a gain of one at U2 21042, the output of which forms a signal negB 21012 of equal amplitude but opposite sign to the signal V_two. Paired combinations of signals from the possible combinations (V_one, V_two, negA, negB) are output to corresponding mixers using CMOS transistors.

[0253]

[0306] Referring again to Figure 21A, V_two 21008 and negA 21010 are input to a multiplexer constructed from NMOS transistors M19 21086, M20 21088, M21 21090, M22 21092 and PMOS transistors M23 21094 and M24 21096. The output of this multiplexer is input to NMOS transistor M6 21060 (Figure 21D).

[0254]

[0307] Similar transformations that occur in signals include: negB 21012 and V_one 21020 are input to a multiplexer constructed from NMOS transistors M11 21070, M12 2072, M13 2074, M14 21076 and PMOS transistors M15 2078 and M16 21080. The output of this multiplexer is input to M5 21058 NMOS transistor (shown in Figure 21D). V_one 21020 and negB 21012 are input to a multiplexer constructed with PMOS transistors M18 21084, M48 21144, M49 21146, and M50 21148, and NMOS transistors M17 21082, M47 21142. The output of this multiplexer is input to M9 PMOS transistor 21066 (shown in FIG. 21D), negA 21010 and V_two 21008 are input to a multiplexer constructed from PMOS transistors M52 21152, M54 21156, M55 21158 and M56 21160 and NMOS transistors M51 21150 and M53 21154. The output of this multiplexer is input to M2 NMOS transistor 21054 (shown in FIG. 21C). negB 21012 and V_one 21020 are input to a multiplexer constructed from NMOS transistors M11 21070, M12 21072, M13 21074 and M14 21076 and PMOS transistors M15 21078 and M16 21080. The output of this multiplexer is input to M10 NMOS transistor 21068 (shown in FIG. 21D). negB 21012 and negA 21010 are input to a multiplexer constructed from NMOS transistors M35 21118, M36 21120, M37 21122 and M38 21124 and PMOS transistors M39 21126 and M40 21128. The output of this multiplexer is input to M27 PMOS transistor 21102 (shown in FIG. 21H). V_two 21008 and V_one 21020 are input to a multiplexer constructed from NMOS transistors M41 21130, M42 21132, M43 21134 and M44 21136 and PMOS transistors M45 21138 and M46 21140. The output of this multiplexer is input to M30 NMOS transistor 21108 (shown in FIG. 21H). V_one 21020 and V_two 21008 are input to a multiplexer constructed from PMOS transistors M58 21162, M60 21166, M61 21168 and M62 21170 and NMOS transistors M57 21160 and M59 21164. The output of this multiplexer is input to M34 PMOS transistor 21116 (shown in FIG. 21H). NegA 21010 and negB 21012 are input to a multiplexer constructed from PMOS transistors M64 21174, M66 21178, M67 21180 and M68 21182 and NMOS transistors M63 21172 and M65 21176. The output of this multiplexer is input to PMOS transistor M33 21114 (shown in FIG. 21H).

[0255]

[0308] A current mirror (transistors M1 21052, M2 21053, M3 21054, and M4 21056) powers the portion of the four-quadrant multiplier shown on the left, made up of transistors M5 21058, M6 21060, M7 21062, M8 21064, M9 21066, and M10 21068. A current mirror (with transistors M25 21098, M26 21100, M27 21102, and M28 21104) powers the right portion of the four-quadrant multiplier, made up of transistors M29 21106, M30 21108, M31 21110, M32 21112, M33 21114, and M34 21116. The multiplication result is obtained from resistor Ro 21022 enabled in parallel with transistor M3 21054 and resistor Ro 21188 enabled in parallel with transistor M28 21104 and is fed to a summer with U3 21044. The output of U3 21044 is fed to a summer with a gain of 7:1 assembled with U5 21048 as shown in FIG. 21I, whose second input is compensated by resistors R1 21024 and R2 21026 and a reference voltage set by buffer U4 21046. The multiplication result is output from the output of U5 21048 via the Mult_Out output 21170.

[0256]

[0309] 21J shows a table 21198 of explanations for the schematic diagrams shown in FIGS. 21A-21I, according to some implementations. U1-U5 are CMOS operational amplifiers. N-channel MOSFET transistors with explicit substrate connections include transistors M1, M2, M25, and M26 (with a shutter ratio of length (L)=2.4u and a shutter ratio of width (W)=1.26u), transistors M5, M6, M29, and M30 (with L=0.36u and W=7.2u), transistors M7, M8, M31, and M32 (with L=0.36u and W=199.98u), transistors M11-M14, M19-M22, M35-M38, and M41-M44 (with L=0.36u and W=0.4u), and transistors M17, M47, M51, M53, M57, M59, M43, and M64 (with L=0.36u and W=0.72u). P-channel MOSFET transistors with explicit substrate connections include transistors M3, M4, M27, and M28 (with a shutter ratio of length (L)=2.4u and a shutter ratio of width (W)=1.26u), transistors M9, M10, M33, and M34 (with L=0.36u and W=7.2u), transistors M18, M48, M49, M50, M52, M54, M55, M56, M58, M60, M61, M62, M64, M66, M67, and M68 (with L=0.36u and W=0.8u), and transistors M15, M16, M23, M24, M39, M40, M45, and M46 (with L=0.36u and W=0.72u). Exemplary resistor ratings include Ro=1 kΩ, Rin=1 kΩ, Rf=1 kΩ, Rc4=2 kΩ, and Rc5=2 kΩ, according to some implementations.

[0257] Example scheme of a sigmoid block

[0310] 22A shows a schematic diagram of a sigmoid block 2200 according to some implementations. The sigmoid function (e.g., modules X1 20078, X3 20082, and X4 20084 described above with reference to FIGS. 20A-20F) is implemented using operational amplifiers U1 2250, U2 2252, U3 2254, U4 2256, U5 2258, U6 2260, U7 2262, and U8 2264 and NMOS transistors M1 2266, M2 2268, and M3 2270. Contact Sigm_in 2206 is the module input, contact input Vdd1 2222 is a positive supply voltage +1.8 V relative to GND 2208, and contact Vss1 2204 is a negative supply voltage −1.0 V relative to GND. In this scheme, U4 2256 has a reference voltage source of −0.2332V, with the voltage set by dividers R10 2230 and R11 2232. U5 2258 has a reference voltage source of 0.4V, with the voltage set by dividers R12 2234 and R13 2236. U6 2260 has a reference voltage source of 0.32687V, with the voltage set by dividers R14 2238 and R15 2240. U7 2262 has a reference voltage source of −0.5V, with the voltage set by dividers R16 2242 and R17 2244. U8 2264 has a reference voltage source of −0.33V, with the voltage set by dividers R18 2246 and R19 2248.

[0258]

[0311] A sigmoid function is formed by applying a corresponding reference voltage to the differential module formed by transistors M1 2266 and M2 2268. The current mirror for the differential stage is formed by active regulation operational amplifier U3 2254 and NMOS transistor M3 2270. The signal from the differential stage is tapped by NMOS transistor M2 and resistor R5 2220 and input to summer U2 2252. The output signal sigm_out 2210 is tapped from the U2 summer 2252 output.

[0259]

[0312] 22B shows a table 2278 of descriptions for the schematic diagram shown in FIG. 22A, according to some implementations. U1-U8 are CMOS op amps. M1, M2, and M3 are N-channel MOSFET transistors with a shutter ratio of length (L)=0.18 u and a shutter ratio of width (W)=0.9 u, according to some implementations.

[0260] Example scheme of hyperbolic tangent blocks

[0313] 23A shows a schematic diagram of a hyperbolic tangent function block 2300 according to some implementations. The hyperbolic tangent function (e.g., modules X2 20080 and X7 20090 described above with reference to FIGS. 20A-20F) is implemented using operational amplifiers (U1 2312, U2 2314, U3 2316, U4 2318, U5 2320, U6 2322, U7 2328, and U8 2330) and NMOS transistors (M1 2332, M2 2334, and M3 2336). In this scheme, node tanh_in 2306 is the module input, node input Vdd1 2304 is at a positive supply voltage of +1.8 V relative to GND 2308, and node Vss1 2302 is at a negative supply voltage of −1.0 V relative to GND. Further, in this scheme, U4 2318 has a reference voltage source of −0.1V, the voltage set by dividers R10 2356 and R11 2358. U5 2320 has a reference voltage source of 1.2V, the voltage set by dividers R12 2360 and R13 2362. U6 2322 has a reference voltage source of 0.32687V, the voltage set by dividers R14 2364 and R15 2366. U7 2328 has a reference voltage source of −0.5V, the voltage set by dividers R16 2368 and R17 2370. U8 2330 has a reference voltage source of −0.33V, the voltage set by dividers R18 2372 and R19 2374. The hyperbolic tangent function is formed by adding the corresponding reference voltage with the differential module made by transistors M1 2332 and M2 2334. The current mirror for the differential stage is obtained with active regulation operational amplifier U3 2316 and NMOS transistor M3 2336. With NMOS transistor M2 2334 and resistor R5 2346, the signal is tapped off from the differential stage and input to summer U2 2314. The output signal tanh_out 2310 is tapped off from the U2 summer 2314 output.

[0261]

[0314] 23B shows a table 2382 of descriptions for the schematic diagram shown in FIG. 23A, according to some implementations. U1-U8 are CMOS op-amps, and M1, M2, and M3 are N-channel MOSFET transistors with a shutter ratio of length (L)=0.18 u and a shutter ratio of width (W)=0.9 u.

[0262] Exemplary scheme of a single neuron OP1 CMOS operational amplifier

[0315] 24A-24C show a schematic diagram of a single neuron OP1 CMOS operational amplifier 2400 according to some implementations. The example is a variant of a single neuron operational amplifier made in CMOS according to the OP1 scheme described herein. In this scheme, contacts V1 2410 and V2 2408 are the inputs of the single neuron, contact Bias 2406 is at a voltage of +0.4 V relative to GND, contact Input Vdd 2402 is at a positive supply voltage of +5.0 V relative to GND, contact Vss 2404 is at GND, and contact Out 2474 is the output of the single neuron. The parameters of the CMOS transistor are determined by the ratio of the geometric dimensions, namely, L (gate channel length) and W (gate channel width). This op-amp has two current mirrors. The current mirror of NMOS transistors M3 2420, M6 2426, and M13 2440 provides the offset current for the differential pair of NMOS transistors M2 2418 and M5 2424. The current mirror of PMOS transistors M7 2428, M8 2430, and M15 2444 provides the offset current for the differential pair of PMOS transistors M9 2432 and M10 2434. In the first differential amplifier stage, NMOS transistors M2 2418 and M5 2424 are for amplification, and PMOS transistors M1 2416 and M4 2422 act as active current loads. From transistor M5 2424, a signal is output to the PMOS gate of transistor M13 2440. From transistor M2 2418, the signal is output to the right input of the second differential amplifier stage by PMOS transistors M9 2432 and M10 2434. NMOS transistors M11 2436 and M12 2438 act as active current loads for transistors M9 2432 and M10 2434. Transistor M17 2448 is switched on according to a scheme with a common source for the positive half-wave of the signal. Transistor M18 2450 is switched on according to a scheme with a common source for the negative half-wave of the signal.To increase the overall load capacitance of the OP amp, an inverter formed by transistors M17 2448 and M18 2450 is enabled at the outputs of transistors M13 2440 and M14 2442.

[0263]

[0316] Figure 24D shows Table 2476, which explains the schematic diagrams shown in FIGS. 24A - 24C according to some implementation modes. The connection weights of a single neuron (having two inputs and one output) are set by the resistance ratios: w1 = (Rp / R1+) - (Rn / R1-); w2 = (Rp / R2+) - (Rn / R2-); w bias = (Rp / Rbias+) - (Rn / Rbias-). The normalization resistors (Rnorm- and Rnorm+) are necessary to obtain the exact equation, namely (Rn / R1-) + (Rn / R2-) + (Rn / Rbias-) + (Rn / Rnorm-) = (Rp / R1+) + (Rp / R2+) + (Rp / Rbias+) + (Rp / Rnorm+). Examples of N-channel MOSFET transistors having explicit substrate connections include transistors M2 and M5 with L = 0.36u and W = 3.6u, transistors M3, M6, M11, M12, M14 and M16 with L = 0.36u and W = 1.8u, and transistor M18 with L = 0.36u and W = 18u. Examples of P-channel MOSFET transistors having explicit substrate connections include transistors M1, M4, M7, M8, M13 and M15 with L = 0.36u and W = 3.96u, transistors M9 and M10 with L = 0.36u and W = 11.88u, and transistor M17 with L = 0.36u and W = 39.6u.

[0264] Exemplary scheme of a single neuron OP3 CMOS op amp

[0317] 25A-25D show schematic diagrams of variants of a single neuron 25000 with operational amplifiers, fabricated in CMOS according to the OP3 scheme, according to some implementations. The single neuron consists of three simple operational amplifiers (opamps), according to some implementations. The unit neuron adder is implemented with two opamps with bipolar power supplies, and the RELU activation function is implemented with an opamp with unipolar power supplies and a gain of 10. Transistors M1 25028-M16 25058 are used to sum the negative connections of the neurons. Transistors M17 25060-M32 25090 are used to add the positive connections of the neurons. The RELU activation function is implemented with transistors M33 25092-M46 25118. In the scheme, contacts V1 25008 and V2 25010 are the inputs of a single neuron, contact Bias 25002 has a voltage of +0.4V relative to GND, contact Input Vdd 25004 has a positive power supply voltage of +2.5V relative to GND, contact Vss 25006 has a negative power supply voltage of -2.5V, and contact Out 25134 is the output of the single neuron. The parameters of the CMOS transistors used in a single neuron are determined by the ratio of the following geometric dimensions: L (length of the gate channel) and W (width of the gate channel). Consider the operation of the simplest operational amplifier contained in a single neuron. Each operational amplifier has two current mirrors. The current mirror formed by NMOS transistors M3 25032 (M19 25064, M35 25096), M6 25038 (M22 25070, M38 25102) and M16 25058 (M32 25090, M48 25122) provides the offset current for the differential pair formed by NMOS transistors M2 25030 (M18 25062, M34 25094) and M5 25036 (M21 25068, M35 25096).The current mirror of PMOS transistors M7 25040 (M23 25072, M39 25104), M8 25042 (M24 25074, M40 25106) and M15 25056 (M31 2588) provides the offset current for the differential pair of PMOS transistors M9 25044 (M25 25076, M41 25108) and M10 25046 (M26 25078, M42 25110). In the first differential amplifier stage, NMOS transistors M2 25030 (M18 25062, M34 25094) and M5 25036 (M21 25068, M37 25100) are used for amplification, and PMOS transistors M1 25028 (M17 25060, M33 25092) and M4 25034 (M20 25066, M36 25098) act as active current loads. From transistor M5 25036 (M21 25068, M37 25100), the signal is input to the PMOS gate of transistor M13 25052 (M29 25084, M45 25116). From transistor M2 25030 (M18 25062, M34 25094), the signal is input to the right input of a second differential amplifier stage by PMOS transistors M9 25044 (M25 25076, M41 25108) and M10 25046 (M26 25078, M42 25110). NMOS transistors M11 25048 (M27 25080, M43 25112) and M12 25048 (M28 25080, M44 25114) act as active current loads for transistors M9 25044 (M25 25076, M41 25108) and M10 25046 (M26 25078, M42 25110). Transistor M13 25052 (M29 25082, M45 25116) is included in a scheme with a common source for the positive half-wave signal. Transistor M14 25054 (M30 25084, M46 25118) is switched on according to a scheme with a common source for the negative half-wave of the signal.

[0265]

[0318] The connection weights of a single neuron (with two inputs and one output) are set by the resistance ratios: w1 = (Rfeedback / R1+) - (Rfeedback / R1-); w2 = (Rfeedback / R2+) - (Rfeedback / R2-); wbias = (Rfeedback / Rbias+) - (Rfeedback / Rbias-); w1 = (Rp*Kamp / R1+) - (Rn*Kamp / R1-); w2 = (Rp*Kamp / R2-); wbias = (Rp*Kamp / Rbias+) - (Rn*Kamp / Rbias-), where Kamp = R1ReLU / R2ReLU. Rfeedback = 100k - used only to calculate w1, w2, and wbias. According to some implementations, example values ​​include: R feedback = 100k, Rn = Rp = Rcom = 10k, K amp ReLU = 1 + 90k / 10k = 10, w1 = (10k * 10 / 22.1k) - (10k * 10 / 21.5k) = -0.126276, w2 = (10k * 10 / 75k) - (10k * 10 / 71.5k) = -0.065268, wbias = (10k * 10 / 71.5k) - (10k * 10 / 78.7k) = 0.127953.

[0266]

[0319] The inputs of the neuron's negative link adders (M1-M17) are received from the neuron's positive link adders (M17-M32) through the Rcom resistors.

[0267]

[0320] 25E shows a table 25136 of explanations for the schematic diagrams shown in FIGS. 25A-25D, according to some implementations. N-channel MOSFET transistors with explicit substrate connections include transistors M2, M5, M18, M21, M34, and M37 with L=0.36 u and W=3.6 u, and transistors M3, M6, M11, M12, M14, M16, M19, M22, M27, M28, M32, M38, M35, M38, M43, M44, M46, and M48 with L=0.36 u and W=1.8 u. P-channel MOSFET transistors with explicit substrate connections include transistors M1, M4, M7, M8, M13, M15, M17, M20, M23, M24, M29, M31, M33, M36, M39, M40, M45, and M47 with L=0.36u and W=3.96u, and transistors M9, M10, M25, M26, M41, and M42 with L=0.36u and W=11.88u.

[0268] Exemplary Method for Analog Hardware Implementation of Trained Neural Networks

[0321] 27A-27J show a flowchart of a method 2700 for hardware implementation 2702 of a neural network, according to some implementations. The method is performed 2704 (e.g., using a neural network transformation module 226) on a computing device 200 having one or more processors 202 and a memory 214 storing one or more programs configured for execution by the one or more processors 202. The method includes obtaining 2706 a neural network topology (e.g., topology 224) and weights (e.g., weights 222) of a trained neural network (e.g., network 220). In some implementations, the trained neural network is trained 2708 using a software simulation to generate the weights.

[0269]

[0322] The method also includes converting the neural network topology to an equivalent analog network of analog components (2710). Referring now to FIG. 27C , in some implementations, the neural network topology includes one or more layers of neurons (2724). Each layer of neurons calculates a respective output based on a respective mathematical function. In such cases, converting the neural network topology to an equivalent analog network of analog components includes performing a series of steps for each of the one or more layers of neurons (2726). The series of steps includes identifying one or more function blocks for each layer based on the respective mathematical function (2728). Each function block has a respective schematic implementation having a block output that matches the output of the respective mathematical function. In some implementations, identifying the one or more function blocks includes selecting one or more function blocks based on the type of the respective layer (2730). For example, a layer can be composed of neurons, and the layer's output is a linear superposition of its inputs. Selecting one or more function blocks is based on this identification of the layer type if the layer's output is a linear superposition or similar pattern identification. Some implementations determine if the number of outputs > 1 and then use either a trapezoidal or pyramidal transform.

[0270]

[0323] 27D, in some implementations, the one or more function blocks include one or more basis function blocks (e.g., basis function block 232) selected 2734 from the group consisting of: (i) a block output

number

number

[0271]

[0324] Referring again to Figure 27C, the series of steps also includes generating (2732) a respective multi-layer network of analog neurons based on arranging one or more function blocks, each analog neuron implementing a respective function of the one or more function blocks, and each analog neuron in a first layer of the multi-layer network being connected to one or more analog neurons in a second layer of the multi-layer network.

[0272]

[0325] Referring again to FIG. 27A , for some networks, such as GRUs and LSTMs, according to some implementations, converting 2710 the neural network topology into an equivalent analog network of analog components requires more complex processing. Referring now to FIG. 27E , the neural network topology may include one or more layers of neurons 2746. Further, each layer of neurons may compute a respective output based on a respective mathematical function. In such cases, converting the neural network topology into an equivalent analog network of analog components may include: (i) decomposing 2748 a first layer of the neural network topology into multiple sublayers, the decomposition including decomposing the mathematical function corresponding to the first layer to obtain one or more intermediate mathematical functions. Each sublayer implements an intermediate mathematical function. In some implementations, the mathematical function corresponding to the first layer includes one or more weights, and decomposing the mathematical function includes adjusting the one or more weights (2750) so that combining one or more intermediate functions results in the mathematical function; and (ii) performing a series of steps for each sub-layer of the first layer of the neural network topology (2752). The series of steps includes selecting one or more sub-function blocks (2754) for each sub-layer based on the respective intermediate mathematical function, and generating a respective multi-layer analog subnetwork of analog neurons based on arranging the one or more sub-function blocks (2756). Each analog neuron implements a respective function of one or more sub-function blocks, and each analog neuron of the first layer of the multi-layer analog subnetwork is connected to one or more analog neurons of a second layer of the multi-layer analog subnetwork.

[0273]

[0326] Referring now to FIG. 27H , assume that the neural network topology includes one or more GRU neurons or LSTM neurons (2768). In that case, converting the neural network topology includes generating one or more signal delay blocks for each recurrent connection of the one or more GRU neurons or LSTM neurons (2770). In some implementations, an external cycle timer activates the one or more signal delay blocks for a fixed period (e.g., 1, 5, or 10 time steps). Some implementations use multiple delay blocks across a signal to generate an additive time shift. In some implementations, the activation frequency of the one or more signal delay blocks is synchronized / synchronized to the network input signal frequency. In some implementations, the one or more signal delay blocks are activated at a frequency that matches a predetermined input signal frequency for the neural network topology (2772). In some implementations, this predetermined input signal frequency may depend on the application, such as human activity recognition (HAR) or PPG. For example, the predetermined input signal frequency may be 30-60 Hz for video processing, approximately 100 Hz for HAR and PPG, 16 KHz for audio processing, and approximately 1-3 Hz for battery management. Some implementations activate different signal delay blocks at different frequencies.

[0274]

[0327] 27I, assume that the neural network topology includes one or more layers of neurons that implement an unconstrained activation function 2774. In some implementations, in such a case, transforming the neural network topology includes applying 2776 one or more transformations selected from the group consisting of: replacing 2778 the unconstrained activation function with a constrained activation function (e.g., replacing ReLU with thresholded ReLU), and adjusting 2780 the connections or weights of the equivalent analog network so that, for one or more given inputs, the difference in output between the trained neural network and the equivalent analog network is minimized.

[0275]

[0328] 27A, the method also includes calculating 2712 a weight matrix for the equivalent analog network based on the weights of the trained neural network, with each element of the weight matrix representing a respective connection between analog components of the equivalent analog network.

[0276]

[0329] The method also includes generating 2714 a schematic model for implementing an equivalent analog network based on the weight matrix, including selecting component values ​​for the analog components. Referring now to FIG. 27B , in some implementations, generating the schematic model includes generating 2716 a resistance matrix of the weight matrix. Each element of the resistance matrix corresponds to a respective weight in the weight matrix and represents a resistance value. In some implementations, the method includes regenerating only the resistance matrix of the resistors of the retrained network. In some implementations, the method further includes obtaining 2718 new weights for the trained neural network, calculating 2720 a new weight matrix of the equivalent analog network based on the new weights, and generating 2722 a new resistance matrix of the new weight matrix.

[0277]

[0330] 27J , in some implementations, the method further includes generating one or more lithography masks (e.g., generating masks 250 and / or 252 using mask generation module 248) for fabricating a circuit implementing the equivalent analog network of analog components based on the resistance matrix (2782). In some implementations, the method includes generating only a mask for the resistors of the retrained network (e.g., mask 250). In some implementations, the method further includes (i) obtaining new weights for the trained neural network (2784), (ii) calculating a new weight matrix for the equivalent analog network based on the new weights (2786), (iii) generating a new resistance matrix for the new weight matrix (2788), and (iv) generating a new lithography mask for fabricating a circuit implementing the equivalent analog network of analog components based on the new resistance matrix (2790).

[0278]

[0331] Referring now to FIG. 27G, the analog components include a plurality of operational amplifiers and a plurality of resistors (2762). Each operational amplifier represents an analog neuron of an equivalent analog network, and each resistor represents a connection between two analog neurons. Some implementations include other analog components such as four-quadrant multipliers, sigmoid and hyperbolic tangent function circuits, delay lines, adders, and / or dividers. In some implementations, selecting component values ​​for the analog components (2764) includes performing gradient descent and / or other weight quantization methods to identify possible resistance values ​​for the plurality of resistors (2766).

[0279]

[0332] 27F, in some implementations, the method further includes digitally implementing a particular activation function (e.g., softmax) in the output layer. In some implementations, the method further includes generating an equivalent digital network of digital components of the one or more output layers of the neural network topology (2758) and connecting outputs of the one or more layers of the equivalent analog network to the equivalent digital network of digital components (2760).

[0280] An exemplary method for constrained analog hardware implementation of neural networks

[0333] 28A-28S show a flowchart of a method 28000 for hardware realization 28002 of a neural network according to hardware design constraints, according to some implementations. The method is executed 28004 on a computing device 200 having one or more processors 202 (e.g., using a neural network transformation module 226), and the memory 214 stores one or more programs configured for execution by the one or more processors 202. The method includes obtaining 28006 a neural network topology (e.g., topology 224) and weights (e.g., weights 222) of a trained neural network (e.g., network 220).

[0281]

[0334] The method also includes calculating 28008 one or more coupling constraints based on analog integrated circuit (IC) design constraints (e.g., constraints 236). For example, the IC design constraints may set a current limit (e.g., 1 A), and the neuron schematic and operational amplifier (opamp) design may set the opamp output current in the range [0-10 mA], which limits the output neuron connections to 100. This means that the neuron has 100 outputs that pass current through 100 connections to the next layer, but because the current at the output of the opamp is limited to 10 mA, some implementations use a maximum of 100 outputs (0.1 mA x 100 = 10 mA). Without this constraint, some implementations increase the number of outputs to more than 100, for example, using current repeaters.

[0282]

[0335] The method also includes converting (28010) the neural network topology (eg, using the neural network conversion module 226) into an equivalent loosely connected network of analog components that satisfies one or more coupling constraints.

[0283]

[0336] In some implementations, transforming the neural network topology involves determining the number of possible input connections N according to one or more connection constraints. i and output coupling degree N o This includes deriving (28012).

[0284]

[0337] Referring now to FIG. 28B, in some implementations, a neural network topology includes at least one densely connected layer (28018) with K inputs (neurons in a previous layer) and L outputs (neurons in a current layer) and a weight matrix U, and transforming (28020) the at least one densely connected layer includes transforming the at least one densely connected layer (28020) with a weight matrix U. i and the output coupling is not more than N o K inputs, L outputs, and

number

[0285]

[0338] Referring now to FIG. 28C, in some implementations, a neural network topology includes at least one densely connected layer (28024) with K inputs (neurons in a previous layer) and L outputs (neurons in a current layer) and a weight matrix U, and transforming (28026) the at least one densely connected layer includes transforming the K inputs, L outputs, and

number

[0286]

[0339] Referring now to FIG. 28D, in some implementations, the neural network topology includes a single loosely connected layer, P, with K inputs and L outputs. i The maximum input coupling, P o The maximum output connectivity of the layer (28030) is K, and the weight matrix of U is U, with missing connections represented by zeros. In such a case, transforming a single loosely connected layer (28032) yields K inputs, L outputs,

number

[0287]

[0340] Referring now to FIG. 28E, in some implementations, the neural network topology includes a convolutional layer (e.g., a depthwise convolutional layer or a separable convolutional layer) with K inputs (neurons in the previous layer) and L outputs (neurons in the current layer) (28036). In such cases, converting the neural network topology to an equivalent loosely connected network of analog components (28038) involves converting the convolutional layer to a network with K inputs, L outputs, P i The maximum input coupling and P o a single loosely connected layer having a maximum output coupling of P i ≦N i and P o ≦N o is.

[0288]

[0341] 28A, the method also includes calculating 28014 a weight matrix of an equivalent loosely coupled network based on the weights of the trained neural network, with each element of the weight matrix representing a respective connection between analog components of the equivalent loosely coupled network.

[0289]

[0342] Referring now to FIG. 28F, in some implementations, the neural network topology includes a recurrent neural layer (28042), and converting the neural network topology to an equivalent loosely connected network of analog components (28044) includes converting the recurrent neural layer to one or more tightly or loosely connected layers with signal delay connections (28046).

[0290]

[0343] Referring now to FIG. 28G, in some implementations, the neural network topology includes a recurrent neural layer (e.g., a long short-term memory (LSTM) layer or a gated recurrent unit (GRU) layer), and converting the neural network topology into an equivalent loosely connected network of analog components includes decomposing the recurrent neural layer into several layers, at least one of which is equivalent to a densely or loosely connected layer with K inputs (neurons in the previous layer) and L outputs (neurons in the current layer) and a weight matrix U, where missing connections are represented by zeros.

[0291]

[0344] 28H, in some implementations, a method includes performing a transformation of a single-layer perceptron having one computational neuron. In some implementations, the neural network topology has K inputs, a weight vector U∈R, K and a single-layer perceptron with computational neurons having activation function F (28054). In such cases, converting the neural network topology to an equivalent loosely connected network of analog components (28056) involves (i) deriving the connectivity N of the equivalent loosely connected network subject to one or more connectivity constraints (28058), and (ii) calculating Eq.

number

number

[0292]

[0345] 28I, in some implementations, a method includes performing a transformation of a single-layer perceptron having L computational neurons. In some implementations, the neural network topology includes K inputs, a single-layer perceptron having L computational neurons, and a weight matrix V including a row of weights for each of the L computational neurons (28068). In such a case, transforming the neural network topology into an equivalent loosely connected network of analog components (28070) includes: (i) deriving a degree of connectivity N for the equivalent loosely connected network according to one or more coupling constraints (28072); (ii) calculating the equation

number

number

[0293]

[0346] Referring now to FIG. 28J, in some implementations, a method includes performing a multilayer perceptron transformation algorithm. In some implementations, a neural network topology includes a multilayer perceptron (28092) having K inputs and S layers, where each layer i of the S layers includes computational neurons L i And, L i The corresponding weight matrix V contains a row of the weights of each computational neuron. i In such a case, converting the neural network topology into an equivalent loosely coupled network of analog components (28094) involves (i) deriving the degree of connectivity N of the equivalent loosely coupled network subject to one or more coupling constraints (28096), (ii) converting the multilayer perceptron into a network with Q=Σ i=1、S (L i ) single-layer perceptron networks (28098), each of which includes a respective computational neuron of Q computational neurons. Decomposing the multi-layer perceptron includes duplicating one or more of the K inputs shared by the Q computational neurons; (iii) for each single-layer perceptron network (28100) of the Q single-layer perceptron networks, (a) Eq.

number

number

number

[0294]

[0347] Referring now to FIG. 28K, in some implementations, a neural network topology includes a convolutional neural network (CNN) with K inputs and S layers (28116), where each layer i of the S layers includes computational neurons L i And, L i The corresponding weight matrix V contains a row of the weights of each computational neuron. i In such a case, converting the neural network topology into an equivalent loosely connected network of analog components (28118) involves (i) deriving the degree of connectivity N of the equivalent loosely connected network subject to one or more coupling constraints (28120), (ii) converting the CNN into a network with Q=Σ i=1、S (L i ) single-layer perceptron networks (28122), each single-layer perceptron network including a respective computational neuron of the Q computational neurons. Decomposing the CNN includes duplicating one or more of the K inputs shared by the Q computational neurons, and (iii) for each single-layer perceptron network of the Q single-layer perceptron networks, (a) Eq.

number

number

number

[0295]

[0348] 28L, in some implementations, a method includes converting two layers into a trapezoid-based network. In some implementations, the neural network topology includes a layer L having K inputs and K neurons. p , layer L with L neurons n and weight matrix W∈R L×K (28140), where R is the set of real numbers and L p Each neuron in layer L n connected to each neuron in layer L n Each neuron in layer L performs an activation function F, which n The output of is the expression Y for the input x. o = F(Wx). In such cases, transforming the neural network topology into an equivalent loosely connected network of analog components (28142) is a trapezoidal transformation that (i) computes the number of possible input connections N, subject to one or more connection constraints. I >1 and possible output coupling N O > 1 (28144) and (ii) K·L <L·N I +K·N O a layer LA with K analog neurons that perform the identity activation function according to the determination that p , which runs the identity activation function

number

[0296]

[0349] Referring now to FIG. 28M, in some implementations, performing a trapezoidal transform is performed by: I +K·N O According to this determination, (i) layer L p Divide (28154) and K'·L≧L·N I +K'·N O A sublayer L with K' neurons is p1 and a sublayer L with (K-K') neurons p2 and (ii) a sublayer L with K′ neurons. p1 (iii) performing the constructing and generating steps for a sublayer L having K-K neurons; p2 and recursively performing the dividing, constructing, and generating steps for the .times. ...

[0297]

[0350] 28N, a method includes transforming a multilayer perceptron into a trapezoid-based network. In some implementations, the neural network topology includes a multilayer perceptron network (28160), and the method further includes iteratively performing a trapezoid transformation (28162) for each pair of successive layers of the multilayer perceptron network and calculating a weight matrix of an equivalent loosely connected network.

[0298]

[0351] Referring now to FIG. 28O, a method includes converting a recurrent neural network to a trapezoid-based network. In some implementations, the neural network topology includes a recurrent neural network (RNN) (28164) that includes (i) a linear combination computation of two fully connected layers, (ii) element-wise summation, and (iii) a nonlinear function computation. In such cases, the method further includes performing a trapezoid transformation (28166) on the (i) two fully connected layers and (ii) the nonlinear function computation, and computing a weight matrix for an equivalent loosely connected network. The element-wise summation is a general operation that can be implemented in networks of any structure, examples of which are provided above. The nonlinear function computation is a per-neuron operation independent of the No and Ni constraints, and is typically computed separately using a "sigmoid" or "tanh" block in each neuron.

[0299]

[0352] Referring now to Figure 28P, the neural network topology includes a long short-term memory (LSTM) network or a gated recurrent unit (GRU) network that includes (i) computation of a linear combination of multiple fully connected layers, (ii) element-wise summation, (iii) Hadamard product, and (iv) multiple nonlinear function computations (sigmoid and hyperbolic tangent operations) (28168). In such cases, the method further includes performing a trapezoidal transform (28170) on the (i) multiple fully connected layers and (ii) multiple nonlinear function computations, and computing a weight matrix for an equivalent loosely connected network. Element-wise summation and Hadamard product are common operations that can be implemented in networks of any of the above structures.

[0300]

[0353] Referring now to FIG. 28Q, the neural network topology includes a convolutional neural network (CNN) 28172 including (i) a plurality of partially connected layers (e.g., a sequence of convolutional layers and pruning layers; each pruning layer is assumed to be a convolutional layer with a stride greater than 1), and (ii) one or more fully connected layers (the sequence terminating with a fully connected layer). In such a case, the method further includes (i) converting the plurality of partially connected layers into equivalent fully connected layers by inserting missing connections with zero weights 28174, and for each pair of the equivalent fully connected layer and one or more successive fully connected layers, iteratively performing a trapezoidal transformation 28176 and calculating the weight matrix of an equivalent loosely connected network.

[0301]

[0354] Referring now to FIG. 28R, the neural network topology has K input, L output neurons and a weight matrix U∈R. L×K where R is a set of real numbers and each output neuron implements an activation function F. In such a case, converting the neural network topology to an equivalent loosely connected network of analog components (28180) involves performing an approximation transformation that includes: (i) determining the number of possible input connections N, subject to one or more connection constraints; I >1 and possible output coupling N O > 1 (28182), (ii) the set

number

number

number

number

number

[0302]

[0355] Referring now to FIG. 28S, in some implementations, a neural network topology has K inputs, S layers, and Li=1,S computational neurons and the weight matrix of the i-th layer

number

[0303]

[0356] Referring again to FIG. 28A, in some implementations, the method further includes generating (28016) a schematic model for implementing an equivalent loosely coupled network utilizing the weight matrix.

[0304] Exemplary Method for Calculating Resistor Values ​​for Analog Hardware Realization of Trained Neural Networks

[0357] 29A-29F show a flowchart of a method 2900 for hardware realization 2902 of a neural network according to hardware design constraints, according to some implementations. The method is performed 2904 (e.g., using the weight quantization module 238) on a computing device 200 having one or more processors 202 and a memory 214 storing one or more programs configured for execution by the one or more processors 202.

[0305]

[0358] The method includes obtaining 2906 a neural network topology (e.g., topology 224) and weights (e.g., weights 222) of a trained neural network (e.g., network 220). In some implementations, weight quantization is performed during training. In some implementations, the trained neural network has been trained 2908 such that each layer of the neural network topology has quantized weights (e.g., specific values ​​from a list of discrete values, e.g., each layer has only three weight values: +1, 0, and −1).

[0306]

[0359] The method also includes converting 2910 (e.g., using neural network conversion module 226) the neural network topology into an equivalent analog network of analog components including a plurality of operational amplifiers and a plurality of resistors, where each operational amplifier represents an analog neuron in the equivalent analog network and each resistor represents a connection between two analog neurons.

[0307]

[0360] The method also includes calculating 2912 a weight matrix for an equivalent analog network based on the weights of the trained neural network, each element of the weight matrix representing a respective connection.

[0308]

[0361] The method also includes generating 2914 a resistance matrix of the weight matrix, where each element of the resistance matrix corresponds to a respective weight in the weight matrix and represents a resistance value.

[0309]

[0362] 29B, in some implementations, generating the resistance matrix of the weight matrix includes a simplified gradient descent-based iterative method for finding the resistor set. In some implementations, generating the resistance matrix of the weight matrix includes: (i) finding a predetermined range of possible resistance values ​​{R 最小 ,R 最大} (2916) and obtain an initial base resistance value R within a predetermined range. ベースFor example, a range and base resistance are selected according to the values ​​of the elements of the weight matrix, the values ​​being determined by the manufacturing process and the range - a quantization of what can actually be manufactured, where large resistors are not preferred. In some implementations, the predetermined range of possible resistance values ​​includes resistors according to nominal series E24 in the range 100 KΩ to 1 MΩ (2918); and (ii) selecting a set of limited-length resistance values ​​within the predetermined range, the set being a set of {R i ,R j}, the range [-R ベース ,R ベース ] the most uniformly distributed possible weights

number

number

[0310]

[0363] Referring now to FIG. 29C , some implementations perform weight reduction. In some implementations, the first one or more weights and the first one or more inputs of the weight matrix represent one or more connections to a first operational amplifier of an equivalent analog network (2930). The method further includes, prior to generating the resistance matrix (2932), (i) modifying the first one or more weights by a first value (2934) (e.g., dividing the first one or more weights by the first value to reduce the weight range or multiplying the first one or more weights by the first value to increase the weight range), and (ii) configuring the first operational amplifier to multiply a linear combination of the first one or more weights and the first one or more inputs by the first value (2936) prior to performing the activation function. Some implementations perform weight reduction to change the multiplication coefficients of the one or more operational amplifiers. In some implementations, a set of resistance values ​​produces a range of weights, and in some portions of this range, the error is higher than in other portions. If there are only two nominal values ​​(e.g., 1 Ω and 4 Ω), these resistors can produce weights [-3;-0.75;0;0.75;3]. Given that the first layer of a neural network has weights of {0,9} and the second layer has weights of {0,1}, some implementations divide the weights of the first layer by 3 and multiply the weights of the second layer by 3 to reduce the overall error. Some implementations consider limiting the weight values ​​during training by adjusting the loss function (e.g., using an l1 or l2 regulator) so that the resulting network does not have weights that are too large for a resistor set.

[0311]

[0364] 29D , the method further includes limiting the weights to intervals, e.g., obtaining 2938 predetermined ranges of the weights such that an equivalent analog network will produce similar outputs as the trained neural network for the same inputs, and updating 2940 the weight matrix according to the predetermined ranges of the weights.

[0312]

[0365] Continuing with reference to FIG. 29E, the method further includes reducing the weight sensitivity of the network. For example, the method further includes retraining (2942) the trained neural network to reduce sensitivity to errors in weights or resistor values ​​that would cause an equivalent analog network to produce different outputs compared to the trained neural network. That is, some implementations include additional training of an already-trained neural network to make it less sensitive to small, randomly distributed weight errors. Quantization and resistor manufacturing produce small weight errors. Some implementations transform the network so that the resulting network is less sensitive to each particular weight value. In some implementations, this is performed by adding a small relative random value to each signal in at least some of the layers during training (e.g., similar to a dropout layer).

[0313]

[0366] Referring now to FIG. 29F, some implementations include reducing the weight distribution range. Some implementations include retraining (2944) the trained neural network to minimize weights in any layer that are greater than a predetermined threshold and greater than the average absolute weight of that layer. Some implementations perform this step via retraining. An exemplary penalty function includes the sum over all layers (e.g., A*max(abs(w)) / mean(abs(w)), where the maximum and average are calculated over the layers. Another example includes an order of magnitude greater than the above. In some implementations, this function affects weight quantization and network weight sensitivity. For example, small relative changes in weights due to quantization can cause high output error. An exemplary approach includes introducing some penalty function during training that penalizes the network when it experiences such weight waste.

[0314] Exemplary methods of optimization for analog hardware implementation of trained neural networks

[0367] 30A-30M show a flowchart of a method 3000 for hardware realization 3002 of a neural network according to hardware design constraints, according to some implementations. The method is performed 3004 (e.g., using the neural network optimization module 246) on a computing device 200 having one or more processors 202 and a memory 214 storing one or more programs configured for execution by the one or more processors 202.

[0315]

[0368] The method includes obtaining (3006) a neural network topology (e.g., topology 224) and weights (e.g., weights 222) of a trained neural network (e.g., network 220).

[0316]

[0369] The method also includes converting 3008 (e.g., using neural network conversion module 226) the neural network topology into an equivalent analog network of analog components including a plurality of operational amplifiers and a plurality of resistors, where each operational amplifier represents an analog neuron in the equivalent analog network and each resistor represents a connection between two analog neurons.

[0317]

[0370] Referring now to FIG. 30L, in some implementations, the method further includes pruning the trained neural network. In some implementations, the method further includes pruning the trained neural network using a neural network pruning technique (3052) and updating the neural network topology and weights of the trained neural network before converting the neural network topology, such that an equivalent analog network contains fewer than a predetermined number of analog components. In some implementations, the pruning is performed iteratively (3054), taking into account the accuracy or level of match of the outputs between the trained neural network and the equivalent analog network.

[0318]

[0371] Referring now to FIG. 30M, in some implementations, the method further includes performing network knowledge extraction (3056) before converting the neural network topology to an equivalent analog network. Knowledge extraction is more deterministic than pruning, unlike stochastic / learning methods like pruning. In some implementations, knowledge extraction is performed independently of the pruning step. In some implementations, before converting the neural network topology to an equivalent analog network, connection weights are adjusted according to a predetermined optimization criterion (e.g., preferring zero weights or a specific range of weights over other weights) by deriving causal relationships between hidden neuron inputs and outputs through knowledge extraction methods. Conceptually, for a single neuron or set of neurons, there may be causal relationships between inputs and outputs for a particular data set, which allows for the weights to be retuned in a way that (1) produces the same network output and (2) is easier to implement with resistors (e.g., more evenly distributed values, more zero values, or no connections). For example, if some neuron output is always 1 for some data set, some implementations remove this neuron's output connection (and the entire neuron) and instead adjust the bias weights of neurons that follow it. Thus, the knowledge extraction step differs from pruning because pruning requires re-training after removing a neuron, and because training is probabilistic, whereas knowledge extraction is deterministic.

[0319]

[0372] Referring again to Figure 30A, the method also includes calculating 3010 a weight matrix for an equivalent analog network based on the weights of the trained neural network, with each element of the weight matrix representing a respective connection.

[0320]

[0373] 30J , in some implementations, the method further includes removing or transforming neurons based on the bias values. In some implementations, the method further includes, for each analog neuron of the equivalent analog network, (i) calculating a respective bias value for the respective analog neuron based on the weights of the trained neural network while calculating the weight matrix (3044), (ii) removing the respective analog neuron from the equivalent analog network in accordance with a determination that the respective bias value is above a predetermined maximum bias threshold (3046), and (iii) replacing the respective analog neuron with a linear junction in the equivalent analog network in accordance with a determination that the respective bias value is below a predetermined minimum bias threshold (3048).

[0321]

[0374] 30K, in some implementations, the method further includes minimizing the number of neurons or compacting the network. In some implementations, the method further includes reducing (3050) the number of neurons of the equivalent analog network by increasing the number of connections (inputs and outputs) from one or more analog neurons of the equivalent analog network before generating the weight matrix.

[0322]

[0375] Referring again to Figure 30A, the method also includes generating a resistance matrix of the weight matrix (3012), where each element of the resistance matrix corresponds to a respective weight in the weight matrix.

[0323]

[0376] The method also includes pruning 3014 the equivalent analog network to reduce the number of operational amplifiers or resistors based on the resistance matrix to obtain an optimized analog network of analog components.

[0324]

[0377] 30B, in some implementations, the method includes replacing insignificant resistances with conductors. In some implementations, pruning the equivalent analog network includes replacing (3016) resistors corresponding to one or more elements of the resistance matrix having a resistance value below a predetermined minimum threshold resistance value with conductors.

[0325]

[0378] 30C, in some implementations, the method further includes removing connections with very high resistance. In some implementations, pruning the equivalent analog network includes removing (3018) one or more connections of the equivalent analog network corresponding to one or more elements of the resistance matrix that exceed a predetermined maximum threshold resistance value.

[0326]

[0379] 30D , in some implementations, pruning the equivalent analog network includes removing 3020 one or more connections of the equivalent analog network that correspond to one or more elements of the weight matrix that are approximately zero. In some implementations, pruning the equivalent analog network further includes removing 3022 one or more analog neurons of the equivalent analog network that do not have any input connections.

[0327]

[0380] Referring now to FIG. 30E, in some implementations, the method includes removing non-essential neurons. In some implementations, pruning the equivalent analog network includes (i) ranking (3024) analog neurons of the equivalent analog network based on detecting usage of the analog neurons in performing computations on one or more datasets, e.g., a training dataset used to train the trained neural network, a representative dataset, or a dataset developed for the pruning procedure. Some implementations perform a ranking of neurons for pruning based on the frequency of usage of a given neuron or block of neurons when subjected to the training dataset. For example, (a) when using a test data set, if there is no signal at all for a given neuron, this means that this neuron or block of neurons has never been used and will be pruned; (b) if a neuron is used very infrequently, it can be pruned without significant loss of accuracy; (c) the neuron is always used, in which case it can not be pruned; (ii) selecting one or more analog neurons of the equivalent analog network based on the ranking (3026); and (iii) removing one or more analog neurons from the equivalent analog network (3028).

[0328]

[0381] Referring now to FIG. 30F, in some implementations, detecting the use of analog neurons includes (i) establishing 3030 a model of an equivalent analog network using modeling software (e.g., SPICe or similar software) and (ii) measuring 3032 the propagation of analog signals (currents) using the model (removing blocks where no signal is propagating when using a special training set) to generate one or more datasets of calculations.

[0329]

[0382] Referring now to FIG. 30G, in some implementations, detecting the use of analog neurons includes (i) establishing a model of the equivalent analog network using modeling software (e.g., SPICe or similar software) (3034); and (ii) measuring (3036) the output signal (current or voltage) of the model by generating one or more data set calculations using the model (e.g., signals at the outputs of some blocks or amplifiers in the SPICe model or in the real circuit and eliminating regions where the output signal of the training set is always zero volts).

[0330]

[0383] Referring now to FIG. 30H, in some implementations, detecting the use of analog neurons includes (i) establishing a model of the equivalent analog network using modeling software (e.g., SPICe or similar software) (3038); and (ii) measuring (3040) the power consumed by the analog neurons (e.g., the power consumed by a particular neuron or block of neurons, represented by an operational amplifier either in the SPICE model or in the real circuit, and eliminating the neuron or block of neurons) by generating one or more data set calculations using the model.

[0331]

[0384] Referring now to FIG. 30I, in some implementations, the method further includes recalculating (3042) a weight matrix of the equivalent analog network after pruning the equivalent analog network and before generating one or more lithography masks for fabricating a circuit that implements the equivalent analog network, and updating the resistance matrix based on the recalculated weight matrix.

[0332] Exemplary Analog Neuromorphic Integrated Circuits and Methods of Fabrication Exemplary Method for Fabricating Analog Integrated Circuits for Neural Networks

[0385] 31A-31Q show a flowchart of a method 3100 of fabricating an integrated circuit 3102 including an analog network of analog components, according to some implementations. The method is performed (e.g., using IC fabrication module 258) on a computing device 200 having one or more processors 202 and a memory 214 storing one or more programs configured for execution by the one or more processors 202. The method includes obtaining (3104) a neural network topology and weights for a trained neural network.

[0333]

[0386] The method also includes converting (3106) the neural network topology (e.g., using the neural network conversion module 226) into an equivalent analog network of analog components including a plurality of operational amplifiers and a plurality of resistors (for recurrent neural networks, signal delay lines, multipliers, Tanh analog blocks, and Sigmoid analog blocks are also used), where each operational amplifier represents a respective analog neuron and each resistor represents a respective connection between a respective first analog neuron and a respective second analog neuron.

[0334]

[0387] The method also includes calculating 3108 a weight matrix for an equivalent analog network based on the weights of the trained neural network, each element of the weight matrix representing a respective connection.

[0335]

[0388] The method also includes generating 3110 a resistance matrix of the weight matrix, each element of the resistance matrix corresponding to a respective weight in the weight matrix.

[0336]

[0389] The method also includes generating (3112) one or more lithography masks for fabricating a circuit implementing an equivalent analog network of analog components based on the resistance matrix (e.g., generating masks 250 and / or 252 using mask generation module 248), and fabricating (3114) the circuit (e.g., IC 262) based on the one or more lithography masks using a lithography process.

[0337]

[0390] Referring now to FIG. 31B, in some implementations, the integrated circuit further includes one or more digital-to-analog converters (3116) (e.g., DAC converter 260) configured to generate analog inputs for an equivalent analog network of analog components based on one or more digital signals (e.g., signals from one or more CCD / CMOS image sensors).

[0338]

[0391] Referring now to FIG. 31C, in some implementations, the integrated circuit further includes an analog signal sampling module (3118) configured to process one-dimensional or two-dimensional analog input at a sampling frequency based on the number of inferences of the integrated circuit (the number of inferences of the IC is determined by the product specifications - we know the sampling rate from the neural network operations and the exact task the chip is intended to solve).

[0339]

[0392] Referring now to FIG. 31D, in some implementations, the integrated circuit further includes a voltage conversion module (3120) for scaling down or up the analog signal to fit the operating range of the multiple operational amplifiers.

[0340]

[0393] Referring now to FIG. 31E, in some implementations, the integrated circuit further includes a tact signal processing module (3122) configured to process one or more frames acquired from the CCD camera.

[0341]

[0394] Referring now to FIG. 31F, in some implementations, the trained neural network is a long short-term memory (LSTM) network, and the integrated circuit further includes one or more clock modules for synchronizing signal tacts and enabling time series processing.

[0342]

[0395] Referring now to FIG. 31G, in some implementations, the integrated circuit further includes one or more analog-to-digital converters (3126) (e.g., ADC converter 260) configured to generate a digital signal based on the output of an equivalent analog network of analog components.

[0343]

[0396] Referring now to FIG. 31H, in some implementations, the integrated circuit includes one or more signal processing modules (3128) configured to process one-dimensional or two-dimensional analog signals obtained from the edge application.

[0344]

[0397] Referring now to FIG. 31I, a trained neural network is trained (3130) using a training dataset including signals from an array of gas sensors (e.g., 2-25 sensors) for different gas mixtures for selective sensing of different gases in a gas mixture containing a predetermined amount of the gas to be detected (i.e., using the trained chip's operations to individually determine each known gas in the gas mixture, regardless of the presence of other gases in the mixture). In some implementations, the neural network topology is a one-dimensional deep convolutional neural network (1D-DCNN) designed to detect three binary gas components based on measurements from 16 gas sensors, and includes a 1D convolution block for each of the 16 sensors, three shared or common 1D convolution blocks, and three dense layers (3132). In some implementations, the equivalent analog network includes (i) up to 100 input / output connections per analog neuron, (ii) delay blocks that can introduce delays of any number of time steps, (iii) a signal limit of 5, (iv) 15 layers, (v) approximately 100,000 analog neurons, and (vi) approximately 4,900,000 connections (3134).

[0345]

[0398] Referring now to FIG. 31J, a trained neural network is trained using a training dataset containing thermal aging time series data for different MOSFETs (e.g., the NASA MOSFET dataset containing thermal aging time series for 42 different MOSFETs, where the data is sampled every 400 ms, typically several hours of data for each device) to predict the remaining useful life (RUL) of a MOSFET device (3136). In some implementations, the neural network topology includes four LSTM layers with 64 neurons in each layer, followed by two dense layers with 64 neurons and one neuron, respectively (3138). In some implementations, the equivalent analog network includes (3140) (i) a maximum of 100 input / output connections per analog neuron, (ii) a signal limit of 5, (iii) 18 layers, (iv) 3,000-3,200 analog neurons (e.g., 3137 analog neurons), and (v) 123,000-124,000 connections (e.g., 123,200 connections).

[0346]

[0399] Referring now to FIG. 31K, a trained neural network is trained 3142 to monitor the state of health (SOH) and state of charge (SOC) of lithium-ion batteries for use in a battery management system (BMS) using a training dataset including time-series data including discharge and temperature data during continuous use of different commercially available lithium-ion batteries (e.g., the NASA Battery Usage Dataset. The dataset presents data on the continuous use of six commercially available lithium-ion batteries, and the network operation is based on an analysis of the battery's discharge curves). In some implementations, the neural network topology includes an input layer, two LSTM layers with 64 neurons in each layer, followed by an output dense layer with two neurons for generating SOC and SOH values ​​3144. An equivalent analog network would include (i) a maximum of 100 input / output connections per analog neuron, (ii) a signal limit of 5, (iii) 9 layers, (iv) 1,200–1,300 analog neurons (e.g., 1,271 analog neurons), and (v) 51,000–52,000 connections (e.g., 51,776 connections) (3146).

[0347]

[0400] Referring now to FIG. 31L, a trained neural network is trained 3148 to monitor the state of health (SOH) of lithium-ion batteries for use in a battery management system (BMS) using a training dataset including time-series data including discharge and temperature data during continuous use of different commercially available lithium-ion batteries (e.g., the NASA Battery Usage Dataset. The dataset presents data on continuous use of six commercially available lithium-ion batteries, and the network operation is based on analysis of the battery discharge curves). In some implementations, the neural network topology includes an input layer with 18 neurons, a simple recurrent layer with 100 neurons, and a dense layer with 1 neuron 3150. In some implementations, the equivalent analog network includes (i) a maximum of 100 input / output connections per analog neuron, (ii) a signal limit of 5, (iii) four layers, (iv) 200-300 analog neurons (e.g., 201 analog neurons), and (v) 2,200-2,400 connections (e.g., 2,300 connections) (3152).

[0348]

[0401] Referring now to FIG. 31M, a trained neural network is trained using a training dataset including speech commands (e.g., the Google Speech Commands Dataset) to identify speech commands (e.g., 10 short spoken keywords including "yes," "no," "up," "down," "left," "right," "on," "off," "stop," and "go") (3154). In some implementations, the neural network topology is a depthwise separable convolutional neural network (DS-CNN) layer with one neuron (3156). In some implementations, an equivalent analog network includes (i) up to 100 input / output connections per analog neuron, (ii) a signal limit of 5, (iii) 13 layers, (iv) approximately 72,000 analog neurons, and (v) approximately 2.6 million connections (3158).

[0349]

[0402] Referring now to FIG. 31N, a trained neural network analyzes photoplethysmography (PPG) data, accelerometer data, temperature data, and galvanic skin response signal data, along with baseline heart rate data obtained from an ECG sensor (e.g., PPG data from the PPG-Dalia dataset (CHECK)), for different individuals performing various physical activities over a predetermined period of time. The system was trained (3160) using a training dataset including a PPG sensor and a 3-axis accelerometer (LICENSE). Data was collected from 15 individuals performing various physical activities over a 1-4 hour period. The wrist-based sensor data included PPG, 3-axis accelerometer, temperature and galvanic skin response signals sampled at 4-64 Hz, and baseline heart rate data obtained from an ECG sensor sampled at approximately 2 Hz. The original data was divided into sequences of 1000 time steps (approximately 15 seconds) with a shift of 500 time steps, resulting in a total of 16,541 samples. The dataset was then expanded to include 13,233 training samples and 3,308 test samples for determining pulse during physical activities (e.g., jogging, fitness exercise, stair climbing) based on the PPG sensor data and 3-axis accelerometer data. The neural network topology includes two Conv1D layers, each with 16 filters and 20 kernels, two LSTM layers with 16 neurons each, and two dense layers with 16 neurons and 1 neuron, respectively, for time-series convolution (3162). In some implementations, the equivalent analog network includes (i) delay blocks for generating any number of time steps, (ii) up to 100 input / output connections per analog neuron, (iii) a signal limit of 5, (iv) 16 layers, (v) 700-800 analog neurons (e.g., 713 analog neurons), and (vi) 12,000-12,500 connections (e.g., 12,072 connections) (3164).

[0350]

[0403] Referring now to FIG. 31O, a trained neural network is trained to classify different objects (e.g., humans, cars, cyclists, scooters) based on pulsed Doppler radar signals (3166), and the neural network topology includes a multi-scale LSTM neural network (3168).

[0351]

[0404] Referring now to FIG. 31P, a trained neural network is trained to perform human activity type recognition (e.g., walking, running, sitting, climbing stairs, exercise, activity tracking) based on inertial sensor data (e.g., 3-axis accelerometer, magnetometer, or gyroscope data from a fitness tracking device, smartwatch, or mobile phone, 3-axis accelerometer data sampled at frequencies up to 96 Hz) (3170). The network was trained on three different publicly available datasets, including "opening and then closing the dishwasher," "drinking while standing," "closing the left-hand door," "jogging," "walking," and "climbing stairs." In some implementations, the neural network topology is configured with 12 filters and 64 cursors, each with 12 filters and 64 cursors. The system includes three channel-wise convolutional networks (3172) each having a channel-dimensional convolutional layer, followed by a max pruning layer and two common dense layers of 1024 neurons and N neurons, respectively, where N is the number of classes. In some implementations, the equivalent analog network includes (3174) (i) delay blocks for generating any number of time steps, (ii) up to 100 input / output connections per analog neuron, (iii) an output layer of 10 analog neurons, (iv) a signal limit of 5, (v) 10 layers, (vi) 1,200-1,300 analog neurons (e.g., 1296 analog neurons), and (vi) 20,000-21,000 connections (e.g., 20,022 connections).

[0352]

[0405] Referring now to FIG. 31Q, the trained neural network is further trained to detect abnormal patterns of human activity based on accelerometer data that is merged with heart rate data using a convolution operation (3176) (to detect pre-stroke or pre-heart attack conditions or signals in case of sudden abnormal patterns caused by injury or malfunction due to medical reasons such as epilepsy).

[0353]

[0406] Some implementations include components that are not integrated into the chip (i.e., they are external elements connected to the chip) selected from the group consisting of speech recognition, video signal processing, image sensing, temperature sensing, pressure sensing, radar processing, LIDAR processing, battery management, MOSFET circuit current and voltage, accelerometer, gyroscope, magnetic sensor, heart rate sensor, gas sensor, volume sensor, liquid level sensor, GPS satellite signal, human body conductance sensor, gas flow sensor, concentration sensor, pH meter, and IR visual sensor.

[0354]

[0407] Examples of analog neuromorphic integrated circuits fabricated according to the above-described processes are provided in the following sections according to some implementations.

[0355] An exemplary analog neuromorphic IC for selective gas detection

[0408] In some implementations, a neuromorphic IC is fabricated according to the process described above. The neuromorphic IC is based on a deep convolutional neural network trained for selective sensing of different gases in a gas mixture containing an amount of the gas to be detected. The deep convolutional neural network is trained using a training dataset including signals from an array of gas sensors (e.g., 2-25 sensors) responsive to different gas mixtures. The integrated circuit (or a chip fabricated according to the techniques described herein) can be used to determine one or more known gases in a gas mixture, regardless of the presence of other gases in the mixture.

[0356]

[0409] In some implementations, the trained neural network is a multi-label 1D-DCNN network used for gas mixture classification. In some implementations, the network is designed to detect three binary gas components based on measurements from 16 gas sensors. In some implementations, the 1D-DCNN includes a 1D convolution block per sensor (16 such blocks), three common 1D convolution blocks, and three dense layers. In some implementations, the 1D-DCNN network performance for this task is 96.3%.

[0357]

[0410] In some implementations, the original network is T-transformed with the following parameters: maximum input / output connections per neuron = 100, delay blocks can generate delays by any number of time steps, signal limit of 5.

[0358]

[0411] In some implementations, the resulting T-network has the following properties: 15 layers, approximately 100,000 analog neurons, approximately 4,900,000 connections.

[0359] An exemplary analog neuromorphic IC for MOSFET failure prediction

[0412] MOSFET on-resistance degradation due to thermal stress is a well-known and serious problem in power electronics. In real-world applications, the temperature of MOSFET devices changes frequently over short periods of time. This temperature sweep causes thermal degradation of the device, which can exhibit exponential behavior. This behavior is typically studied by power cycling, which generates a temperature gradient that causes MOSFET degradation.

[0360]

[0413] In some implementations, a neuromorphic IC is fabricated according to the process described above. The neuromorphic IC is based on a network discussed in the paper "Real-time Deep Learning at the Edge for Scalable Reliability Modeling of SI-MOSFET Power Electronics Converters" to predict the remaining useful life (RUL) of MOSFET devices. Using a neural network, the remaining useful life (RUL) of a device can be determined with an accuracy of over 80%.

[0361]

[0414] In some implementations, the network is trained on the NASA MOSFET dataset, which contains thermal aging time series for 42 different MOSFETs. The data is sampled every 400 ms and typically contains several hours of data for each device. The network contains four LSTM layers of 64 neurons each, followed by two dense layers of 64 neurons and 1 neuron each.

[0362]

[0415] In some implementations, the network was T-transformed with the following parameters: maximum input / output connections per neuron was 100, and a signal limit of 5. The resulting T-network had the following properties: 18 layers, approximately 3,000 neurons (e.g., 137 neurons), and approximately 120,000 connections (e.g., 123,200 connections).

[0363] An Exemplary Analog Neuromorphic IC for Lithium-Ion Battery Health and SoC Monitoring

[0416] In some implementations, a neuromorphic IC is fabricated according to the process described above. The neuromorphic IC can be used for predictive analysis of lithium-ion batteries for use in a battery management system (BMS). BMS devices typically offer features such as overcharge and overdischarge protection, state of health (SOH) and state of charge (SOC) monitoring, and load balancing of several cells. SOH and SOC monitoring typically requires a digital data processor, which increases the cost of the device and consumes power. In some implementations, the integrated circuit can be used to obtain accurate SOC and SOH data without implementing a digital data processor on the device. In some implementations, the integrated circuit determines SOC with greater than 99% accuracy and SOH with greater than 98% accuracy.

[0364]

[0417] In some implementations, the network calculations are based on an analysis of the battery's discharge curve and temperature, and / or the data is presented in a time series. Some implementations use data from the NASA Battery Usage Dataset. The dataset presents data on the continuous usage of six commercially available lithium-ion batteries. In some implementations, the network includes an input layer, two LSTM layers of 64 neurons each, and an output dense layer of two neurons (SOC and SOH values).

[0365]

[0418] In some implementations, the network is T-transformed with the following parameters: maximum input / output connections per neuron = 100 and a signal limit of 5. In some implementations, the resulting T-network has the following characteristics: 9 layers, approximately 1,200 neurons (e.g., 1,271 neurons), and approximately 50,000 connections (e.g., 51,776 connections). In some implementations, the network operation is based on analysis of battery discharge curves and temperature. The network is trained using the Network IndRnn, disclosed in the paper "State-of-Health Estimation of Li-ion Batteries in Electric Vehicles Using IndRNN under Variable Load Condition," which was designed to process data from the NASA battery usage dataset. The dataset presents data on the continuous use of six commercially available lithium-ion batteries. The IndRnn network includes an input layer with 18 neurons, a simple recurrent layer of 100 neurons, and a dense layer of 1 neuron.

[0366]

[0419] In some implementations, the IndRnn network is T-transformed with the following parameters: maximum input / output connections per neuron = 100 and a signal limit of 5. In some implementations, the resulting T-network had the following characteristics: 4 layers, approximately 200 neurons (e.g., 201 neurons), and approximately 2,000 connections (e.g., 2,300 connections). Some implementations output only the SOH with an estimation error of 1.3%. In some implementations, the SOC is obtained similarly to how the SOH is obtained.

[0367] An exemplary analog neuromorphic IC for keyword spotting

[0420] In some implementations, a neuromorphic IC is fabricated according to the process described above. The neuromorphic IC can be used for keyword spotting.

[0368]

[0421] The input network is a neural network with 2D convolutional layers and 2D deep convolutional layers, and has an input audio mel-spectrogram of size 49 x 10. In some implementations, the network includes five convolutional layers, four deep convolutional layers, an average pruning layer, and a final dense layer.

[0369]

[0422] In some implementations, the network is pre-trained to recognize 10 short spoken keywords ("yes," "no," "up," "down," "left," "right," "on," "off," "stop," and "go") from the Google Speech Commands Dataset with 94.4% recognition accuracy.

[0370]

[0423] In some implementations, an integrated circuit is fabricated based on a depthwise separable convolutional neural network (DS-CNN) for voice command recognition. In some implementations, the DS-CNN network is T-transformed with the following parameters: maximum input / output connections per neuron = 100 and a signal limit of 5. In some implementations, the resulting T-network had the following characteristics: 13 layers, approximately 72,000 neurons, and approximately 2.6 million connections.

[0371] An example DS-CNN keyword spotting network

[0424] In one example, a keyword spotting network is converted into a T-network according to some implementation aspects. The network is a neural network with 2D convolutional layers and 2D deep convolutional layers, and has an input audio mel-spectrogram of size 49 x 10. The network includes five convolutional layers, four deep convolutional layers, an average pruning layer, and a final dense layer. The network is pre-trained to recognize 10 short spoken keywords ("yes," "no," "up," "down," "left," "right," "on," "off," "stop," and "go") from the Google Speech Commands Dataset (https: / / ai.googleblog.com / 2017 / 08 / launching-speech-commands-dataset.html). There are two additional classes corresponding to "silence" and "unknown." The network output is a softmax of length 12.

[0372]

[0425] The trained neural network (input to the transform) had a recognition accuracy of 94.4% according to some implementations. In the neural network topology, each convolutional layer was followed by a BatchNorm layer and a ReLU layer, with the ReLU activations unbounded and containing approximately 2.5 million multiply-add operations.

[0373]

[0426] After the conversion, the converted analog network was tested on a test set of 1000 samples (100 each of the spoken commands). All test samples were also used as test samples for the original dataset. The original DS-CNN network gave a recognition error of close to 5.7% on this test set. The network was converted into a T-network of trivial neurons. The BatchNormalization layer in "test" mode produces a simple linear signal transformation, which can be interpreted as a weight multiplier + some additional bias. The convolutional, mean pruning, and dense layers are T-transformed fairly straightforwardly. The softmax activation function is not implemented in the T-network, but is applied separately to the T-network output.

[0374]

[0427] The resulting T-network had 12 layers, including the input layer, approximately 72,000 neurons, and approximately 2.5 million connections.

[0375]

[0428] 26A-26K show example histograms 2600 of absolute weights for layers 1-11, respectively, according to some implementations. A weight distribution histogram (of absolute weights) was calculated for each layer. The dashed lines in the chart correspond to the mean absolute weight values ​​for the respective layer. After transformation (i.e., T-transform), the mean output absolute error (computed on the test set) of the transformed network versus the original network is calculated to be 4.1e-9.

[0376]

[0429] Various examples for setting the network limit of the transformed network are described herein according to some implementations. Regarding signal limiting, some implementations use signal limiting for each layer because the ReLU activation used in the network is unbounded. This can potentially affect mathematical equivalence. For this reason, some implementations use a signal limit of 5 for all layers, which corresponds to a power supply voltage of 5 for the input signal range.

[0377]

[0430] To quantize the weights, some implementations use a nominal set of 30 resistors: [0.001, 0.003, 0.01, 0.03, 0.1, 0.324, 0.353, 0.436, 0.508, 0.542, 0.544, 0.596, 0.73, 0.767, 0.914, 0.985, 0.989, 1.043, 1.101, 1.149, 1.157, 1.253, 1.329, 1.432, 1.501, 1.597, 1.896, 2.233, 2.582, 2.844].

[0378]

[0431] Some implementations select the R- and R+ values ​​(see above) for each layer individually. For each layer, some implementations select the values ​​that provide the best weight accuracy. In some implementations, all weights (including biases) in the T-network are then quantized (e.g., set to the closest value that can be achieved with the inputs or selected resistors).

[0379]

[0432] Some implementations transform the output layer as follows: The output layer is a dense layer with no ReLU activation. According to some implementations, the layer is not implemented with a T transform and has a softmax activation reserved for the digital portion. Some implementations do not perform any additional transformations.

[0380] An exemplary analog neuromorphic IC for acquiring heart rate

[0433] PPG is an optically acquired plethysmogram that can be used to detect changes in blood volume in tissue microvascular beds. PPG is often acquired by using a pulse oximeter, which illuminates the skin and measures changes in light absorption. PPG is often processed to determine heart rate in devices such as fitness trackers. Deriving heart rate (HR) from PPG signals is an essential task in edge device computing. PPG data acquired from wrist-based devices typically allows reliable HR acquisition only when the device is stable. When a person is engaged in physical activity, deriving HR from PPG data produces poor results unless combined with inertial sensor data.

[0381]

[0434] In some implementations, an integrated circuit based on a combination of convolutional neural networks and LSTM layers can be used to bias data from a photoplethysmography (PPG) sensor and a three-axis accelerometer to accurately determine pulse rate. The integrated circuit can be used to suppress motion artifacts in the PPG data and determine pulse rate during physical activities such as jogging, fitness exercise, and stair climbing with greater than 90% accuracy.

[0382]

[0435] In some implementations, the input network was trained with PPG data from the PPG-Dalia dataset. Data was collected from 15 individuals performing various physical activities over a set duration (e.g., 1-4 hours each). The training data, which included wrist-based sensor data, included PPG, triaxial accelerometer, temperature, and galvanic skin response signals sampled at 4-64 Hz, as well as baseline heart rate data obtained from an ECG sensor sampled at approximately 2 Hz. The original data was divided into sequences of 1000 time steps (approximately 15 seconds) with a shift of 500 time steps, generating a total of 16,541 samples. The dataset was divided into 13,233 training samples and 3,308 test samples.

[0383]

[0436] In some implementations, the input network included two Conv1D layers with 16 filters each performing time-series convolution, two LSTM layers of 16 neurons each, and two dense layers of 16 and 1 neurons. In some implementations, the network produced an MSE error of less than 6 beats per minute on the test set.

[0384]

[0437] In some implementations, the network is T-transformed with the following parameters: delay blocks can generate delays by any number of time steps, maximum input / output connections per neuron = 100, and a signal limit of 5. In some implementations, the resulting T-network had the following characteristics: 15 layers, approximately 700 neurons (e.g., 713 neurons), and approximately 12,000 connections (e.g., 12,072 connections).

[0385] Processing example PPG data with a T-transformed LSTM network

[0438] As described above, for recurrent neurons, some implementations use a signal delay block added to each recurrent connection of the GRU neuron and the LSTM neuron. In some implementations, the delay block has an external cycle timer (e.g., a digital timer) that activates the delay block at a fixed period dt. This activation generates an output of x(t-dt), where x(t) is the input signal of the delay block. Such an activation frequency may correspond, for example, to the network input signal frequency (e.g., the output frequency of an analog sensor processed by a T-transformed network). Typically, all delay blocks are activated simultaneously with the same activation signal. Some blocks can be activated simultaneously with one frequency and other blocks with another frequency. In some implementations, these frequencies have a common multiplier, and the signals are synchronized. In some implementations, multiple delay blocks are used to generate additive time shifts for one signal. Examples of delay blocks are described above with reference to FIG. 13B, which shows two examples of delay blocks according to some implementations.

[0386]

[0439] According to some implementations, the network for processing the PPG data uses one or more LSTM neurons, an example of which is described above with reference to FIG. 13A, according to some implementations.

[0387]

[0440] The network also uses Conv1D, a convolution performed over the time coordinate, an example implementation of Conv1D is described above with reference to Figures 15A and 15B, according to some implementations.

[0388]

[0441] Details of PPG data are described herein according to several implementation aspects. PPG is an optically acquired plethysmogram that can be used to detect changes in blood volume in the microvascular bed of tissue. PPG is often acquired by using a pulse oximeter, which illuminates the skin and measures changes in light absorption. PPG is often processed to determine heart rate in devices such as fitness trackers. Deriving heart rate (HR) from PPG signals is an essential task in edge device computing.

[0389]

[0442] Some implementations use PPG data from the Capnobase PPG dataset. The data includes raw PPG signals from 42 individuals, each 8 minutes in duration, sampled at 300 samples per second, and baseline heart rate data acquired from an ECG sensor, sampled at 1 sample per second. For training and evaluation, some implementations split the original data into a sequence of 6000 time steps, shifted by 1000 time steps, to obtain a total set of 5838 samples.

[0390]

[0443] In some implementations, a NN-based, input-trained neural network allows for 1-3% accuracy in deriving heart rate...

Claims

1. 1. A method for recognizing human activity, comprising: Tracking a user's activity using one or more sensors, including obtaining a plurality of electrical signals from the one or more sensors; forming a feature vector by extracting a plurality of features from the plurality of electrical signals, the features corresponding to inputs of a neural network model trained to generate a plurality of descriptors for a plurality of predefined human activities; applying an analog neuro-computing hardware device to the feature vector to generate an embedding vector that specifies a descriptor, the analog neuro-computing hardware device implementing the trained neural network model; and applying a trained machine learning classifier to the embedding vector to classify the activity of the user as one of the predefined human activities; A method comprising:

2. The method of claim 1 , wherein the trained neural network model is an autoencoder that includes an encoder and a decoder.

3. The method of claim 1 , wherein the trained machine learning classifier is a KNN (K Nearest Neighbor) classifier.

4. The method of claim 3 , wherein the number of neighbors of the KNN classifier is equal to five.

5. The method of claim 1 , wherein the trained machine learning classifiers are trained separately for each of the predefined human activities using binary classification.

6. The method of claim 1 , wherein the one or more sensors include one or more of an IMU, a camera, a microphone, and a biofeedback device.

7. The method of claim 1 , further comprising smoothing the output of the trained machine learning classifier to obtain underlying classes of activity.

8. the analog neurocomputing hardware device comprises: Obtaining a neural network topology and weights of the trained neural network model; converting the neural network topology into an equivalent analog network of analog components; calculating a weight matrix of the equivalent analog network based on the weights of the trained neural network model, each element of the weight matrix representing one or more connections between analog components of the equivalent analog network; generating a schematic model for implementing the equivalent analog network based on the weighting matrix, including selecting component values ​​for the analog components; fabricating an integrated circuit according to said schematic model using a lithographic process; The method of claim 1 , wherein the ion exchange layer is formed by the steps comprising:

9. 9. The method of claim 8, wherein generating the rough model includes generating a resistance matrix of the weight matrix, each element of the resistance matrix corresponding to a respective weight in the weight matrix and representing a resistance value.

10. 10. The method of claim 1 , wherein the trained machine learning classifier is implemented using one or more digital components, and wherein the trained machine learning classifier can be retrained for new users.

11. 1. A method for recognizing human activity, comprising: acquiring a sequence of electrical signals from one or more sensors that track a user's activity; forming a plurality of feature vectors by extracting features from the sequence of electrical signals, the features corresponding to inputs of a neural network model trained to generate a plurality of descriptors for a plurality of predefined human activities; applying an analog neurocomputing hardware device to the plurality of feature vectors to generate a plurality of embedding vectors, each of which specifies a corresponding descriptor; using the plurality of embedding vectors to classify the activity of the user as one of the predefined human activities; A method comprising:

12. receiving a set of descriptors from the user that describe a particular physical activity; classifying the activity of the user as one of the particular physical activities using the set of descriptors and the plurality of embedding vectors; and The method of claim 11 further comprising:

13. The method of claim 12 , further comprising generating statistics of the user's personal daily routine based on classifying the activity of the user as one of the specific physical activities.

14. storing the plurality of embedding vectors for the user as describing a particular activity; using the plurality of embedding vectors to classify subsequent activities of the user as the particular activity; The method of claim 11 further comprising:

15. receiving a set of descriptors describing a particular activity from a trainer different from the user; providing feedback to the user if the activity matches the particular activity based on the plurality of embedding vectors and the set of descriptors; and The method of claim 11 further comprising:

16. A human activity recognition device, comprising: an integrated circuit for human activity recognition, the integrated circuit including: an analog network of analog components configured to implement a trained neural network model, the trained neural network model being trained to generate a plurality of descriptors for a plurality of predefined human activities based on a plurality of features extracted from a plurality of electrical signals from one or more sensors; one or more digital components configured to classify a human activity as one of the plurality of predefined human activities according to the plurality of descriptors generated by the integrated circuit; A human activity recognition device comprising:

17. The human activity recognition device of claim 16 , further comprising the one or more sensors configured to collect the plurality of electrical signals during the human activity.

18. The human activity recognition device of claim 16 , wherein the trained neural network model is an autoencoder including an encoder and a decoder.

19. 17. The human activity recognition device of claim 16, wherein the one or more digital components implement a trained machine learning classifier that is a KNN (K Nearest Neighbor) classifier that can be retrained.

20. The human activity recognition device of claim 19 , wherein the number of neighbors of the KNN classifier is equal to five.

21. 20. The human activity recognition device of claim 19, wherein the trained machine learning classifier is trained separately for each of the plurality of predefined human activities using binary classification.

22. 20. The human activity recognition device of claim 19, wherein the one or more digital components are further configured to smooth the output of the trained machine learning classifier to obtain fundamental classes of activities.

23. 17. The human activity recognition device of claim 16, wherein the one or more sensors include one or more of an IMU, a camera, a microphone, and a biofeedback device.

24. The integrated circuit comprises: Obtaining a neural network topology and weights of the trained neural network model; converting the neural network topology into an equivalent analog network of analog components; calculating a weight matrix of the equivalent analog network based on the weights of the trained neural network model, each element of the weight matrix representing one or more connections between analog components of the equivalent analog network; generating a schematic model for implementing the equivalent analog network based on the weighting matrix, including selecting component values ​​for the analog components; fabricating said integrated circuit according to said schematic model using a lithographic process; 17. The human activity recognition device of claim 16, made by steps comprising:

25. 25. The human activity recognition device of claim 24, wherein generating the rough model includes generating a resistance matrix of the weight matrix, each element of the resistance matrix corresponding to a respective weight in the weight matrix and representing a resistance value.

Citation Information

Patent Citations

  • Pattern recognizing device

    JP1994052317A

  • Arithmetic circuit and its operation control method

    JP2005122467A

  • Memristor crossbar arrays to activate processors

    US20180364785A1

  • Instruction set for hybrid CPU and analog in-memory artificial intelligence processor

    US20200242459A1

  • Analog hardware realization of neural networks

    WO2021262023A1