Hybrid fixed / flexible neural network architecture

The hybrid neuromorphic analog signal processor addresses the limitations of conventional hardware by combining fixed and flexible parts, achieving efficient, low-power neural network implementations suitable for edge environments with retraining capabilities.

JP2026517759APending Publication Date: 2026-06-02POLYN TECHNOLOGY LIMITED

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
POLYN TECHNOLOGY LIMITED
Filing Date
2024-05-10
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Conventional hardware systems struggle to keep up with the complexity of neural networks, leading to high power consumption, limited computing speed, and high costs due to the inability to reconfigure hardware for retraining, especially in edge environments.

Method used

A hybrid neuromorphic analog signal processor is developed, comprising a fixed part with fixed weights and a flexible part that can be reconfigured, utilizing analog circuits and digital processors to implement neural networks efficiently, allowing for low power consumption and flexibility.

Benefits of technology

The hybrid approach reduces energy consumption and network load by moving initial processing onto the chip, enabling efficient edge computing and retraining capabilities with improved parallelism and resilience to noise and temperature changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026517759000001_ABST
    Figure 2026517759000001_ABST
Patent Text Reader

Abstract

A hybrid analog-digital hardware device and a method for realizing such a device are provided. The hardware device includes an analog circuit comprising multiple operational amplifiers and multiple resistors. The analog circuit is configured to receive analog signals from one or more sensors and to compute an analog output based on the analog signals by executing a portion of a trained neural network. In some implementations, the hardware device includes an analog-to-digital converter coupled to the analog circuit, which is configured to receive the analog output and convert it to a digital input. The hardware device also includes a classifier or regression circuit coupled to the analog circuit. The classifier or regression circuit is configured to receive an output (e.g., a set of embeddings) from the analog circuit and to classify the output according to a machine learning model to obtain a result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications

[0001] This application is a continuation application of PCT application PCT / RU2020 / 000306 filed on June 25, 2020, entitled "Analog Hardware Realization of Neural Networks", a partial continuation application of US Patent Application Publication No. 17 / 189,109 filed on March 1, 2021, entitled "Analog Hardware Realization of Neural Networks", and a continuation application of US Patent Application Publication No. 18 / 196,412 filed on May 11, 2023, entitled "Hybrid Fixed / Flexible Neural Network Architecture", each of which is hereby incorporated by reference in its entirety. US Patent Application Publication No. 17 / 189,109 is also a partial continuation application of PCT application PCT / EP2020 / 067800 filed on June 25, 2020, entitled "Analog Hardware Realization of Neural Networks", which is hereby incorporated by reference in its entirety.

[0002] Technical Field

[0002] The disclosed implementations generally relate to neural networks, and more specifically, to hybrid neural network hardware including an initial fixed part (e.g., analog) and a second flexible part (e.g., digital).

Background Art

[0003] Background

[0003] Conventional hardware has not been able to keep up with the innovations in neural networks and the growing popularity of machine learning-based applications. With the stagnation of progress in digital microprocessors, the complexity of neural networks continues to exceed the computing power of CPUs and GPUs. Neuromorphic processors based on spike neural networks such as Loihi and True North have limitations in their applications. In GPU-like architectures, power and speed are limited by the data transmission rate. Data transmission can consume up to 80% of the chip power and can have a significant impact on computing speed. Edge applications require low power consumption, but there are currently no known high-performance hardware implementations that have the required low power consumption (e.g., consuming less than 50 milliwatts of power).

[0004]

[0004] The neural network training process presents inherent challenges to the hardware implementation of neural networks. A trained neural network is used for specific inference tasks such as classification or regression. Once a neural network is trained, a hardware equivalent is manufactured. When a neural network is retrained, the hardware manufacturing process is repeated, increasing costs. While some reconfigurable hardware solutions exist, such hardware cannot be easily mass-produced and costs significantly more (e.g., five times or more) than non-reconfigurable hardware. Conventional neuromorphic analog signal processors have fixed weights that cannot be adjusted after chip manufacturing. [Overview of the Initiative] [Problems that the invention aims to solve]

[0005] overview

[0005] Therefore, there is a need for methods, circuits and / or interfaces to address at least some of the shortcomings identified above. Analog circuits that model trained neural networks and are manufactured according to the techniques described herein can provide improved performance per watt and may be useful for implementing hardware solutions in edge environments, and can be used for various applications such as drone navigation and autonomous vehicles. The cost advantages provided by these manufacturing methods and / or analog network architectures are even more pronounced with larger neural networks. Analog hardware implementations of neural networks also provide improved parallelism and neuromorphism. Furthermore, neuromorphic analog components are less sensitive to noise and temperature changes compared to their digital counterparts. [Means for solving the problem]

[0006]

[0006] Chips manufactured according to the techniques described herein offer significant improvements over conventional systems in terms of size, output, and performance, making them ideal for edge environments, including for retraining purposes. Such analog neuromorphic chips can be used to implement edge computing applications or in Internet of Things (IoT) environments. Analog hardware allows initial processing that can consume more than 80-90% of the power (e.g., the formation of descriptors for image recognition) to be moved onto the chip, thereby reducing energy consumption and network load in new applications.

[0007]

[0007] According to several implementations, hybrid approaches toward neuromorphic computing are described herein. Similar to the human brain, an artificial neural network may include a fixed part and a flexible part. The flexible part may be modified for a new classification or regression task. According to several implementations, a hybrid neuromorphic analog signal processor combines (i) a fixed part to support fixed weights and (ii) a flexible part responsible for classification or regression. The flexible part can manipulate the output generated by the fixed part. The flexible part may be modified based on updated needs after manufacturing. The flexible part may be implemented as an array of memristors and / or an array of SuperFlash memory having several determined architectures.

[0008]

[0008] In machine learning, after hundreds of training cycles (sometimes called epochs), a deep convolutional neural network typically maintains fixed weights and structure for the first 80-90% of its layers. In subsequent cycles, only the last few layers responsible for classification or regression continue to change their weights. This property is also used in transfer learning. This property can be used to implement the hybrid architecture described herein. A fixed neural network responsible for pattern detection (embedding) is combined with a subsequent flexible algorithm (e.g., a flexible neural network) responsible for pattern interpretation. According to some implementations, the hybrid core includes a fixed neuromorphic analog core configured to generate the embedding. This part has very low power consumption and provides low latency. The hybrid core also includes a flexible part which can be used for the final classification or regression.

[0009]

[0009] In some implementations, the hardware device includes an analog circuit and a classifier or regression circuit. The analog circuit corresponds to a part of a trained neural network. The analog circuit is configured to acquire one or more analog signals from one or more sensors and to compute an analog output based on one or more analog signals. The classifier or regression circuit is coupled to the analog circuit. The classifier or regression circuit is configured to (1) acquire an input signal based on the analog output and (2) apply a machine learning model to the input signal to either (i) classify the input signal according to a plurality of discrete categories or (ii) assign an output on a predetermined continuous scale.

[0010]

[0010] In some implementations, the classifier or regression circuit includes a digital circuit, and the hardware device further includes an analog-to-digital converter (ADC) coupled to the analog circuit. The ADC is configured to receive an analog output and convert it to a digital input.

[0011]

[0011] In some implementations, the analog output includes a set of latent embeddings, and a classifier or regression circuit applies a machine learning model to the latent embeddings.

[0012]

[0012] In some implementations, the analog circuit includes multiple operational amplifiers and multiple resistors. The resistance values ​​of the multiple resistors are based on the weights of neurons in a portion of the trained neural network. The resistors are configured to connect the multiple operational amplifiers. In some implementations, the analog circuit includes sputtered resistors in the back-end of line (BEOL).

[0013]

[0013] In some implementations, the classifier or regression circuit includes one or more digital computing units selected from the group consisting of CPUs, GPUs, RISCs, FPGAs, and ASICs.

[0014]

[0014] In some implementations, the classifier or regression circuit further includes a processor configured to function as a digital controller that provides signals to one or more interfaces and multiplexes power within the hardware device.

[0015]

[0015] In some implementations, the classifier or regression circuit includes a compute-in-memory component and one or more programmable memory tiles.

[0016]

[0016] In some implementations, the classifier or regression circuit includes a network of memristors.

[0017]

[0017] In some implementations, the trained neural network is an autoencoder comprising an encoder portion having multiple hidden layers that compute each representation of each input vector in a lower-dimensional space than the input space of each input vector, and a decoder portion that reconstructs each input vector. The analog circuit corresponds to the encoder portion, and the classifier or regression circuit corresponds to the decoder portion.

[0018]

[0018] In some implementations, the classifier or regression circuit can be reconfigured to train a machine learning model for a new set of inputs different from the set of inputs used to train the trained neural network.

[0019]

[0019] In some implementations, one or more sensors include at least one analog sensor. The analog sensor is a microphone, piezoelectric sensor, PPG sensor, IMU sensor, chemical sensor, lidar sensor, radar sensor, or CMOS matrix sensor.

[0020]

[0020] In some implementations, the analog circuit is configured to generate an embedding that encodes several types of human activity, and the analog signal includes a 3-axis accelerometer signal.

[0021]

[0021] In some implementations, the analog circuit is configured to generate compressed data that encodes vibration sensor data based on vibration characteristics from the vibration sensor, and the analog signal includes a 3-axis accelerometer signal. In some implementations, the vibration sensor is configured to be installed in machinery, automobiles, railway tracks, rail vehicles, wind turbines, or oil and gas pumps, and the analog signal is acquired wirelessly from the vibration sensor.

[0022]

[0022] In some implementations, an analog circuit is configured to generate embeddings that encode a first set of keywords, and a classifier or regression circuit is configured to be retrained for a second set of keywords different from the first set of keywords.

[0023]

[0023] In some implementations, the analog circuit is configured to generate pseudo-labels for unlabeled data for self-supervised representation learning.

[0024]

[0024] In another embodiment, a method for partitioning a neural network is provided. The method comprises obtaining a multilayer neural network having multiple neuron layers. The method also comprises selecting a set of layers of the multilayer neural network. The set of layers includes a first neuron layer and ends with a candidate neuron layer. The method also comprises generating an embedding output by the candidate neuron layer by inputting a set of input vectors into the multilayer neural network. The method also comprises training a classifier or regressor to classify or regress the embedding. The method also comprises evaluating the classifier or regressor using a test set to determine a performance metric for classification or regression. The method also comprises, in accordance with the determination that the performance metric for classification is above a predetermined threshold, repeating the process of selecting a new set of layers based on the set of layers, generating a new embedding using the new set of layers, training a classifier or regressor to classify or regress the new embedding, and evaluating the classifier using a test set until the performance metric falls below a predetermined threshold.

[0025]

[0025] In some implementations, selecting a set of layers and selecting a new set of layers are based on determining whether (i) the number of operations, (ii) the number of neurons, and (iii) the resulting embedding dimension each fall below respective predetermined thresholds.

[0026]

[0026] In some implementations, selecting a set of layers and selecting a new set of layers are based on calculating the energy per operation by simulating a multi-layer neural network.

[0027]

[0027] In some implementations, selecting a set of layers and selecting a new set of layers are based on estimating the energy per operation based on the supply voltage, propagation time, and average current consumption per neuron of a multi-layer neural network.

[0028]

[0028] In some implementations, the method further includes repeating the steps over a predetermined number of iterations.

[0029]

[0029] In some implementations, the method further includes using a new classifier to classify the new embedding after repeating the steps over a predetermined number of iterations.

[0030]

[0030] In some implementations, the plurality of neuron layers includes a first neuron layer for receiving an input, and each neuron layer of the plurality of neuron layers is connected to a subsequent neuron layer of the plurality of neuron layers.

[0031]

[0031] In some implementations, the computer system has one or more processors and memory. One or more programs include instructions for performing any of the methods described herein.

[0032]

[0032] In some implementations, a non-temporary computer-readable storage medium stores one or more programs configured to be executed by a computer system having one or more processors and memory. The one or more programs include instructions for performing any of the methods described herein.

[0033]

[0033] Accordingly, methods, systems and devices used for hardware implementation of neural networks are disclosed.

[0034] Brief explanation of the drawing

[0034] For a better understanding of the aforementioned systems and methods for fixed / flexible hybrid implementations of neural networks, as well as additional systems and methods, please refer to the following descriptions of implementations in conjunction with the following drawings, where throughout the drawings, similar reference numbers refer to corresponding parts. [Brief explanation of the drawing]

[0035] [Figure 1A]

[0035] This is a schematic diagram of a system for realizing a neural network in hardware using hybrid hardware components in several implementation forms. [Figure 1B]

[0036] This is a conceptual block diagram of hybrid hardware for realizing neural networks using several implementation methods. [Figure 1C]

[0037] This is a schematic diagram illustrating exemplary methods for implementing an autoencoder-based classifier using several different implementation configurations. [Figure 1D]

[0038] This document provides schematic diagrams comparing (i) process flows for classification using conventional digital neural network models and (ii) process flows for classification using hybrid hardware based on the techniques described herein, using several implementation configurations. [Figure 2]

[0039] This is a block diagram of computing devices for partitioning a neural network to realize a hybrid hardware implementation of a neural network, using several different implementation configurations. [Figure 3A]

[0040] This provides schematic diagrams of the process for partitioning exemplary keyword spotting neural networks using several implementation methods. [Figure 3B]

[0040] A schematic diagram of the process of partitioning an exemplary keyword spotting neural network is provided for several implementation forms. [Figure 3C]

[0040] A schematic diagram of the process for partitioning exemplary keyword spotting neural networks in several implementation forms is provided. [Figure 4]

[0041] This flowchart shows several implementation methods for splitting a neural network to realize a hybrid hardware architecture. [Modes for carrying out the invention]

[0036]

[0042] Here, an example is shown in the attached drawings. The following description provides many specific details to give a detailed understanding of the present invention. However, it will be apparent to those skilled in the art that the present invention can be carried out without these specific details.

[0037] Description of implementation

[0043] Several implementations realize the neural network in hardware by dividing it into two parts. The first fixed part includes fixed weights and is implemented using analog circuitry. In some implementations, the fixed circuitry is implemented using a neuromorphic analog signal processor. The second flexible part includes programmable weights. In some implementations, the second flexible part is implemented using a digital processor, which may be included as part of the neuromorphic analog signal processor chip or may be an external processor or device. In some implementations, the second flexible part uses an array of memristors and / or an array of SuperFlash memory with some determined architecture.

[0038]

[0044] In this way, the advantages of neuromorphic analog signal processors, such as low latency and high power efficiency, can be combined with the flexibility of the second flexible component.

[0039]

[0045] This specification describes exemplary hardware and techniques for dividing a neural network into two parts, in several implementation forms.

[0040]

[0046] In machine learning, after many training cycles (sometimes called epochs), a deep convolutional neural network model maintains fixed weights and structure for the first 80-90% of its layers. In subsequent training cycles, the weights change only in the last few layers of the neural network (e.g., the layers responsible for classification). This property is also used in transfer learning techniques. This property or feature of neural networks forms the basis of the hybrid hardware described herein. In some implementations, the fixed neural network is responsible for pattern detection (generating a high-density set of latent embeddings or descriptors). This is then combined with a subsequent flexible algorithm. In some implementations, this algorithm includes an additional flexible neural network responsible for pattern interpretation, depending on the nature of the application.

[0041]

[0047] An embedding (also called a latent embedding or descriptor) is a densely packed representation of information about sensory input. Embeddings are formed by neural networks similar to those in the biological nervous system. Embeddings are found in visual neurobiology. For example, the retina in the eye compresses and encodes visual sensory signals from the visual cortex. The visual cortex can then classify and extract meaning for further decision-making. Embeddings are formed in the hidden layers of neural networks. Embeddings contain important information about the input data. Embeddings are used as input data for further efficient processing, such as data classification and interpretation.

[0042]

[0048] Figure 1A is a schematic diagram of system I00 for hardware implementation of neural network 114 using hybrid hardware components in several implementation forms. In this example, the neural network 114 (sometimes called a multilayer neural network) includes two parts or sets of layers 108 and 110. The first set of layers 108 includes neuron layers (circles in Figure 1A), where the neuron layers take input vectors and produce intermediate outputs. The very first layer is sometimes called the input layer, and the other layers (except the output layer) are called the hidden layers of the neural network. The second set of layers 110 receives the output (sometimes called the embedding) of the last layer 109 (the last layer 109 is sometimes called the candidate layer) as input and produces an output in the output layer. In this example, the output layer includes a single neuron 115, but it does not necessarily have to. A layer can have any number of neurons. The neural network includes weights that are trained and associated with the edges that connect the neurons. The techniques described herein may be used to implement a neural network using hybrid hardware that includes a fixed part 102 (e.g., an analog circuit or processor with fixed weights or weights that are unique or cannot be changed after the chip is fabricated) coupled (via an interface 106 such as an analog-to-digital converter) to a flexible part 104 (e.g., a classifier or regression circuit or digital processor with flexible weights). Methods for partitioning (112) the neural network are also described herein, according to some implementations. Partitioning, according to some implementations, involves identifying candidate layers 109 of the neural network 114.

[0043]

[0049] Several implementations include (i) a fixed neuromorphic analog core configured to generate embeddings with ultra-low power and low latency, and (ii) a flexible digital core for final classification or regression. Figure 1B is a block diagram of hybrid hardware 116 for implementing a neural network in several implementations. In some implementations, the fixed neuromorphic analog core includes operational amplifiers 120 representing the nodes of the neural network and resistors 118 representing the connections of the neural network. The values ​​of these analog components may be fixed during manufacturing. For example, resistance values ​​are calculated using a compiler and cannot be changed after chip manufacturing. In some implementations, connections are represented by sputtered resistors at the BEOL of the chip. Exemplary methods for realizing a neural network and / or manufacturing or fabricating such a chip using analog hardware (e.g., using operational amplifiers, resistors, and estimating resistance values) are described in detail in U.S. Patent Application Publication No. 17 / 189,109, filed March 1, 2021, entitled “Analog Hardware Realization of Neural Networks,” which is incorporated herein by reference in its entirety. The hybrid hardware described herein includes one or more interfaces (e.g., an analog-to-digital converter 122) configured to couple analog hardware (e.g., an analog circuit including an operational amplifier 120 and resistors 118 interconnecting the operational amplifiers) with classifier or regression hardware 124 (e.g., a digital processor including a CPU). The analog circuit is configured to implement a set of first layers 108 using analog components. For example, the values ​​of the operational amplifiers and / or resistance values ​​may be determined based on the weight values ​​of the input neural network and its topology. The classifier or regression hardware 124 (sometimes called the classifier or regression circuit) is configured to implement a second set of layers 110 following the candidate layer 109. The regression circuit maps the input signals to continuous digital outputs according to the machine learning model.A regression circuit typically implements a regression algorithm that maps analog outputs to continuous digital values ​​using embeddings. A classifier or regression circuit can perform a classification or regression function and be reconfigured to have different weights for different applications, different instances of an application, different users, and / or use cases.

[0044]

[0050] Multiple machine learning tasks require flexible weights. The fixed portion 102 can be implemented using sputtered resistors on the chip's BEOL. In some implementations, the flexible portion 104 is implemented using a digital microcontroller unit (MCU) coupled to a neuromorphic analog signal processor (fixed portion). In some implementations, the flexible portion is implemented using a RISC V processor, which can be an integral part of the neural analog signal processor. The flexible portion can be a neural network or algorithm, such as k-nearest neighbors (KNN).

[0045]

[0051] A classification or regression task can be considered to have two stages. In the first stage, v = G(x, WG), where x is the input data and x ∈ R N (R is a set of real numbers, N is the number of dimensions of the input data), G is a neural network for building the embedding, WG is a set of trainable parameters for G, and v is v∈R M This is an embedding in (where M is the number of dimensions of the output data). Typically, M is much smaller than N. In the second stage, y = C(v, WC), where C is the classifier, WC is the set of trainable parameters for C, and y is the classification result corresponding to the input vector x. In some cases, C is a discrete classifier, assigning one of a set of categorical values ​​to each input vector x. In other cases, C calculates regression values ​​based on a set of embeddings v and assigns classification results on a continuous scale. A regression calculator can be considered a continuous classifier.

[0046]

[0052] In machine learning, the majority of neural network weights (e.g., 80-90%) typically remain unchanged after several training epochs. Transfer learning can be used to realize such neural networks. For example, a new neural network may be implemented with a fixed part for the majority of the network and a flexible part (e.g., the remaining 10-20%) that can be trained separately. Some implementations combine the fixed part of the neural network (e.g., using resistors) and the flexible part of the network (e.g., implemented using a digital processor in the device's MCU, RISC-V processor, FPGA, or CPU) as part of a neuromorphic analog signal processor chip. In some implementations, the flexible part of the network is implemented using conventional techniques and programming languages ​​used in software development, such as Python, C / C++, and assembly code, as well as specialized frameworks such as TensorFlow or Torch. The implementation then runs on conventional digital computing units such as CPUs, GPUs, RISC processors, or FPGAs, depending on the target device and / or application.

[0047] An exemplary method for dividing a neural network into a fixed part and a flexible part.

[0053] Figure 1C is a schematic diagram of an exemplary method 126 for realizing an autoencoder-based classifier using the techniques described herein, in several implementation forms. The autoencoder constructs an output vector x 136 that closely matches the input vector x 134 after a nonlinear transformation is performed by a hidden layer. The autoencoder consists of an encoder 128 and a decoder 130. The encoder includes one or more hidden layers. The encoder computes a representation 164 of the input vector 134 in a space having fewer dimensions than the input space. The decoder 130 processes the representation 164 and reconstructs the input. The representation 164 is an embedding obtained by projecting the input vector 134 onto the representation space. The embedding 164 may also be further processed by a classifier 132 to obtain a result vector 138 (for example, using discrete or continuous classification or regression).

[0048]

[0054] In some implementations, the first stage involves training an autoencoder. The number of training epochs is determined by the reconstruction error calculated between the output and input vectors. After training the autoencoder, it is used to transform the input vectors into a representation space. In the second stage, a classifier is trained for a specific task. This task space is the same as the space in which the autoencoder was trained. All vectors from this task space are transformed into a representation space by the encoder. The classifier is constructed in this representation space. The classifier processes the vectors after the transformation by the encoder. The number of training epochs for the classifier is determined by the resulting accuracy. In this way, the encoder is trained once and implemented in the fixed part of the system. The classifier is trained for a specific task and implemented in the flexible part of the system.

[0049]

[0055] In some implementations, self-supervised representation learning (SSRL) provides deep feature learning without requiring large annotated datasets. In some implementations, binary classification neural networks are trained to compare pairs of input vectors. If both inputs are the same or are transformations of the same basis vectors, the classifier outputs a first class (e.g., "1"). If the two input vectors are completely different, the classifier outputs a second class (e.g., "0").

[0050]

[0056] Typically, this type of neural network includes two branches (one for each of the two inputs) with shared weights for processing two input vectors. The embeddings obtained at the output layers of these branches are then fed to a classifier for comparison. These parts are trained end-to-end. The number of training epochs is determined by the target binary classification error. The branches that generate the embeddings serve the same role as the encoders in the autoencoder described above. Similar to the exemplary method described above for the autoencoder, the classifier is trained separately for a specific downstream classification task. The number of training epochs for the classifier is determined by the classification or regression accuracy. The branches that generate the embeddings are trained once and implemented in the fixed part of the system. The classifier is trained for a specific task and implemented in the flexible part of the system.

[0051]

[0057] Figure 1D provides a schematic diagram comparing a conventional process flow 168 for classification using a conventional digital neural network model in several implementation forms with a hybrid process flow 166 for classification using hybrid hardware based on the techniques described herein. In the standard flow 168, analog signals 162 from one or more analog sensors 142 are input to an analog-to-digital converter (ADC) 144, which generates digital data 146. The digital data 146 is input to a neural network digital simulation 148, which extracts embeddings 170. The embeddings are classified or analyzed (150) and further processed for algorithm-based decisions 152. Approximately 80% of the computational work is performed in the digital simulation 148.

[0052]

[0058] On the other hand, in the hybrid process flow 166, analog signals 162 from one or more sensors 142 are input to a neuromorphic analog signal processor 154, which may implement the analog circuits described above. The neuromorphic analog signal processor 154 generates an embedding 172 (which may be similar to an embedding 170). The embedding is input to a classification or analysis circuit 156 (sometimes called a classifier), which may be functionally similar to a digital classification module 150. Unlike the standard flow 168, most of the calculations use hybrid hardware 160, which includes the neuromorphic analog signal processor 154 and the classification and analysis circuit 156. The algorithm-based decision module 158 for the hybrid process flow 166 may be similar to the algorithm-based decision module 152 for the standard process flow 168. The neuromorphic analog signal processor 154 has ultra-low power consumption and / or very low latency. The classification and decision algorithms may be digital. These modules can run with significantly reduced resources, power consumption, and / or using ultra-small microcontroller unit (MCU) cores. When neural networks are simulated on digital processors (from CPUs to GPUs and / or Tensor Processing Units (TPUs)), most resources are used for primary data processing to extract embeddings (see extraction process 148 of standard process 168). If the input signals have rich data, primary data processing consumes up to 80% of capacity or computational resources. Classification and decision-making consume fewer resources. Primary data processing is fixed after training. Classification and decision-making are constantly improving and changing as data accumulation and learning progress.

[0053]

[0059] Figure 2 is a block diagram of a computing device 200 for partitioning a neural network to realize a hybrid hardware neural network in several implementation forms. The computing device 200 may include one or more processing units 202 (e.g., CPU, GPU), one or more network interfaces 204, one or more memory units 206, and one or more communication buses 208 for interconnecting these components (e.g., chipset).

[0054]

[0060] Memory 206 includes high-speed random-access memory such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some implementations, memory includes non-volatile memory such as one or more magnetic disk storage devices, one or more optical disk storage devices, one or more flash memory devices, or one or more other non-volatile solid-state storage devices. In some implementations, memory 206 includes one or more storage devices located away from one or more processing units 112. Memory 206, or alternatively, non-volatile memory within memory 206, includes a non-temporary computer-readable storage medium. In some implementations, memory 206, or the non-temporary computer-readable storage medium of memory 206, stores the following programs, modules, and data structures, or subsets or supersets thereof: Operating system 210, which includes procedures for handling various basic system services and procedures for performing hardware-dependent tasks. A network communication module 212 connects the computing device 200 to other computing devices via one or more network interfaces 204 (wired or wireless). • A neural network partitioning module 214 configured to partition a multilayer neural network 215. • An embedding generation module 216 that generates an embedding 220 by inputting a vector 218 into the layers of a multilayer neural network 215, and / or A classifier or regression module 222 constructs a machine learning model for classifying an input vector according to a set of discrete categories that compute values ​​on a continuous scale. The classifier or regression module includes a training module 224 for training the classifier (e.g., a machine learning model) and an evaluation module 226 for evaluating the output of the classifier (e.g., for determining whether the output meets a predetermined performance metric, such as a predetermined accuracy level of 98%).

[0055]

[0061] The operation of the modules and data structures shown and described above will be further explained with reference to Figures 1C, 1D, 3A, 3B, 3C, and 4, according to several implementations, with reference to Figures 2.

[0056] Exemplary neural network applications

[0062] Figures 3A–3C provide schematic diagrams of process 300 for splitting exemplary keyword spotting neural networks in several implementation forms. Keyword spotting networks, such as those shown in Figures 3A–3C, recognize spoken words from a given list of words. Assume the list contains 10 words. The convolutional neural network takes a representation of the spoken word as input and outputs a class for this word (e.g., a matching word from the list). In Figures 3A–3C, the layers of the network are represented by rectangles. Each rectangle contains the name and type of the layer, input parameters, and output parameters. The parameters include the data dimension and the number of filters. In this example, the output layer has 11 outputs; that is, one output for each word from the list, plus an additional ("other") output if the input word does not belong to the list. The line 302 shown in Figure 3C splits this network into a fixed part and a flexible part. The fixed part is before the line and ends at the candidate layer 304. The fixed part is trained once. The flexible part is below this line. The flexible part is trained for specific tasks (e.g., different word lists).

[0057] Exemplary hardware implementation of the flexible part

[0063] The flexible part can be implemented in hardware using different methods. In some implementations, the flexible part is performed by a RISC V processor, which may be an integral part of an analog neuromorphic signal processor. In this case, the flexible part may be a digital controller that provides signals to the interface and multiplexes power signals within the analog neuromorphic signal processor. Some implementations use in-memory computing and / or programmable memory tiles (e.g., flash memory, memristors, or other types of programmable memory). Some implementations use a CPU to run a neural network or classification algorithm for classification. The fixed part is typically computationally limiting, but classification tends to be less resource-intensive than the fixed part.

[0058]

[0064] Some implementations separate the neural network into a fixed part and a flexible part, realizing the fixed part using resistors for weights, fabricating the resistors on the BEOL, and realizing the flexible part with a coupled MCU or RISC-V. According to some implementations, the fixed and flexible parts are integral components of the neuromorphic analog signal processor chip.

[0059]

[0065] Some implementations use transfer learning techniques to divide a neural network into a fixed part and a flexible part. For example, a convolutional neural network is trained for data classification. The convolutional part computes feature representations (embeddings) of the input data. These embeddings are further processed by classifiers specifically trained for other classification or regression tasks from the same data space. The weights of the part of the neural network that generates the embeddings are fixed.

[0060]

[0066] Several implementations train a deep neural network for representing input data features using autoencoders, self-supervised representation learning systems, or generative adversarial networks. The neural network's weights are used to implement the fixed part of the system. Some implementations train the fixed part of the neural network using transfer learning techniques, and then train the flexible part. Some implementations generate embeddings using the network's fixed part. These embeddings are then analyzed by the flexible part using algorithms or neural network-based analysis.

[0061] Examples of applications of human activity recognition

[0067] Some implementations generate embeddings for human activity recognition based on 3-axis accelerometer signals. Some implementations use autoencoder neural networks. The autoencoder encodes several types of human activity as 16-byte strings (embeddings) and then decodes them without loss of accuracy. Some implementations use analyzer neural networks to decode the human activity encoded in the embeddings.

[0062]

[0068] In some implementations, the encoder portion is implemented using fixed neurons of a neuromorphic analog signal processor chip. The flexible portion (analyzer) is implemented using a digital processor. In some implementations, the digital processor is an external CPU or RISC-V processor of the neuromorphic analog signal processor chip. In some implementations, the digital processor is used for input, output, and power management. In some implementations, the embedding is achieved using the fixed analog portion of the neuromorphic analog signal processor chip. In some cases, this accounts for approximately 90% of the total workload. In some implementations, activity recognition is performed using a digital analyzer implemented with a conventional CPU (typically 10% of the workload).

[0063]

[0069] Embeddings generated by encoder neural networks play a crucial role. For example, when a user performs a new physical activity (e.g., riding a bicycle), a unique descriptor is formed. This descriptor is likely to be different from embeddings of other classes. In a multidimensional space (e.g., a 16-dimensional space where each dimension corresponds to a different feature), the embedding is likely to be compact and specific to cycling. If the user marks this activity as cycling, the activity may be recognized as cycling the next time. In some implementations, new classes are encoded even if the class does not exist during the neural network's training.

[0064] Examples of predictive maintenance applications

[0070] In some implementations, neuromorphic analog signal processors based on the techniques described herein are used in predictive maintenance applications such as vibration control. Typically, this involves large amounts of data flowing from vibration sensors installed in machinery, automobiles, railway tracks, railcars, wind turbines, and oil and gas pumps.

[0065]

[0071] The data can be transmitted wirelessly to the analytical instrument. The large data flow shortens the battery life of battery-powered sensors.

[0066]

[0072] Some implementations use the encoder / decoder techniques described above to compress the data flow from the vibration sensor (e.g., to 1 / 1000th of its original size). The resulting embedding is transmitted over long range (LoRa). An advantage of the techniques described herein is that it is possible to create new classes to describe the features of the vibration sensor, even if the network has not been taught to distinguish such types of features.

[0067]

[0073] Some implementations use an encoder network to obtain fixed weights, which are then used to implement the fixed portion of a neuromorphic analog signal processor. In this way, it is possible to obtain entire manifolds of different vibration characteristics from different vibration sensors. These different vibration characteristics can then be analyzed by a digital analyzer of the flexible portion, which recognizes instances of machine malfunction. Both embedded and encoder methods have the advantage of being applicable regardless of the type of sensor signal.

[0068] Examples of keyword spotting applications

[0074] Keyword spotting typically requires the recognition of different word sets (e.g., from different languages). Neuromorphic analog signal processors need to be adaptable. Changing the chip architecture for each new word set is not practical. Therefore, some implementations include a fixed part (performing about 90% of the computation) and a flexible part (performing the remaining 10%). The fixed part distinguishes different words from a specific dataset. For other word sets, the fixed part can generate embeddings. A second flexible network is implemented to distinguish different word sets.

[0069]

[0075] In some implementations, the hardware device includes an analog circuit (e.g., a fixed part 102 which is a circuit including an operational amplifier 120 interconnected using resistor 118) configured to receive one or more analog signals from one or more sensors and compute an analog output based on one or more analog signals by executing part of a neural network. In some implementations, one or more sensors are integrated into the hardware device (not shown). For example, one or more sensors may be connected to resistor 118 (see Figure 1B). In some implementations, one or more sensors include analog sensors such as microphone sensors, piezoelectric sensors, PPG sensors, IMU sensors, chemical sensors, lidar sensors, radar sensors, or CMOS matrix sensors, examples of which are described above with reference to Figure 1D.

[0070]

[0076] The hardware device also includes a classifier or regression circuit (e.g., flexible portion 104 in Figure 1B) coupled to the analog circuit (e.g., via interface 106). The classifier or regression circuit is configured to acquire an input signal based on the analog output and to classify the input signal according to a machine learning model to obtain the result.

[0071]

[0077] In some implementations, the classifier or regression circuit includes a digital circuit, examples of which are described above with reference to Figures 1A and 1B. The hardware device further includes an analog-to-digital converter 122, which is coupled to an analog circuit and configured to receive an analog output and convert it to a digital input. The digital circuit is configured to (i) receive a digital input and (ii) classify the digital input to obtain a result.

[0072]

[0078] In some implementations, the analog output represents an embedding, and a classifier or regression circuit uses the embedding to classify or regress the analog output.

[0073]

[0079] In some implementations, the analog circuit includes multiple operational amplifiers 120 and multiple resistors 118. The resistance values ​​of the resistors are based on some of the weights of a trained neural network. The resistors are configured to connect multiple operational amplifiers. In some implementations, the analog circuit includes sputtered resistors formed on the back-end obline (BEOL).

[0074]

[0080] In some implementations, the classifier or regression circuit includes one or more digital computing units such as a CPU, GPU, RISC processor, FPGA, and ASIC.

[0075]

[0081] In some implementations, the classifier or regression circuit includes a processor further configured to function as a digital controller that provides signals to one or more interfaces and multiplexes power within the hardware device. For example, the CPU in Figure 2 also controls the overall operation of the hardware device.

[0076]

[0082] In some implementations, the classifier or regression circuit includes a compute-in-memory component and one or more programmable memory tiles.

[0077]

[0083] In some implementations, the classifier or regression circuit includes a network of memristors.

[0078]

[0084] In some implementations, the classifier or regression circuit includes a processor configured to run a neural network for data classification or regression. The neural network differs from a pre-trained neural network. For example, the fixed part implements the first set of layers 108 of the neural network 114, while the flexible part implements a classifier different from the neural network 114.

[0079]

[0085] In some implementations, the neural network is an autoencoder, which includes an encoder portion and a decoder portion. The encoder portion performs nonlinear transformations in a hidden layer. The analog circuit corresponds to the encoder portion of the autoencoder and is configured to compute a representation of the input vector in a lower-dimensional space than the input space of the input vector.

[0080]

[0086] In some implementations, a classifier or regression circuit can be reconfigured to train a machine learning model for a new set of inputs different from the set of inputs used to train the neural network.

[0081]

[0087] In some implementations, the analog circuitry is configured to generate embeddings that encode one or more types of human activity. The analog signals include 3-axis accelerometer signals.

[0082]

[0088] In some implementations, the analog circuitry is configured to generate compressed data that encodes vibration sensor data based on vibration characteristics from the vibration sensor. The analog signal includes a 3-axis accelerometer signal. In some implementations, the vibration sensor is configured to be installed in machinery, automobiles, railway tracks, rail vehicles, wind turbines, or oil and gas pumps, and the analog signal is acquired wirelessly from the vibration sensor.

[0083]

[0089] In some implementations, an analog circuit is configured to generate an embedding that encodes a first set of keywords. A classifier or regression circuit is configured to be retrained for a second set of keywords that is different from the first set of keywords.

[0084]

[0090] In some implementations, the analog circuit is configured to generate pseudo-labels for unlabeled data for self-supervised representation learning.

[0085] An exemplary method for dividing a neural network into a fixed part and a flexible part.

[0091] Figure 4 provides flowcharts of a method 400 for partitioning a neural network for hybrid hardware implementation of a neural network, according to several implementations. According to several implementations, the method may be performed by a neural network partitioning module 214 and / or other modules of a computing device 200. The method includes obtaining a neural network 215 having multiple neuron layers (402). The method also includes selecting an initial set 108 of layers for the neural network (404). The initial set of layers includes a first neuron layer and ends with a candidate neuron layer (e.g., the last layer of the initial set 108 of layers) (404). The method also includes generating embeddings 220 output by the candidate neuron layer (e.g., by an embedding generation module 216) (406) by inputting a set of input vectors 218 into the neural network. The method also includes training a classifier or regression model (e.g., by a training module 224) (408) to map the embeddings to output values.

[0086]

[0092] The method also includes evaluating the classifier or regression model using a test set (e.g., by evaluation module 226) (410) to determine the accuracy level and / or performance level. The test set (sometimes called a dataset) contains samples. Each sample is input data. For each sample, there is typically a real-world output, which is considered ground truth. The dataset typically includes a test set (used for validation) and a larger training set. Any common labeled dataset (e.g., CIFAR, COCO, or Imagenet) can be used. A given threshold specifies the target performance metric (e.g., 95% accuracy if determining whether there is a car in an image). One goal is to optimize power efficiency and flexibility. Typically, a larger digital portion provides greater flexibility but lower power efficiency of the system.

[0087]

[0093] The method also includes, if the accuracy level (or performance metric) does not meet a predetermined threshold, repeating (for example, by the neural network partitioning module 214) the following: selecting a new set of initial layers based on the set of layers, generating a new embedding using the new set of initial layers, training a classifier or regression model according to the new embedding, and evaluating the classifier using a test set (412).

[0088]

[0094] In some implementations, the neural network partitioning module 214 reduces the number of analog layers, provided that the classification or regression performance metric exceeds a threshold. The goal is to provide flexibility for classifying inputs to a given domain or application, while maximizing the analog portion and minimizing the flexible portion as much as possible.

[0089]

[0095] In some implementations, selecting an initial set of layers and selecting a new initial set of layers is based on determining whether (i) the number of operations, (ii) the number of neurons, and (iii) the dimensions of the resulting embedding are below predetermined thresholds.

[0090]

[0096] In some implementations, selecting an initial set of layers and a new initial set of layers is based on calculating the energy per operation by simulating the neural network. Commercial software such as Cadence Virtuoso can be used for the simulation. Since the classifier cannot be smaller than a certain predetermined size, the analog or fixed parts must be at least of a certain size. Typically, the smaller the classifier, the smaller the set of different classes it can classify.

[0091]

[0097] In some implementations, selecting an initial set of layers and a new initial set of layers is based on estimating the energy per operation based on the neural network's supply voltage, propagation time, and average current used per neuron.

[0092]

[0098] In some implementations, the method further includes repeating the step a predetermined number of times (e.g., 5 times).

[0093]

[0099] In some implementations, the method further includes using a new classifier to classify new embeddings after repeating the steps over a predetermined number of iterations.

[0094]

[0100] In some implementations, multiple neuron layers include a first neuron layer for receiving input. Each neuron layer among the multiple neuron layers is connected to a subsequent neuron layer among the multiple neuron layers.

[0095] Examples of applications for dividing a neural network into fixed and flexible parts.

[0101] When selecting the fixed part layers, some implementations take into account not only the performance metrics of the classifier or regression model (e.g., accuracy), but also the complexity of the fixed part (e.g., the number of actions and / or neurons) as well as the dimensions of the embedding.

[0096]

[0102] Assume a multilayer neural network has five layers and processes input data samples of length 10. Also assume that the first layer contains 10 neurons, the second layer contains 20 neurons, the third layer contains 15 neurons, the fourth layer contains 10 neurons, and the last layer contains 5 neurons. Furthermore, assume that a particular task has the following requirements: the complexity of the fixed part must be less than 650 operations, the number of neurons must be less than 50, and the embedding dimension must be 20 or less. If four layers of this neural network are selected, they will produce a 10-dimensional (number of neurons in the fourth layer) embedding. The computational complexity of these four layers is defined by the number of operations performed to compute the embedding, i.e., 750 operations resulting from (10 × 10) + (10 × 20) + (20 × 15) + (15 × 10). The number of neurons is the total number of neurons in the selected layers, which is 10 + 20 + 15 + 10, i.e., 55 neurons. These parameters do not meet the requirements (for the number of actions). Therefore, this method removes the fourth layer from the selection. For the remaining three layers, the complexity is (10 × 10) + (10 × 20) + (20 × 15), i.e., 600 actions. The number of neurons is 10 + 20 + 15, i.e., 45, and the embedding dimension is 15. Since these parameters meet the requirements, these three layers are considered as a fixed part. The method involves generating embeddings for the input data using these three layers and training a classifier for these embeddings. We assume that the accuracy of the trained classifier is 98%.

[0097]

[0103] The method then removes the third layer and checks the requirements of the first two layers. Here, the complexity is (10 × 10) + (10 × 20), i.e., 300 operations. The number of neurons is 10 + 20, i.e., 30, and the embedding dimension is 20. Since these parameters also satisfy the requirements, these two layers are considered as the fixed part. The method involves generating embeddings for the input data using these two layers as candidate layers for splitting and training a classifier for these embeddings. Here, we assume the accuracy of the second classifier is 95%. Since the classification accuracy when three layers are selected is higher than when two layers are selected, the method selects three layers (i.e., the first three layers) as the fixed part.

[0098]

[0104] In some implementations, there are several constraints on the configuration of the analog portion. Therefore, some implementations select candidate layers for the analog portion to satisfy these constraints and maximize the classification accuracy of the classifier that classifies the embeddings. These constraints may include energy consumption, the number of neurons in the analog portion, and the embedding dimension. Whether these constraints are met depends on the parameters of the layers selected as the fixed portion (number of actions, number of neurons, and number of neurons in the last layer).

[0099]

[0105] The technical terms used in this description of the invention are intended solely to describe specific implementations and are not intended to limit the invention. Where used in this description and in the appended claims, the singular forms “a,” “an,” and “it” are intended to include the plural form unless the context clearly indicates otherwise. Where used herein, the terms “and / or” will also be understood to refer to and include any possible combination of one or more of the enumerated items relating to the invention. Where used herein, the terms “including” and / or “containing” will also be understood to refer to the presence of the described features, steps, actions, elements, and / or components, but not to exclude the presence or addition of one or more other features, steps, actions, elements, components, and / or groups thereof.

[0100]

[0106] The above description is written with reference to specific implementations for illustrative purposes. However, the above exemplary considerations are not intended to be exhaustive or to limit the invention to the exact form disclosed. In light of the above teachings, many modifications and variations are possible. The implementations have been selected and described to best illustrate the principles of the invention and its practical applications, so that those skilled in the art can best utilize the invention and its various implementations with various modifications suitable for specific intended applications.

Claims

1. An analog circuit that corresponds to a part of a trained neural network, Acquiring one or more analog signals from one or more sensors, The process involves calculating an analog output based on one or more analog signals. An analog circuit configured to perform the following: A classifier or regression circuit coupled to the aforementioned analog circuit, The input signal is acquired based on the aforementioned analog output, Applying a machine learning model to the input signal to either (i) classify the input signal according to a plurality of discrete categories, or (ii) assign an output on a predetermined continuous scale. A classifier or regression circuit configured to perform the following: Hardware devices including...

2. The classifier or regression circuit includes a digital circuit, The hardware device according to claim 1, further comprising an analog-to-digital converter coupled to the analog circuit, the analog-to-digital converter configured to receive the analog output and convert it to a digital input.

3. The hardware device according to claim 1, wherein the analog output includes a set of latent embeddings, and the classifier or regression circuit applies the machine learning model to the latent embeddings.

4. The analog circuit includes a plurality of operational amplifiers and a plurality of resistors, The resistance values ​​of the plurality of resistors are determined based on the weights of the neurons in the portion of the trained neural network, The hardware device according to claim 1, wherein the plurality of resistors are configured to connect the plurality of operational amplifiers.

5. The hardware device according to claim 4, wherein the analog circuit includes a sputtering resistor in the back-end of line (BEOL).

6. The hardware device according to claim 1, wherein the classifier or regression circuit includes one or more digital computing units selected from the group consisting of a CPU, GPU, RISC, FPGA, and ASIC.

7. The hardware device according to claim 1, wherein the classifier or regression circuit further includes a processor configured to function as a digital controller that provides signals to one or more interfaces and multiplexes power within the hardware device.

8. The hardware device according to claim 1, wherein the classifier or regression circuit includes a computing-in-memory component and one or more programmable memory tiles.

9. The hardware device according to claim 1, wherein the classifier or regression circuit includes a network of memristors.

10. The trained neural network is an autoencoder that includes an encoder portion having multiple hidden layers that compute the respective representation of each input vector in a lower-dimensional space than the input space of each input vector, and a decoder portion that reconstructs each of the input vectors. The analog circuit corresponds to the encoder section, The hardware device according to claim 1, wherein the classifier or regression circuit corresponds to the decoder portion.

11. The hardware device according to claim 1, wherein the classifier or regression circuit is reconfigurable to train the machine learning model for a new set of inputs different from the set of inputs used to train the trained neural network.

12. The hardware device according to claim 1, wherein the one or more sensors include an analog sensor selected from the group consisting of a microphone, piezoelectric sensor, PPG sensor, IMU sensor, chemical sensor, lidar sensor, radar sensor, and CMOS matrix sensor.

13. The hardware device according to claim 1, wherein the analog circuit is configured to generate embeddings that encode several types of human activity, and the analog signal includes a three-axis accelerometer signal.

14. The hardware device according to claim 1, wherein the analog circuit is configured to generate compressed data that encodes vibration sensor data based on vibration characteristics from a vibration sensor, and the analog signal includes a three-axis accelerometer signal.

15. The hardware device according to claim 14, wherein the vibration sensor is configured to be installed in machinery, automobiles, railway tracks, railway vehicles, wind turbines, or oil and gas pumps, and the analog signal is acquired wirelessly from the vibration sensor.

16. The hardware device according to claim 1, wherein the analog circuit is configured to generate embeddings that encode a first set of keywords, and the classifier or regression circuit is configured to be retrained for a second set of keywords different from the first set of keywords.

17. The hardware device according to claim 1, wherein the analog circuit is configured to generate pseudo-labels for unlabeled data for self-supervised representation learning.

18. A method for dividing a neural network into a fixed part and a flexible part, Obtaining a neural network with multiple hidden layers, The process involves selecting an initial set of layers for the neural network, wherein the initial set of layers includes a first layer of the neural network and terminates at a candidate layer. For each test vector in the set of input test vectors, an embedding is generated that is output by the candidate layer, Training a regression model to map embeddings to output values, The regression model is evaluated according to the set of input test vectors, and the accuracy level is determined. In accordance with the determination that the accuracy level does not meet a predetermined threshold, the following steps are repeated until the accuracy level meets the predetermined threshold: selecting a new set of initial layers, generating a new embedding using the new set of initial layers, training a new regression model, and evaluating the new regression model using the set of input test vectors. A method that includes this.

19. The method according to claim 18, wherein selecting an initial set of layers and selecting a new set of initial layers is based on determining whether (i) the number of operations, (ii) the number of neurons, and (iii) the dimensions of the resulting embedding are below a predetermined threshold, respectively.

20. The method according to claim 18, wherein selecting the set of initial layers and selecting the new set of initial layers is based on calculating the energy per operation by simulating the neural network.

21. The method according to claim 18, wherein selecting the set of initial layers and selecting the new set of initial layers is based on estimating the energy per operation based on the supply voltage, propagation time and average current used per neuron of the neural network.

22. The method according to claim 18, further comprising repeating the step over a predetermined number of iterations.

23. A method for dividing a neural network into a fixed part and a flexible part, Obtaining a neural network with multiple hidden layers, Selecting a set of candidate hidden layers from the aforementioned neural network, Selecting a set of test input vectors for the aforementioned neural network, For each of the candidate hidden layers, the sum of errors for dividing the neural network at each of the candidate hidden layers is calculated. Specifying each fixed portion of the neural network, including each of the candidate hidden layers and the layers up to each of the candidate hidden layers, Applying the respective fixed parts to each of the aforementioned test input vectors generates each set of test embeddings, Training each classifier using each set of the aforementioned test embeddings, Using the respective trained classifiers and the set of test input vectors, the respective total error for each candidate hidden layer is calculated. Including, Selecting a split layer as a candidate layer with the smallest total error, A method that includes this.