Simulated systems using balanced propagation for learning
By using linear programmable network layer and nonlinear activation layer in the simulation system, combined with equalization propagation technology, the challenge of weight determination in hardware is solved, the target output matching of machine learning is achieved, and the learning efficiency of the simulation architecture is improved.
Patent Information
- Application Number
- CN202080063888.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-03
- Filing Date
- 2020-07-29
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2040-07-29
AI Technical Summary
When performing machine learning in hardware, it is challenging to determine the appropriate weights to match the final output signal to a specific target value, especially with the difficulty of converting mathematical models into the device.
Using a simulation system, combining a linear programmable network layer and a nonlinear activation layer, the gradient is determined and weighted is adjusted using equalization propagation technology, and machine learning is realized through the interlaced linear programmable network layer and a nonlinear activation layer.
The weights are effectively adjusted in the simulation system, and the matching of the target output signal of machine learning and the target value is achieved, improving the efficiency and performance of machine learning in the simulation architecture.
Smart Images

Figure CN114586027B_ABST
Abstract
Description
[0001] Cross-references to other applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 886,800, filed on August 14, 2019, entitled “METHOD AND SYSTEM FOR PERFORMING ANALOG EQUILIBRIUM PROPAGATION,” which is incorporated herein by reference for all purposes. Background Art
[0003] To perform machine learning in hardware, particularly supervised learning, a desired output is achieved from a specific set of input data. For example, input data is provided to the first layer. In this layer, the input data is multiplied by a matrix of values or weights. The output signal for the layer is the result of the matrix multiplication in that layer. The output signal is provided as the input signal to the matrix multiplication of the next layer. This process can be repeated for a large number of layers. The final output signal of the last layer is expected to match a specific set of target values. To perform machine learning, the weights (e.g., resistors) in one or more layers are adjusted so that the final output signal is closer to the target value. Although this process can theoretically change the weights of the layers to provide the target output, in practice, it is challenging to determine an appropriate set of weights. Various mathematical models exist to help determine the weights. However, it may be difficult or impossible to convert such a model into a device. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Various embodiments of the invention are disclosed in the following detailed description and accompanying drawings.
[0005] Figures 1A to 1C is a block diagram depicting an embodiment of a simulation system for performing machine learning.
[0006] Figures 2A to 2B Described are embodiments of a simulation system for performing machine learning.
[0007] Figure 3 is a flow chart depicting an embodiment of a method for performing machine learning.
[0008] Figure 4 is a block diagram depicting an embodiment of a simulation system for performing machine learning using balanced propagation.
[0009] Figure 5 is a diagram depicting an embodiment of a simulation system for performing machine learning using balanced propagation.
[0010] Figure 6 is a block diagram depicting an embodiment of a simulation system for performing machine learning.
[0011] Figure 7 is a diagram depicting an embodiment of a portion of a simulation system for performing machine learning using balanced propagation.
[0012] Figures 8A to 8B is a diagram depicting an embodiment of nanofibers.
[0013] Figure 9 is a diagram depicting an embodiment of a system for performing machine learning using balanced propagation. DETAILED DESCRIPTION
[0014] The present invention may be implemented in a variety of ways, including as: a method; an apparatus; a system; a composition of matter; a computer program product embodied on a computer-readable storage medium; and / or a processor, such as a processor configured to execute instructions stored on and / or provided by a memory coupled to the processor. In this specification, these implementations or any other form that the invention may take may be referred to as techniques. In general, the order of the steps of the disclosed processes may be changed within the scope of the present invention. Unless otherwise specified, a component such as a processor or memory described as being configured to perform a task may be implemented as a general component that is temporarily configured to perform a task at a given time or as a specific component that is manufactured to perform a task. As used herein, the term "processor" refers to one or more devices, circuits, and / or processing cores that are configured to process data, such as computer program instructions.
[0015] A detailed description of one or more embodiments of the present invention is provided below together with the accompanying figures illustrating the principles of the present invention. The present invention is described in relation to such embodiments, but the present invention is not limited to any embodiment. The scope of the present invention is limited only by the claims, and the present invention encompasses many replacements, modifications, and equivalents. In the following description, many specific details are set forth to provide a thorough understanding of the present invention. These details are provided for illustrative purposes, and the present invention can be implemented according to the claims without some or all of these specific details. For the purpose of clarity, technical materials known in the technical field related to the present invention are not described in detail so as not to unnecessarily obscure the present invention.
[0016] In order to perform machine learning in a hardware system, a layer of matrix multiplication is used. In each layer, the input signal (e.g., an input vector) is multiplied by a matrix of values or weights. This matrix multiplication can be performed by a crossbar array, where the weights are the resistors that connect each input to each output at each intersection of the array. The output signal for a layer is the result of the matrix multiplication in that layer. The output signal is provided as an input signal to the next layer (e.g., another crossbar array) that performs another matrix multiplication. This process can be repeated. In order to match the final output signal of the last layer with a set of target values, the weights in one or more layers are adjusted. Although this process can theoretically change the weights of the layer to provide the target output, in practice, it is challenging to determine the weights based on the output. Various mathematical models exist to help determine the weights. However, it may be difficult or impossible to convert such a model into an analog architecture.
[0017] A system for performing learning is described. In some embodiments, the system is an analog system. The system includes a linear programmable network layer and a nonlinear activation layer. The linear programmable network layer includes an input, an output, and a linear programmable network component interconnected between the input and the output. The nonlinear activation layer is coupled to the output. The linear programmable network layer and the nonlinear activation layer are configured to have a static state at a minimum as a generalized function of power dissipation, commonly known as the "content" or "common content" of the system. In some embodiments, multiple programmable network layers are interleaved with one or more nonlinear activation layers. In such embodiments, the nonlinear activation layer is connected to the output of one linear programmable network layer and the input of an adjacent linear programmable network layer.
[0018] In some embodiments, the nonlinear activation layer further comprises a nonlinear activation module and a regeneration module coupled to the output of the linear programmable network layer and the nonlinear activation module. The regeneration module is configured to scale the output signal from the output. In some embodiments, the regeneration module comprises a bidirectional amplifier. In some embodiments, the nonlinear activation module comprises a plurality of diodes.
[0019] The linear programmable network layer may include a programmable resistance network layer. In some embodiments, the programmable resistance network layer includes a fully connected programmable resistance network layer. For example, in some embodiments, a crossbar array with programmable resistors (e.g., memristors) is used. In some embodiments, the programmable resistance network layer includes a sparsely connected programmable resistance network layer. For example, the programmable resistance network layer may include a partially connected crossbar array. In some embodiments, the programmable resistance network layer includes nanofibers and electrodes. In some embodiments, each nanofiber has a conductive core and a memristive layer surrounding at least a portion of the conductive core. A portion of the memristive layer is between the conductive cores of the multiple nanofibers and the multiple electrodes. In some embodiments, each nanofiber has a conductive core and an insulating layer surrounding at least a portion of the conductive core. The insulating layer has holes. In such embodiments, at least a portion of each memristive plug is located in each hole. Therefore, the electrodes can be sparsely connected through the nanofibers.
[0020] A learning system can be used to perform machine learning. To this end, an input signal is provided to the learning system comprising a linear programmable network layer interleaved with (multiple) nonlinear activation layers. The linear programmable network layer and the (multiple) nonlinear activation layers are configured to have a static state at a minimum value of the content of the learning system. The input signal thus causes an output signal corresponding to the static state. The output of the first linear programmable network layer is perturbed. Therefore, the perturbation input signal is provided to the output of the first linear programmable network. In some embodiments, the perturbation input signal corresponds to a second set of output signals that are closer to the target output than the output signal. As a result, a perturbation output signal is generated at the input of the second linear programmable network layer. Based on the perturbation output signal and the output signal, gradient(s) for the linear programmable network components of the second linear programmable network layer are determined. These gradient(s) can be determined using balanced propagation. One or more linear programmable network components in the second linear programmable network layer are adjusted based on the gradient(s). This process can be performed iteratively to adjust the weights in one or more linear programmable network layers to achieve the target output signal from the input signal.
[0021] Figures 1A to 1C 1 is a block diagram illustrating embodiments 100A, 100B, and 100C, respectively, of simulation systems for performing machine learning. Other and / or additional components may be present in some embodiments. Learning systems 100A, 100B, and / or 100C may utilize balanced propagation to perform learning. Although described in the context of a linear programmable network layer, in some embodiments, a nonlinear programmable network (e.g., a network layer including nonlinear programmable components) may be used.
[0022] Reference Figure 1ALearning system 100A includes a linear programmable network layer 110A and a nonlinear activation layer 120A. Linear programmable network layer 110A and nonlinear activation layer 120A include analog circuitry. In some embodiments, multiple linear programmable network layers 110A may be used, interleaved with nonlinear activation layer(s) 120A. For simplicity, a single layer of 110A and 120A is shown. Linear programmable network layer 110A includes linear programmable components. As used herein, a linear programmable component has a linear relationship between voltage and current over at least a portion of its operating range. For example, passive linear components include resistors and memristors. Programmable components have a changeable relationship between voltage and current. For example, a memristor can have a different resistance depending on the current previously driven through the memristor. Thus, linear programmable network layer 110A includes linear programmable components (e.g., programmable resistors and / or memristors) interconnected between input 111A and output 113A. In linear programmable network layer 110A, the linear programmable components can be fully connected (each component is connected to all of its neighbors). In some embodiments, linear programmable components are sparsely connected (not all components are connected to all of their neighbors). Although described in the context of programmable resistors and memristors, in some embodiments, other components with linear impedances may be used in addition to or in place of programmable resistors and / or memristors.
[0023] The nonlinear activation layer 120A can be used to provide activation functionality for the linear programmable network layer 110A. In some embodiments, the nonlinear activation layer 120A can include one or more rectifiers. For example, a plurality of diodes can be used. In other embodiments, other and / or additional nonlinear components capable of providing activation functionality can be used.
[0024] The linear programmable network layer 110A and the nonlinear activation layer 120A are configured such that, in response to an input voltage, the output voltage from the nonlinear activation layer 120A has a static state at a minimum of the contents of the linear programmable network layer 110A and the nonlinear activation layer 120A. In other words, for a specific set of input voltages provided at input 111A, the linear programmable network layer 110 and the nonlinear activation layer 120 settle at a minimum corresponding to a general form of the power dissipated by the system and a function of the input voltage ("content"). In some embodiments, such as some linear networks, the content corresponds to the power dissipated by the network. The content corresponds to a function of the difference in voltage between the corresponding input node 111A and output node 113A (e.g., the square of the voltage difference) and the resistance between nodes 111A and 113A. Thus, the learning system 100A minimizes characteristics of the learning system 100A that depend on the input and output voltages and the impedance of the linear programmable network 110A.
[0025] Figure 1B A learning system 100B is depicted that includes a linear programmable network layer 110B and a nonlinear activation layer 120B. The linear programmable network layer 110B and the nonlinear activation layer 120B include analog circuits and are similar to the linear programmable network layer 110A and the nonlinear activation layer 120A, respectively. In some embodiments, multiple linear programmable network layers 110B can be used, interleaved with (multiple) nonlinear activation layers 120B. For simplicity, single layers of 110B and 120B are shown. The linear programmable network layer 110B includes linear programmable components. In the linear programmable network layer 110B, the linear programmable components can be fully connected or sparsely connected. Although described in the context of programmable resistors and memristors, in some embodiments, other components with linear impedance can be used in addition to or instead of programmable resistors and / or memristors.
[0026] The nonlinear activation layer 120B can be used to provide activation functionality for the linear programmable network layer 110B. In the illustrated embodiment, the nonlinear activation layer 120B includes a nonlinear activation module 122B and a regeneration module 124B. The nonlinear activation module 122B is similar to the nonlinear activation layer 120A. Thus, the nonlinear activation module 122B may include one or more diodes. In other embodiments, other and / or additional nonlinear components capable of providing activation functionality may be used. The regeneration module 124B can be used to account for the reduction in amplitude of the output voltage from the linear programmable network layer 110B across multiple layers. Thus, in some embodiments, the regeneration module 124B is used to scale voltage and current. In some embodiments, the regeneration module 124B is an amplifier, such as a bidirectional amplifier. For example, in some embodiments, the regeneration module 124B scales the voltage in the forward direction (toward the output voltage / next layer) by a gain factor of G and scales the current in the reverse direction (toward the input voltage) by a factor of 1 / G. In embodiments where G = 1, the regeneration module acts as a short circuit. Thus, in some embodiments, the regeneration module 124B may not affect the dynamics of the components of the linear programmable network layer 110B and the nonlinear activation module 122B. Instead, the regeneration module 124B scales the voltage and current in an up / down manner.
[0027] The linear programmable network layer 110B and the nonlinear activation layer 120B are configured such that, in response to an input voltage, the output voltage from the nonlinear activation layer 120B has a static state at a minimum value of the content for the linear programmable network layer 110B and the nonlinear activation layer 120B. In other words, for a specific set of input voltages provided at the input 111B, the linear programmable network layer 110A and the nonlinear activation layer 120B settle at a minimum value of the content corresponding to the input voltage. The content corresponds to the square of the difference in voltage between the corresponding input node 111B and output node 113B, and the resistance between nodes 111B and 113B. Thus, the learning system 100B minimizes characteristics of the learning system 100B that depend on the input and output voltages, as well as the impedance of the linear programmable network 110B.
[0028] Figure 1CA learning system 100C is depicted that includes multiple linear programmable network layers 110C-1, 110C-2, and 110C-3, and nonlinear activation layers 120C-1 and 120C-2. The linear programmable network layers 110C-1, 110C-2, and 110C-3, as well as the nonlinear activation layers 120C-1 and 120C-2, comprise analog circuitry and are similar to the linear programmable network layer(s) 110A / 110B and the nonlinear activation layer(s) 120A / 120B, respectively. Thus, in some embodiments, multiple linear programmable network layers 110B interleaved with the nonlinear activation layer(s) 120B may be used. Each of the linear programmable network layers 110C-1, 110C-2, and 110C-3 includes a linear programmable component. Within the linear programmable network layers 110C-1, 110C-2, and 110C-3, the linear programmable components may be fully connected or sparsely connected. Although described in the context of programmable resistors and memristors, in some embodiments, other components with linear impedance may be used in addition to or instead of programmable resistors and / or memristors.
[0029] Nonlinear activation layers 120C-1 and 120C-2 can be used to provide activation functionality for linear programmable network layers 110C-1, 110C-2, and 110C-3. In the illustrated embodiment, each nonlinear activation layer 120C-1 and 120C-2 includes a nonlinear activation module 122C-1 and 122C-2 and a regeneration module 124C-1 and 124C-2. Nonlinear activation modules 122C-1 and 122C-2 are similar to nonlinear activation module 122B and nonlinear activation layer 120A. Regeneration modules 124C-1 and 124C-2 are similar to regeneration module 124B. Therefore, in some embodiments, regeneration modules 124C-1 and 124C-2 are used to scale voltage and current. In some embodiments, regeneration modules 124C-1 and 124C-2 are similar to regeneration module 124B and, therefore, can include bidirectional amplifiers.
[0030] The linear programmable network layers 110C-1, 110C-2, and 110C-3 and the nonlinear activation layers 120C-1 and 120C-2 are configured such that, in response to input voltages, the output voltages at nodes 132-1, 134-1, 136-1, 132-2, 134-2, 136-2, 132-3, 134-3, and 136-3 have a static state at a minimum value of the content for the linear programmable network layers 110C-1, 110C-2, and 110C-3 and the nonlinear activation layers 120C-1 and 120C-2. In other words, for a specific set of input voltages provided to the linear programmable network layer 110C-1, the learning system 100C stabilizes at a minimum value of the content corresponding to the input voltage. The content corresponds to the square of the difference between the voltages at the corresponding input 111A and output node 113B. Therefore, the learning system 100C minimizes characteristics of the learning system 100C that depend on the input and output voltages and the impedance of the linear programmable network 110C.
[0031] Learning systems 100A, 100B, and 100C can utilize balanced propagation to perform machine learning. Balanced propagation stipulates that for certain functions (referred to herein as "energy functions"), the gradients of network parameters can be derived from the values of specific parameters at nodes. Although referred to as energy functions, balanced propagation does not indicate that the "energy function" corresponds to specific physical properties of the simulated system. It has been determined that balanced propagation can be performed for simulated systems (e.g., impedance networks) that have "energy functions" that correspond to content. More specifically, pseudo power can be used for balanced propagation. In some embodiments, the pseudo power corresponds to the content. The pseudo power, and therefore the content, is minimized by the system.
[0032] The pseudo power can be given by:
[0033] (1)
[0034] where g ij is the conductance of the resistor connecting node i to node j, v i is the voltage at node i, and v j is the voltage at node j. In some embodiments, the pseudo power can be given by a function similar to equation (1). For a two-terminal component, the pseudo power is 0.5 times the square of the voltage drop across the component divided by the resistance. In other words, the pseudo power is half the power dissipated by the two-terminal component. Given a fixed boundary node voltage v i, the internal node voltages are stabilized in a configuration that minimizes the above "energy" function (e.g., content or pseudo power). For the purposes of machine learning, the factor 1 / 2 can be ignored. Therefore, in some embodiments, minimizing the content corresponds to minimizing the power dissipated by the network. Because the content is naturally minimized in a stable state, the learning systems 100A, 100B, and 100C allow balanced propagation to be utilized for various purposes. For example, the learning systems 100A, 100B, and 100C allow balanced propagation to be used in performing machine learning. Therefore, the learning systems 100A, 100B, and 100C allow the weights (impedances) of the linear programmable networks 110A, 110B, 110C-1, 110C-2, and 110C-3 to be determined using balanced propagation in combination with the input signal provided to the learning system, the output signal obtained for the linear programmable network layer, the perturbation input signal provided to the output of the learning system, and the perturbation output signal obtained at the input to the linear programmable network.
[0035] For example, the learning system 100C can be used to perform machine learning. An input signal (e.g., an input voltage) is provided to a linear programmable network layer 110C-1. As a result, a first set of output signals is generated at nodes 132-1, 134-1, and 136-1. This first set of output signals is input to a linear programmable network layer 110C-2 and results in a second set of output signals at nodes 132-2, 134-2, and 136-2. The second set of output signals is input to a linear programmable network layer 110C-3 and results in a set of final output signals at nodes 132-3, 134-3, and 136-3. The output signals at nodes 132-1, 134-1, 136-1, 132-2, 134-2, 136-2, 132-3, 134-3, and 136-3 correspond to minimum values in the content of the learning system 100C (e.g., by the linear programmable network layers 110C-1, 110C-2, and 110C-3). Output nodes 132-3, 134-3, and 136-3 are perturbed. For example, a perturbation input signal (e.g., a perturbation input voltage) is provided to outputs 132-3, 134-3, and 136-3. In other words, outputs 132-3, 134-3, and 136-3 are clamped at perturbation voltages. These perturbation voltages can be selected to be closer to desired target voltages for outputs 132-3, 134-3, and 136-3. These perturbation signals propagate backward through the learning system 100C and result in a first set of perturbation output signals (voltages) at nodes 132-2, 134-2, and 136-2. This first set of perturbation output voltages is provided to the output of the linear programmable network layer 110C-2. These perturbation signals propagate backward through the learning system 100C and result in a second set of perturbation output signals (voltages) at nodes 132-1, 134-1, and 136-1. These perturbation signals propagate backward through the learning system 100C and result in a final set of perturbation output signals (voltages) at the input of the linear programmable network layer 110C-1. The perturbation output signals (voltages) at nodes 132-1, 134-1, 136-1, 132-2, 134-2, 136-2, 132-3, 134-3, 136-3 correspond to the minimum value among the contents of the perturbation voltages provided by the learning system 100C (e.g., via the linear programmable network layers 110C1, 110C-2, and 110C-3) at the output nodes 132-3, 134-3, and 136-3.
[0036] Using the output signals on nodes 132-1, 134-1, 136-1, 132-2, 134-2, 136-2, 132-3, 134-3, 136-3 in combination with balanced propagation for the input voltages and using the perturbed output signals on nodes 132-1, 134-1, 136-1, 132-2, 134-2, 136-2, 132-3, 134-3, 136-3 for the perturbed input signals, gradients for weights (e.g., impedances) for the linear programmable networks 110C-1, 110C-2, and 110C-3 may be determined.
[0037] For example, an input voltage X representing input data is provided to an input node (boundary voltage) of the learning system 100C, and internal nodes 132-1, 134-1, 136-1, 132-2, 134-2, 136-2, 132-3, 134-3, and 136-3 (including output nodes 132-3, 134-3, and 136-3) are stabilized to the minimum value of the energy function (e.g., content). These output signals (node voltages) are represented as S free The output nodes 132-3, 134-3, and 136-3 of the learning system 100C are then "pushed" in the direction of a set of target signals (e.g., target voltages) Y. In other words, the perturbation voltages are provided to the output nodes 132-3, 134-3, and 136-3. For example, Y can be a true label in a classification task. The learning system 100C settles to a new energy minimum, and the new "weakly clamped" node voltage (e.g., the perturbed output voltage) is denoted as S clamped .
[0038] According to the balanced propagation, the gradient of the network parameters (the impedance value of each linear programmable component in each linear programmable network layer) with respect to the error or loss function L can be directly obtained from S free and S clamping This gradient can then be used to modify the conductance (and hence the impedance) of the linear programmable component (e.g., changing g ij For example, the memristor can be programmed by driving an appropriate current through it. Thus, learning system 100C can be trained to optimize a well-defined objective function L. In other words, learning system 100C can utilize balanced propagation to perform machine learning. For similar reasons, learning systems 100A and 100B can also utilize balanced propagation to perform machine learning.
[0039] Thus, balanced propagation can be used to determine the modification of the weights (impedances) in the learning systems 100A, 100B, and / or 100C. As indicated in the learning system 100C, the modification of the impedance can be determined for a learning system comprising multiple layers. Thus, balanced propagation can be used to perform backpropagation and train the learning systems 100A, 100B, and / or 100C using the output signals for the learning systems 100A, 100B, and 100C. Adding nonlinear activation layers 120A, 120B, 120C-1, and 120C-2 also allows for more complex separation of data by the learning systems 100A, 100B, and 100C. Thus, machine learning can be better able to be performed using analog architectures that can be easily implemented.
[0040] Figures 2A to 2B Depicts embodiments of learning systems 200A and 200B that can perform machine learning using balanced propagation. Although described in the context of a linear programmable network layer, in some embodiments, a nonlinear programmable network (e.g., a network layer including nonlinear programmable components) can be used. Figure 2A , a learning system 200A similar to the learning system 100A is shown. The learning system 200A includes a linear programmable network layer 210 and a nonlinear activation layer 220A, which are similar to the linear programmable network layer 110A and the nonlinear activation layer 120A, respectively. In some embodiments, multiple linear programmable network layers 210 interleaved with (multiple) nonlinear activation layers 220A can be used. Further, the learning system 200A can be replicated in parallel to provide a first layer in a more complex learning system. In some embodiments, such a first layer can be replicated in series, where the output of one layer is the input to the next layer. In other embodiments, the linear programmable network layers and / or the nonlinear activation layers need not be identical. For simplicity, a single layer of 210 and 220A is shown. Also explicitly shown is a voltage input 202, which provides an input voltage (e.g., an input signal) to the input of the linear programmable network layer 210.
[0041] Linear programmable network layer 210 includes linear programmable components. More specifically, linear programmable network layer 210 includes programmable resistors 212, 214, and 216. In some embodiments, programmable resistors 212, 214, and 216 are memristors. However, in some embodiments, other and / or additional programmable passive components may be used.
[0042] The nonlinear activation layer 220A can be used to provide an activation function for the linear programmable network layer 210. Thus, the activation layer 220A comprises a two-terminal circuit element whose IV curve is weakly monotonic. In the illustrated embodiment, the nonlinear activation layer 220A comprises diodes 222 and 224. Thus, the diodes 222 and 224 are used to create an S-shaped nonlinearity as an activation function. In other embodiments, more complex layers with additional resistors and / or different resistor arrangements including more nodes and multiple activation functions can be used.
[0043] In operation, an input signal (e.g., an input voltage) is provided to the linear programmable network layer 210. As a result, an output signal is generated at node 230A. The output signal propagates to subsequent layers (not shown). The output signal at the final node and at the internal node (e.g., at node 230A) corresponds to the minimum value on the energy dissipated by the learning system 200A for the input voltage. The output (e.g., node 230A or a subsequent output) is perturbed. (Multiple) perturbation voltages are selected to be closer to the desired target voltage for the output. These perturbation signals propagate backward through the learning system 200C and cause a perturbation output signal (voltage) at node 202. The perturbation output signal (voltage) at the input 202 and any internal node (e.g., node 230A for a multi-layer learning system) corresponds to the minimum value on the energy dissipated by the learning system 200A for the perturbation voltage.
[0044] By using the output voltages at the internal and external nodes for the input voltages and the perturbation output signals at the input 202 and internal nodes for the perturbation input signals in combination with balanced propagation, it is possible to determine the gradient of the weight (e.g., impedance) for each programmable resistor (e.g., programmable resistors 212, 214, and 216) in each linear programmable network layer (e.g., layer 210). Therefore, machine learning can be performed more easily in the learning system 200A. Consequently, the performance of the analog learning system 200A can be improved.
[0045] Figure 2BA circuit diagram depicts an embodiment of a learning system 200B that can perform machine learning using balanced propagation. Learning system 200B is similar to learning system 100B and learning system 200A. Learning system 200B includes a linear programmable network layer 210 and a nonlinear activation layer 220A, similar to linear programmable network layers 110A / 210 and nonlinear activation layers 120A / 220A, respectively. In some embodiments, multiple linear programmable network layers 210 can be used, interleaved with nonlinear activation layer(s) 220B. Further, learning system 200B can be replicated in parallel to provide a first layer in a more complex learning system. In some embodiments, such a first layer can be replicated in series, where the output of one layer is the input for the next layer. Therefore, an additional linear programmable network layer 240 is shown with resistors 242, 244, and 246. In other embodiments, each linear programmable network layer need not be identical. Voltage input 202 is also explicitly shown, providing an input voltage (e.g., an input signal) to the input of linear programmable network layer 210B.
[0046] Linear programmable network layer 210B includes programmable resistors 212, 214, and 216. In some embodiments, programmable resistors 212, 214, and 216 are memristors. However, in some embodiments, other and / or additional programmable passive components may be used.
[0047] Non-linear activation layer 220B includes a non-linear activation module 221 and a regeneration module 226. Non-linear activation module 221 can be used to provide activation functionality for linear programmable network layer 210 and is similar to non-linear activation layer 220A. Non-linear activation module 211 thus includes diodes 222 and 224. In other embodiments, more complex layers with additional resistors and / or different resistor arrangements including more nodes and multiple activation functions can be used.
[0048] The input voltage to input 202 can have a mean of zero and a unit standard deviation. In practice, the output voltage of the resistor layer has a significantly smaller standard deviation than the input (the voltage will be closer to zero). To prevent signal attenuation, the output voltage can be amplified at each layer. Therefore, a regeneration module 226 is used. The regeneration module 226 is a feedback amplifier that can act as a buffer between the input and output. However, the backward effect is used to propagate gradient information. Therefore, in the illustrated embodiment, the regeneration module 226 is a bidirectional amplifier. The voltage in the forward direction is amplified by a gain factor of A. The current in the backward direction is amplified by a gain factor of 1 / A. If the gain is set to A=1, then the amplifier 226 behaves as a short circuit and does not affect the solution to the minimization problem. If the gain is set to a higher value, such as A=4, the dynamics of the output are simply rescaled by a factor of approximately four. In this scheme, the voltage variable carries the input information forward to perform inference, and the current variable carries the gradient information backward.
[0049] The voltage amplification performed by amplifier 226 in the forward direction can be performed by a voltage-controlled voltage source (VCVS). The current amplification in the backward direction can be performed by a current-controlled current source (CCCS). The control current entering the CCCS is given by the current sourced from the VCVS. The CCCS reflects this current backward and reduces it by the same factor as the forward gain. In this way, the injected current at the output node can be propagated backward, carrying the correct gradient information. In some embodiments, a more complex layer 220B with additional resistors and / or different resistor arrangements including more nodes and multiple activation functions can be used.
[0050] In operation, an input signal (e.g., an input voltage) is provided to the linear programmable network layer 210. As a result, an output signal is generated at node 230B. The output signal propagates to subsequent layers (e.g., the linear programmable network layer 240). The output signal at the final node and at the internal nodes (e.g., at node 230B) corresponds to the minimum value on the energy dissipated by the learning system 200B for the input voltage. The output (e.g., the subsequent output) is perturbed. (Multiple) perturbation voltages are selected to be closer to the desired target voltage for the output. These perturbation signals propagate backward through the learning system 200B and cause perturbation output signals (voltages) at node 202 and other internal nodes. The perturbation output signal (voltage) at the input 202 and any internal node (e.g., node 230B for a multi-layer learning system) corresponds to the minimum value on the energy dissipated by the learning system 200B for the perturbation voltage.
[0051] By using the output voltages at the internal and external nodes for the input voltages and the perturbation output signals at the input 202 and internal nodes for the perturbation input signals in combination with balanced propagation, it is possible to determine the gradient of the weight (e.g., impedance) for each programmable resistor (e.g., programmable resistors 212, 214, 216, 242, 244, and 246) in each linear programmable network layer (e.g., layers 210 and 240). Therefore, machine learning can be performed more easily in the learning system 200B. Consequently, the performance of the analog learning system 200B can be improved.
[0052] Figure 3 is a flow chart depicting an embodiment of a method 300 for performing machine learning using balanced propagation. For clarity, only some steps are shown. In some embodiments, other and / or additional processes may be performed. Although described in the context of a linear programmable network layer, in some embodiments, method 300 may be utilized in the context of a nonlinear programmable network (e.g., a network layer including nonlinear programmable components).
[0053] At 302, an input signal (e.g., an input voltage) is provided to the input of a first linear programmable network layer. The input signal propagates rapidly through the learning system. As a result, output signals are generated at internal nodes (e.g., the output of each linear programmable network layer) and at the output of the learning system (e.g., the output of the last linear programmable network layer). The output signals at the final node and at the internal nodes correspond to a static state with a minimum value for the energy dissipated by the learning system for the input voltage provided in 302. At 304, these output signals for the internal and external nodes are determined.
[0054] At 306, the output is disturbed. In some embodiments, a disturbance signal (e.g., voltage) is applied to the output of the last linear programmable network layer. In some embodiments, the disturbance signal is applied at one or more internal nodes. The (multiple) disturbance signals provided at 306 are selected to be closer to the desired target voltage for the output. These disturbance signals are propagated backward through the learning system and cause disturbance output signals (voltages) on the input and other internal nodes. The disturbance output signals (voltages) on the input 202 and any internal node (e.g., node 230B for a multi-layer learning system) correspond to the minimum value on the energy dissipated by the learning system 200B for the disturbance voltage. At 308, these disturbance output signals are determined.
[0055] At 310, gradients for weights (e.g., impedances) for each linear programmable component in each linear programmable network layer are determined using the output voltages at the internal and external nodes for the input voltages and the perturbed output signals at the input and internal nodes for the perturbed input signals, in combination with balanced propagation. At 312, the linear programmable components are reprogrammed based on the determined gradients. For example, at 312, the impedances of the linear programmable components (e.g., memristors) may be changed. At 314, 302, 304, 306, 308, 310, and 312 may be repeatedly iterated until the appropriate weights for the target outputs are obtained. Thus, machine learning may be more easily performed in an analog learning system. Thus, method 300 may be used to improve the performance of an analog learning system.
[0056] Figure 4 An embodiment of a learning system 400 that can perform machine learning using balanced propagation is depicted. The learning system 400 includes a fully connected linear programmable network layer 410 and a nonlinear activation layer 420, similar to the linear programmable network layer 110B and the nonlinear activation layer 120B, respectively. In some embodiments, multiple linear programmable network layers 410 can be used, interleaved with (multiple) nonlinear activation layers 420. Although described in the context of linear programmable network layers, in some embodiments, nonlinear programmable networks (e.g., network layers that include nonlinear programmable components) can be used.
[0057] The linear programmable network layer 410 includes fully connected linear programmable components. Therefore, each linear programmable component is connected to all of its neighbors. For example, Figure 5 Depicts a crossbar array 500 that can be used for a fully connected linear programmable network layer 410. The crossbar array 500 includes horizontal lines 510-1 to 510-(n+1), vertical lines 530-1 to 530-m, and programmable conductances 520-11 to 520-nm. In some embodiments, the programmable conductances 520-11 to 520-nm are memristors. In some embodiments, the programmable conductances 520-11 to 520-nm can be memristive fibers arranged in the crossbar array. As in Figure 5 As can be seen in FIG, each horizontal line 510-1 to 510-(n+1) is connected to each vertical line 530-1 to 530-m at each intersection through programmable conductance 520-11 to 520-nm. Thus, crossbar array 500 is a fully connected network that can be used for programmable network layer 410.
[0058] Return to reference Figure 4, a nonlinear activation layer 420 can be used to provide an activation function for the linear programmable network layer 410. The activation layer 420 includes (multiple) nonlinear activation modules 422 and (multiple) linear regeneration modules 424. The nonlinear activation modules can include one or more activation modules, such as module 221. Similarly, the linear regeneration modules 424 can include one or more (multiple) regeneration modules 226. The learning system 400 functions in a similar manner to the learning systems 100A, 100B, 100C, 200A, and 200B and can utilize the method 300. Therefore, the performance of the simulated learning system 400 can be improved.
[0059] Figure 6 An embodiment of a learning system 600 that can perform machine learning using balanced propagation is depicted. The learning system 600 includes a sparsely connected linear programmable network layer 610 and a nonlinear activation layer 620, similar to a linear programmable network layer 610B and a nonlinear activation layer 620B, respectively. In some embodiments, multiple linear programmable network layers 610 can be used, interleaved with (multiple) nonlinear activation layers 620. Although described in the context of linear programmable network layers, in some embodiments, nonlinear programmable networks (e.g., network layers including nonlinear programmable components) can be used.
[0060] The linear programmable network layer 610 includes sparsely connected linear programmable components. Therefore, not every linear programmable component is connected to all of its neighbors. A nonlinear activation layer 620 can be used to provide an activation function for the linear programmable network layer 610. The activation layer 620 includes (multiple) nonlinear activation modules 622 and (multiple) linear regeneration modules 624. The nonlinear activation modules can include one or more activation modules, such as module 221. Similarly, the linear regeneration modules 624 can include one or more (multiple) regeneration modules 626. The learning system 600 functions in a similar manner to the learning systems 100A, 100B, 100C, 200A, and 200B and can utilize the method 300. Therefore, the performance of the simulated learning system 600 can be improved.
[0061] Figure 7 A plan view depicting an embodiment of a sparsely connected network 710 that can be used in the linear programmable network layer 610. The network 710 includes nanofibers 720 and electrodes 730. The electrodes 730 are sparsely connected by the nanofibers 720. The nanofibers 720 can be arranged on the underlying layer. The nanofibers 720 can be covered in an insulator, and the electrodes 730 are provided in through-holes in the insulator.
[0062] Figure 8A and Figure 8BDepicts an embodiment in which nanofibers 800A and 800B may be used as nanofibers 720. In some embodiments, only nanofiber 800A is used. In some embodiments, only nanofiber 800B is used. In some embodiments, nanofibers 800A and 800B are used. Figure 8A Nanofiber 800A includes a core 812 and a memristive layer 814A. Other and / or additional layers may be present. Figure 8A Electrode 830 is also shown. In some embodiments, the diameter of the conductive core 812 may be no greater than the nanometer range. In some embodiments, the diameter of the core 812 is on the order of tens of nanometers. In some embodiments, the diameter of the core 812 may be no more than ten nanometers. In some embodiments, the diameter of the core 812 is at least one nanometer. In some embodiments, the diameter is at least ten nanometers and less than one micron. However, in other embodiments, the diameter of the core 812 may be larger. For example, in some embodiments, the diameter of the core 812 may be 1 to 2 microns or larger. In some embodiments, the length of the nanofiber 810 along the axis is at least one thousand times the diameter of the core 812. In other embodiments, the length of the nanofiber 810 may not be limited based on the diameter of the conductive core 812. In some embodiments, the cross-sections of the nanofiber 810 and the conductive core 812 are not circular. In some such embodiments, the transverse dimension(s) of the core 812 are the same as the diameter described above.
[0063] The conductive core 812 can be monolithic (including a single continuous piece) or can have multiple components. For example, the conductive core 812 can include a plurality of conductive fibers (not shown separately) that can be woven or otherwise connected together. The conductive core 812 can be a metallic element or alloy and / or other conductive material. In some embodiments, for example, the conductive core 812 can include at least one of the following: Cu, Al, Ag, Pt, other precious metals and / or other materials that can be formed into a nanofiber core. For example, in some embodiments, the conductive core 812 can include or be composed of one or more conductive polymers (e.g., PEDOT:PSS, polyaniline) and / or one or more conductive ceramics (e.g., indium tin oxide / ITO).
[0064] In some embodiments, the memristive layer 814A surrounds the core 812 along its axis. In other embodiments, the memristive layer 814A may not completely surround the core 812. In some embodiments, the memristive layer 814A comprises HfO x 、TiO x(where x indicates various stoichiometric amounts) and / or additional memristive materials. In some embodiments, memristive layer 814A is composed of HfO. Memristive layer 814A can be monolithic, comprising a single memristive material. In other embodiments, multiple memristive materials can be present in memristive layer 814A. In other embodiments, other configurations of (multiple) memristive materials can be used. However, it is desirable that memristive layer 814A be between electrode 830 and core 812. Thus, nanofiber 800A has a programmable resistance between the electrodes.
[0065] Nanofiber 800B includes a core 812 and an insulator 814B. Also shown are a memristive plug 820 and an electrode 870. The core 812 of nanofiber 800B is similar to the core 812 of nanofiber 800A. The insulator 814B coats the conductive core 812, but has a hole 816 therein. In some embodiments, the insulator 814B is thick enough to electrically insulate the conductive core 812 in the area where the insulator 814B covers the conductive core 812. For example, the insulator 814B can be at least a few nanometers to tens of nanometers thick. In some embodiments, the insulator 814B can be hundreds of nanometers thick. Other thicknesses are possible. In some embodiments, the insulator 814B surrounds the sides of the conductive core 812 except at the hole 816. In other embodiments, the insulator 814B can only surround the side portions of the conductive core 812. In such an embodiment, another insulator (not shown) can be used to insulate the conductive core 812 from its surroundings. For example, in such an embodiment, during the preparation of a device incorporating nanofibers 800B, an insulating layer can be deposited on the exposed portion of the conductive core 812. In some embodiments, a barrier layer can be provided in the hole 816. Such a barrier layer is between the conductive core 812 and the memristive plug 820. Such a barrier layer can reduce or prevent material migration between the conductive core 812 and the memristive plug 820. However, such a barrier layer is conductive so as to facilitate connection between the conductive core 812 and the electrode 830 through the memristive plug 820. In some embodiments, the insulator 114 includes one or more of SiO2, HfO2, Ta2O5, Al2O3, and polyvinyl pyrrolidone (PVP).
[0066] The memristive plug 820 is located within the hole 816. In some embodiments, the memristive plug 820 is completely within the hole 816. In other embodiments, a portion of the memristive plug 820 is outside the hole 816. In some embodiments, the memristive plug 820 may include HfO. x 、TiO x(where x indicates various stoichiometric amounts) and / or additional memristive materials. In some embodiments, the memristive plug 820 is composed of HfO. The memristive plug 820 can be monolithic, comprising a single memristive material. In other embodiments, multiple memristive materials can be present in the memristive plug 920. For example, the memristive plug 820 can include multiple layers of memristive material. In other embodiments, other configurations of (multiple) memristive materials can be used.
[0067] Figure 9 Depicts an embodiment of a sparsely connected crossbar array 900 that can be used for a sparsely connected linear programmable network layer 610. The crossbar array 900 includes horizontal lines 910-1 to 910-(n+1), vertical lines 930-1 to 930-m, and programmable conductances 920-11 to 920-nm. In some embodiments, the programmable conductances 920-11 to 920-nm are memristors. In some embodiments, the programmable conductances 920-11 to 920-nm can be memristive fibers arranged in the crossbar array. As in Figure 9 As can be seen in FIG. 1 , some conductances are missing. Thus, not all horizontal lines 910-1 to 910-(n+1) are connected to all vertical lines 930-1 to 930-m at each intersection via programmable conductances 920-11 to 920-nm. For example, line 930-2 is not connected to line 910-2. Similarly, line 910-n is not connected to line 930-n. Thus, crossbar array 900 is a sparsely connected network that can be used in programmable network layer 910. Thus, sparsely connected networks 900 and / or 700 can be used in a linear programmable network layer configured for use with balanced propagation. Thus, system performance can be improved.
[0068] Although the foregoing embodiments have been described in some detail for purposes of clarity of understanding, the present invention is not limited to the details provided. There are many alternative ways to implement the present invention. The disclosed embodiments are illustrative and not restrictive.
Claims
1. A system for performing learning, comprising: a linear programmable network layer comprising a plurality of inputs, a plurality of outputs, and a plurality of linear programmable network components interconnected between the plurality of inputs and the plurality of outputs; as well as a nonlinear activation layer coupled to the plurality of outputs; an additional linear programmable network layer comprising an additional plurality of inputs, an additional plurality of outputs, and an additional plurality of linear programmable network components interconnected between the additional plurality of inputs and the additional plurality of outputs, the nonlinear activation layer being coupled to the additional plurality of inputs, wherein the linear programmable network layer, the nonlinear activation layer, and the additional linear programmable network layer are configured to have a static state at a minimum value for the content of the system; wherein the plurality of input signals provided to the plurality of inputs generate a plurality of output signals at the plurality of outputs and generate an additional plurality of output signals at the additional plurality of outputs and correspond to a static state; The additional plurality of outputs are configured to receive a plurality of disturbances, the plurality of disturbances generating a plurality of disturbance output signals at the plurality of outputs, the plurality of outputs comprising a plurality of nodes configured to measure the plurality of disturbance output signals, the gradients of the plurality of linear programmable network components of the linear programmable network layer being determined based on the plurality of disturbance output signals and the plurality of output signals using balanced propagation, and at least one of the plurality of linear programmable network components in the linear programmable network layer being configured to be reprogrammed based on the gradients.
2. The system of claim 1 , wherein the nonlinear activation layer further comprises: Non-linear activation module; as well as A regeneration module is coupled to the plurality of outputs and to the nonlinear activation module, the regeneration module being configured to scale a plurality of output signals from the plurality of outputs.
3. The system of claim 2, wherein the regeneration module comprises a bidirectional amplifier.
4. The system of claim 1, wherein the linear programmable network layer comprises a programmable resistive network layer.
5. The system of claim 4, wherein the programmable resistor network layer comprises a fully connected programmable resistor network layer.
6. The system of claim 5, wherein the fully connected programmable resistor network layer comprises a crossbar array, the crossbar array comprising a plurality of programmable resistors. The system of claim 6 , wherein the plurality of programmable resistors comprises a plurality of memristors.
8. The system of claim 4, wherein the programmable resistor network layer comprises a sparsely connected programmable resistor network layer.
9. The system of claim 8, wherein the programmable resistor network layer comprises an array of partially connected crossbars.
10. The system of claim 8, wherein the programmable resistor network layer comprises: a plurality of nanofibers, each of the plurality of nanofibers having a conductive core and a memristive layer surrounding at least a portion of the conductive core; as well as A plurality of electrodes are provided, and a portion of the memristive layer is between the conductive cores of the plurality of nanofibers and the plurality of electrodes.
11. The system of claim 8, wherein the programmable resistor network layer comprises: a plurality of nanofibers, each of the plurality of nanofibers having a conductive core and an insulating layer surrounding at least a portion of the conductive core, the insulating layer having a plurality of pores therein; a plurality of memristive plugs for the plurality of wells, at least a portion of each of the plurality of memristive plugs being located in each of the plurality of wells; as well as A plurality of electrodes are provided, and the plurality of memristive plugs are between the conductive core and the plurality of electrodes.
12. The system of claim 1, wherein the nonlinear active layer comprises a plurality of diodes.
13. A system for performing learning, comprising: a plurality of linear programmable network layers, each of the plurality of linear programmable network layers comprising a plurality of inputs, a plurality of outputs, and a plurality of linear programmable network components interconnected between the plurality of inputs and the plurality of outputs; as well as at least one nonlinear activation layer inserted between the plurality of linear programmable network layers, each of the at least one nonlinear activation layer being coupled to the plurality of outputs of a linear programmable network layer in the plurality of linear programmable network layers and to the plurality of inputs of a next linear programmable network layer in the plurality of linear programmable network layers, each of the at least one nonlinear activation layer comprising a nonlinear activation module and a regeneration module, the regeneration module being configured to scale a plurality of output signals from the plurality of outputs, the plurality of linear programmable network layers and the at least one nonlinear activation layer being configured to minimize content for the system, wherein the plurality of linear programmable network layers comprises a first linear programmable network layer, a second linear programmable network layer, and a third linear programmable network layer; wherein a plurality of input signals are provided to a plurality of inputs of a first linear programmable network layer, the plurality of input signals generating a plurality of third output signals at a plurality of outputs of a third linear programmable network layer and corresponding to a static state of the system; The multiple outputs of the third linear programmable network layer are configured to receive multiple disturbances, the multiple disturbances generate multiple disturbance output signals at the multiple outputs of the second linear programmable network layer, the multiple outputs of the second linear programmable network layer include multiple nodes for measuring the multiple disturbance output signals, the gradient of the second linear programmable network layer is determined based on the multiple disturbance output signals and the multiple third output signals using balanced propagation, and at least one of the multiple linear programmable network components in the linear programmable network layer is configured to be reprogrammed based on the gradient.
14. The system of claim 13, wherein each of the plurality of linear programmable network layers comprises a programmable resistive network layer.
15. The system of claim 14, wherein the programmable resistor network layer comprises a fully connected programmable resistor network layer.
16. The system of claim 14, wherein the programmable resistor network layer comprises a sparsely connected programmable resistor network layer.
17. The system of claim 14, wherein the programmable resistance network layer comprises a plurality of memristive devices.
18. A method for performing learning, comprising: providing a plurality of input signals to a learning system, the learning system comprising a plurality of linear programmable network layers and at least one nonlinear activation layer, each of the plurality of linear programmable network layers comprising a plurality of inputs, a plurality of outputs, and a plurality of linear programmable network components interconnected between the plurality of inputs and the plurality of outputs, the at least one nonlinear activation layer being interposed between the plurality of linear programmable network layers, each of the at least one nonlinear activation layer being coupled to the plurality of outputs of a linear programmable network layer in the plurality of linear programmable network layers and to the plurality of inputs of a next linear programmable network layer in the plurality of linear programmable network layers, the plurality of linear programmable network layers and the at least one nonlinear activation layer being configured to have a static state at a minimum value of content of the learning system, the plurality of input signals resulting in a plurality of output signals corresponding to the static state; perturbing the plurality of outputs for a first linear programmable network layer among the plurality of linear programmable network layers to provide a plurality of perturbed output signals at the plurality of inputs of a second linear programmable network layer among the plurality of linear programmable network layers; determining gradients of the plurality of linear programmable network components for a second linear programmable network layer using balanced propagation based on the plurality of perturbed output signals and the plurality of output signals; as well as At least one of the plurality of linear programmable network components in the second linear programmable network layer is reprogrammed based on the gradient.
19. The method of claim 18, wherein perturbing further comprises: A plurality of perturbation input signals are provided to the plurality of outputs of the first linear programmable network layer, the plurality of perturbation input signals corresponding to a second plurality of outputs that are closer to a plurality of target outputs than the plurality of output signals.
20. The method of claim 18, further comprising: Providing an input signal, perturbing the plurality of outputs, determining a gradient, and reprogramming are performed iteratively.
Citation Information
Patent Citations
Scalable silicon based resistive memory device
US10290801B2
Neuromorphic processing devices
US20160379110A1
Quantum statistic machine
US20190005402A1