A method and device for training weights of a memristor convolutional neural network

CN115310581BActive Publication Date: 2026-09-08GUILIN UNIV OF ELECTRONIC TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110491512.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-06
Publication Date
2026-09-08
Estimated Expiration
2041-05-06

AI Technical Summary

Technical Problem

[0010]1.忆阻器神经网络的常规学习方案过于复杂:网络需要在更新之前计算出所需的电脉冲的数量,这将大大增加系统计算的复杂度,尤其对于大规模的卷积神经网络,其参数量往往数以百万计甚至数以亿计,这一非线性计算过程将消耗巨量的计算资源和能耗

Benefits of technology

[0065] By adopting the above technical solution, the weight training method is presented in the form of computer-readable code and stored on a computer storage medium. When the processor runs the computer code on the medium, the steps of the above method can reduce the error caused by the nonlinear update characteristics of the memristor and effectively improve the online learning performance of the memristor neural network. At the same time, it can also take into account the high energy efficiency and low power consumption characteristics of the overall system, and has good hardware friendliness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115310581B_ABST
    Figure CN115310581B_ABST
Patent Text Reader

Abstract

The application relates to an artificial intelligence network, and particularly discloses a memristor convolutional neural network weight training method and device. A plurality of differential circuits are used to form a memristor array of a neural network hardware system, each two devices form a differential circuit, and at least one of the two devices comprises a memristor. The method comprises the following steps: binding the two devices forming each differential circuit to form a neuron; mapping the weight of the neuron according to the difference between the electrical parameters of the two devices; updating the neuron weight by applying a single-step pulse to obtain the update direction of the neuron weight; when the update direction of the weight is positive, applying a positive single-step pulse to the memristor to update the neuron weight again; when the update direction of the weight is negative, applying a negative single-step pulse to the memristor to update the neuron weight again; and when the weight update iteration of all neurons reaches a convergence condition, the iteration is ended to complete the training. The weight update calculation process of the memristor neural network system is simplified, and the consumption of the calculation resources in the process is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of spiking neural network technology, and in particular to a method and apparatus for training weights in a memristor convolutional neural network. Background Technology

[0002] Memristors are a novel type of two-port nanodevice with a certain degree of memory in their conductance, hence the name memory resistors (i.e., memristors). Memristor arrays can utilize Ohm's law and Kirchhoff's laws to achieve high-performance, low-power multiplication and accumulation, which is highly compatible with the core operation (matrix multiplication) in neural network computing systems. Therefore, memristors provide a new research approach for the hardware implementation of neural networks and have attracted widespread attention from researchers.

[0003] like Figure 1 The diagram illustrates the inference process (forward propagation) of a neural network. Memristor networks can easily perform matrix multiplication using Ohm's law and Kirchhoff's laws. For the network training process (backward propagation), the weights to be updated can be calculated using the stochastic gradient descent (SGD) algorithm, and the synaptic weights can be updated by applying a corresponding number of electrical pulses to the device terminals.

[0004] However, memristor devices and arrays have almost unavoidable drawbacks such as random fluctuations in conductance and nonlinear weight update characteristics, making them difficult to apply directly to the online learning process of neural networks. This also significantly reduces their accuracy in offline inference processes.

[0005] The nonlinear update characteristic of a memristor refers to the phenomenon that the change in conductance of a memristor is nonlinearly related to the number of pulses in its conductance-pulse modulation curve (as follows). Figure 2 (As shown). Generally, the nonlinearity follows an exponential function form. The conductance state function of a memristor can be represented by five parameters: Gmax (maximum conductance), Gmin (minimum conductance), Pmax (number of conductance states), Ap (nonlinearity during the weighted LTP process) - see [link to relevant documentation]. Figure 2 (a) Ad (Nonlinearity in the weight reduction LTD process) - see Figure 2 (b). Can be expressed as follows:

[0006]

[0007] In the above formula, x represents the current state parameter of the device, and 0≤x≤Pmax; A is equal to Ap in the weight increase process (LTP process) and equal to Ad in the weight decrease process (LTD process).

[0008] The online learning scheme for memristor neural networks is roughly as follows: Figure 3 As shown. First, the error is calculated by forward propagation. Then, the required update amount for the current weights is calculated according to the standard neural network back propagation algorithm and gradient descent (SGD) algorithm. Next, the required number of pulses is calculated according to the nonlinear update characteristic curve of the memristor device. Then, a certain number of electrical pulses are applied across the memristor by gating to adjust the current conductance value of the memristor.

[0009] Current conventional learning schemes based on memristor networks have at least the following drawbacks:

[0010] 1. Conventional learning schemes for memristor neural networks are too complex: the network needs to calculate the number of electrical pulses required before updating, which greatly increases the computational complexity of the system. Especially for large-scale convolutional neural networks, the number of parameters often reaches millions or even hundreds of millions. This nonlinear calculation process will consume huge amounts of computing resources and energy.

[0011] 2. Memristor network arrays cannot perform the task of updating electrical pulses on their own; a large number of external control logic circuits are required to implement this. For example, the calculation of the required weight update amount and number of pulses is often provided by a computer with calculation software installed, and then the calculation and conversion of the device's conductance weight update are assisted by software and hardware control methods such as MATLAB and microcontroller (MCU).

[0012] 3. No effective measures were taken to address the nonlinear update characteristics inherent in the memristor: After applying a certain number of electrical pulses to update the conductance weights, the actual conductance state of the memristor deviates from the ideal state, resulting in low learning performance of the memristor neural network. Summary of the Invention

[0013] To simplify the weight update calculation process of memristor neural networks and reduce the consumption of computational resources in this process, in a first aspect, this application provides a weight update method for memristor convolutional neural networks, adopting the following technical solution:

[0014] A memristor array is constructed using several differential circuits to form a neural network hardware system, with each pair of devices constituting a differential circuit, and at least one of the two devices including a memristor; the method includes:

[0015] The two devices that make up each differential circuit are bound together to form a neuron;

[0016] The weights of neurons are mapped based on the difference in electrical parameters between the two devices;

[0017] The direction of neuron weight update is obtained by updating neuron weights by applying single-step pulses;

[0018] When the weight update direction is positive, a positive single-step pulse is applied to the memristor to update the neuron weights again;

[0019] When the weight update direction is negative, a negative single-step pulse is applied to the memristor to update the neuron weights again;

[0020] The training process ends when the weight updates of all neurons converge.

[0021] By adopting the above technical solution, calculations of learning rate, weight update amount, and required number of pulses are eliminated compared to related technologies. After obtaining the direction of weight update through the gradient descent algorithm, a single-pulse scheme is used for weight update. Knowing the direction of weight update in each learning process, a pulse is applied to the corresponding memristor device to complete the weight update for this time. During each weight update (learning) process, by applying a single-step pulse to memristors with large conductance changes, the weight parameters can be made to approach the optimal value with a certain step size until they reach or infinitely approach the optimal value. This system, which continuously and automatically adjusts the processing method according to the different characteristics of the processed data during the data processing process, so that it is always in or approaching the optimal operating state, is called an adaptive system, and the learning method based on the adaptive system is called an adaptive learning method. For memristor neural networks, they need a certain learning ability and the ability to continuously adjust network parameters using the network's capabilities. Only providing parameter update direction suggestions (whether the weights need to be increased or decreased) based on the gradient descent algorithm can undoubtedly significantly reduce the computational complexity required for updates and is also more conducive to hardware implementation.

[0022] Optionally, as one embodiment of the above scheme, both devices constituting the differential circuit include memristors; the parameters of the two memristors forming the differential circuit are different to satisfy the requirement that a predictable trend of change in conductance be obtained after an electrical pulse is applied; the method includes:

[0023] The two memristors that make up each differential circuit are bound together to form a neuron;

[0024] The weights of neurons are mapped based on the difference in conductance between two memristors;

[0025] The direction of neuron weight update is obtained by updating neuron weights by applying single-step pulses;

[0026] When the weight update direction is positive, a positive single-step pulse is applied to the memristor with a large change in conductance to update the neuron weights again;

[0027] When the weight update direction is negative, a negative single-step pulse is applied to the memristor with a large change in conductance to update the neuron weights again;

[0028] The training process ends when the weight updates of all neurons converge.

[0029] By adopting the above scheme, a differential circuit architecture for a neural network hardware system is constructed using two memristors. By utilizing the directional trend influence of the memristor parameters on the magnitude of conductance changes, single-step electrical pulses are selectively applied to memristors with larger conductance changes based on the direction of weight updates, and the weights are iteratively updated. This allows the weights of each neuron in the neural network to approach the optimal value within a high probability range. Compared to a neural network with a single memristor differential circuit architecture, it is more flexible and can flexibly adjust the step size according to specific needs, thereby accelerating the iteration process.

[0030] Optionally, each differential unit consists of two independent memristors; the step of mapping the neuron weights based on the conductance values ​​of the two memristors includes:

[0031] Half of the neurons are mapped to positive weights; the other half are mapped to negative weights.

[0032] In the step of applying a positive single-step pulse to the memristor with a large change in conductance to update the neuron weights again when the weight update direction is positive:

[0033] Memristors with large variations in conductivity include those with large variations in weights during positive weighting mapping.

[0034] In the step of applying a negative single-step pulse to the memristor with a large change in conductance to update the neuron weights again when the weight update direction is negative:

[0035] Memristors with large variations in conductivity include those with large variations in weight in the negative weighting mapping.

[0036] By adopting the above technical solution, in the neural network, half of the weight parameters are initially preset to positive values, while the rest are negative values. Since the online learning process of a neural network is essentially a process of solving a system of multivariable equations, the number of unknowns far exceeds the number of equations, resulting in an infinite number of solutions that satisfy the conditions. This is the source of the robustness of the neural network. Therefore, presetting half of the weight parameters to positive values ​​and the other half to negative values ​​will not affect the final performance of the network.

[0037] In the preset positive definite weights, the circuit scheme of memristor w1-memristor w2 (i.e., w1-w2) is adopted; in the preset negative definite weights, the circuit scheme of memristor w2-memristor w1 (i.e., w2-w1) is adopted. The device parameters (such as the cross-sectional area of ​​the device) of memristor w1 and memristor w2 are different. After obtaining the direction of weight update, based on the direction of change, for memristors with large conductance changes in differential circuits with positive weight mapping, after receiving a positive electrical pulse, the weight changes in the mapped neurons are also large, and they move closer to the optimization goal. Therefore, applying a positive single-step pulse to the memristor can quickly achieve network weight optimization iteration within a high probability range. Based on negative changes, for memristors with large conductance changes in differential circuits with negative weight mapping, after receiving a negative electrical pulse, the weight changes in the mapped neurons are also large, and they move closer to the optimization goal. Therefore, applying a negative single-step pulse to the memristor can quickly achieve network weight optimization iteration within a high probability range.

[0038] Optionally, the differential circuit is composed of two memristors with different cross-sectional areas; wherein the weight change of the first memristor is relatively larger than that of the second memristor after a single pulse is applied; wherein the conductance of the first memristor is W1 and the conductance of the second memristor is w2.

[0039] The steps for mapping positive weights include:

[0040] The difference between the conductance value W1 of the first memristor and the conductance value W2 of the second memristor is used to map the positive weight of the neuron: W1-W2;

[0041] The steps for mapping negative weights include:

[0042] The negative weight of the neuron is mapped by the difference between the conductance value W2 of the second memristor and the conductance value W1 of the first memristor: W2-W1.

[0043] By adopting the above technical solution, the two memristors are designed with different area sizes from the outset (for example, the area of ​​w1 can be designed to be twice the area of ​​w2). Therefore, the intrinsic conductivity of device w1 will (with a high probability) be greater than that of w2. In this case, device w = w1 - w2 is set as the positive weight mapping of the network, and device w = w2 - w1 is set as the negative weight mapping. It is worth noting that in the above settings, due to the random initialization of the weights, the weight values ​​in the positive weight mapping are only likely to be positive, not guaranteed to be positive. The same applies to the negative weight mapping. This introduces a certain degree of probabilistic update behavior, which can serve as a compensation mechanism for network weight updates.

[0044] Optionally, when the weight update direction is positive, the step of applying a positive single pulse to the memristor with a large weight change amplitude in the positive weight mapping includes:

[0045] Apply a positive single-step pulse to both the first and second memristors, or

[0046] A positive single-step pulse is applied to the first memristor, and a negative single-step pulse is applied to the second memristor.

[0047] By adopting the above technical solution, in each update of the network weights, the two devices constituting the difference pair change in the same direction, but their change amplitudes are unequal. This means it cannot be guaranteed that the new weight change obtained after each pulse application will be consistent with the expected weight change. However, appropriate parameter design (mainly the area parameters of devices w1 and w2) can ensure that the above requirement is met with a high probability. When the device area of ​​device w1 is twice that of device w2, the change amplitude of device w1 is likely to be greater than that of device w2. Therefore, the weight values ​​in the network are likely updated according to the requirements given by the algorithm, meaning that the update of the weights in the network is likely successful. This setting introduces another probabilistic update behavior, which can serve as another compensation mechanism for network weight updates. Compared to the above scheme, when the weight update direction is positive, applying positive single-step pulses to both the first and second memristors simultaneously is suitable for weight updates with small step amplitudes. Applying positive single-step pulses to the first memristor and negative single-step pulses to the second memristor is suitable for weight updates with larger step amplitudes. Based on different applicable scenarios, multiple training methods are provided to improve the adaptability and flexibility of the training method.

[0048] Optionally, the step of applying a negative single pulse to memristors with large weight changes in the negative weight mapping when the weight update direction is negative includes:

[0049] Apply a negative single-step pulse to both the first and second memristors, or

[0050] A negative single-step pulse is applied to the first memristor, and a positive single-step pulse is applied to the second memristor.

[0051] By adopting the above technical solution, similarly, compared to the above solution, when the weight update direction is negative, applying negative single-step pulses to both the first and second memristors simultaneously is suitable for weight updates with small step amplitudes. Applying negative single-step pulses to the first memristor and positive single-step pulses to the second memristor is suitable for weight updates with large step amplitudes. Based on different applicable scenarios, multiple training methods are provided, improving the adaptability and flexibility of the training method.

[0052] Optionally, as another embodiment of the above scheme, another device constituting the differential circuit includes a shared resistor; the step of updating the neuron weights by applying a single-step pulse to obtain the update direction of the neuron weights further includes:

[0053] The resistance value of the shared resistance in all neurons is randomly initialized to R, where R is a constant value;

[0054] The conductance value of the memristor is W', and the step of mapping the neuron weights based on the difference in electrical parameters between the two devices includes:

[0055] The weight of the neuron is mapped based on the difference between the conductance of the memristor and the initial value of the shared resistance. The weight W of the neuron is: W' – R.

[0056] By adopting the above technical solution, the circuit area of ​​the neural network hardware system is significantly reduced compared to the previous implementation. This is because the 1T1R-1T1R structure requires twice the number of components, while different differential pairs in the 1T1R-1R structure can share the same resistor. Theoretically, this circuit structure could reduce the circuit area consumption by 50% for the same network. Furthermore, it eliminates the need to calculate the specific number of voltage pulses. What is needed is to determine the sign of the weight update Δw during training. This simple learning rule enables the memristor network to update its parameters, thereby achieving self-optimization capabilities. This can improve the performance, area, and power consumption of memristor neural network chips.

[0057] Optionally, as another embodiment of the above scheme, another device constituting the differential circuit includes a shared resistor; the step of updating the neuron weights by applying a single-step pulse to obtain the update direction of the neuron weights further includes:

[0058] The resistance value R of the shared resistance in all neurons is predetermined, where R is the maximum conductance value G of the memristor. max With minimum conductivity G min Half of the sum, R = (G max / 2+G min / 2);

[0059] The conductance value of the memristor is W', and the step of mapping the weights of the neuron based on the difference in electrical parameters between the two devices includes:

[0060] The neuron's weight is mapped based on the difference between the memristor's conductance and the initial value of the shared resistance. The neuron's weight W is: W' – (G max / 2+G min / 2).

[0061] By adopting the above technical solution, the hardware architecture is the same as that of the previous embodiment. The difference is that the resistance value of the shared resistor of all row memristors is determined to be the same. Compared with the previous embodiment, which randomly initialized the power supply resistance value of each row memristor, this method is conducive to obtaining a uniform and symmetrical neuron weight distribution during the iterative calculation process. Compared with the randomly initialized weight, it is conducive to high-precision online learning, and the error between the optimized weight and the optimal weight is relatively small.

[0062] Secondly, this application also provides a memristor convolutional neural network weight training device, which adopts the following technical solution, including:

[0063] The memory stores the weight training program.

[0064] The processor executes the steps of the above method when running the weight training program.

[0065] By adopting the above technical solution, the weight training method is presented in the form of computer-readable code and stored on a computer storage medium. When the processor runs the computer code on the medium, the steps of the above method can reduce the error caused by the nonlinear update characteristics of the memristor and effectively improve the online learning performance of the memristor neural network. At the same time, it can also take into account the high energy efficiency and low power consumption characteristics of the overall system, and has good hardware friendliness. Attached Figure Description

[0066] Figure 1 This is a diagram of the forward propagation process of a memristor neural network in related technologies.

[0067] Figure 2 (a) is a diagram of the update model for the nonlinear weighting process of memristors in related technologies.

[0068] Figure 2 (b) is a diagram of the update model for the nonlinear weight reduction process of memristors in related technologies.

[0069] Figure 3 This is a flowchart of the memristor neural network weight training process in related technologies.

[0070] Figure 4 This is a flowchart of the memristor convolutional neural network weight training method provided in Embodiment 1 of this application.

[0071] Figure 5 This is a hardware framework diagram for implementing the memristor convolutional neural network weight training method provided in Embodiment 2 of this application.

[0072] Figure 6 This is a flowchart of the memristor convolutional neural network weight training method provided in Embodiment 2 of this application.

[0073] Figure 7 (a) is the differential circuit architecture used in the memristor neural network in the weight training method of the memristor convolutional neural network provided in Embodiment 2 of this application.

[0074] Figure 7 (b) is a graph showing the trend of the conductance of two memristors under the applied electrical pulse state in the memristor convolutional neural network weight training method provided in Embodiment 2 of this application.

[0075] Figure 8 (a) is a diagram of the nonlinear weight increase process update model of two memristors in the memristor convolutional neural network weight training method provided in Embodiment 2 of this application.

[0076] Figure 8 (b) is a diagram of the nonlinear weight reduction process update model of the two memristors in the memristor convolutional neural network weight training method provided in Embodiment 2 of this application.

[0077] Figure 9 This is a graph showing the ratio of the maximum conductance values ​​of the two memristors in Example 2 and the probability of successful weight updates each time.

[0078] Figure 10 This is a hardware framework diagram of the memristor convolutional neural network weight training method provided in Embodiment 3 of this application;

[0079] Figure 11 This is a flowchart of the memristor convolutional neural network weight training method provided in Embodiment 3 of this application;

[0080] Figure 12 This is a diagram illustrating the accuracy of neural network weight optimization using different algorithms.

[0081] Figure 13 This is a schematic diagram illustrating the accuracy of LeNet-5 networks trained using different schemes under different nonlinearities Ap and Ad. Detailed Implementation

[0082] The present application will be further described in detail below with reference to the accompanying drawings.

[0083] This specific embodiment is merely an explanation of this application and is not intended to limit it. Those skilled in the art, after reading this specification, can make modifications to this embodiment without contributing any inventive step, but such modifications are protected by patent law as long as they fall within the scope of the claims of this application. To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application.

[0084] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.

[0085] The embodiments of the vision testing terminal of this application will be described in further detail below with reference to the accompanying drawings.

[0086] Example 1

[0087] The hardware architecture of a memristor neural network uses differential circuits. A memristor array is constructed from several differential circuits, with each pair of devices forming a differential circuit. At least one of these two devices includes a memristor. Figure 4 As shown, the above method includes:

[0088] Step S1: Bind the two devices that make up each differential circuit to form a neuron;

[0089] In the software system, each component of the differential circuit is bound to a neuron, and a convolutional neural network is formed through the connections between neurons. The connections between neurons are represented by weights.

[0090] Step S2: Map the weights of the neurons based on the difference in electrical parameters between the two devices;

[0091] It should be noted that the electrical parameters here are different in different differential circuit architectures, specifically resistive parameters. For memristors, the conductance value is used, and for resistors, the resistance value is used.

[0092] Step S3: Update the neuron weights by applying a single-step pulse to obtain the update direction of the neuron weights;

[0093] Compared to related technologies, the key feature of this application is that it applies a single pulse to each row of memristors in the memristor array to update the weights of each neuron. The method for calculating the neuron weights during the update process can employ existing gradient descent algorithms, such as linear SGD or nonlinear SGD. Unlike related technologies, this application inputs a single-step pulse for each update, eliminating the need to calculate the number of electrical pulses applied to each memristor. Instead, it only needs to obtain the weight update direction for each neuron, and then determine the direction for the next single-step pulse application based on this direction. This significantly simplifies the weight update process and reduces computational load, thereby saving computational resources.

[0094] Step S4: When the weight update direction is positive, apply a positive single-step pulse to the memristor to update the neuron weights again;

[0095] When a positive weight update direction is obtained, a weight closer to the optimal value is obtained by applying a positive single-step pulse to both ends of the memristor.

[0096] Step S5: When the weight update direction is negative, apply a negative single-step pulse to the memristor to update the neuron weights again.

[0097] When obtaining a negative weight update direction, a weight closer to the optimal value is obtained by applying a negative single-step pulse to the memristor terminals.

[0098] Step S6: When the weight update iterations of all neurons reach the convergence condition, the iteration ends and the training is completed.

[0099] Repeat steps S2-S5, continuously updating and iterating the weights. When the convergence condition is met, the iteration ends and the network weight training is complete. The convergence condition can be the set weight error or the number of iterations, etc.

[0100] Example 2

[0101] As a first preferred embodiment, the hardware system of the memristor convolutional neural network weight training method, such as Figure 5 As shown, the hardware architecture includes a weight training device and a memristor convolutional neural network. The weight training device can be understood as a computing device outside the network system. This device includes at least a memory and a processor. The memory stores a software program, and the processor executes the memristor convolutional neural network weight training method when running the program.

[0102] The following section provides a detailed explanation of the implementation of the memristor convolutional neural network weight training method using the aforementioned system:

[0103] The hardware architecture of memristor convolutional neural networks is as follows: Figure 5 As shown, this scheme utilizes several differential circuits (see...). Figure 5 The memristor array (shown in the dashed box) constitutes the hardware system of the neural network. Each pair of memristors forms a differential circuit. The parameters of the two memristors forming the differential circuit are different to ensure that the change in conductance is predictable after an electrical pulse is applied. In this scheme, the memristor neural network adopts a differential circuit architecture. Since the conductance of memristors is not negative, while the weight parameters in the neural network may be negative (generally, the number of positive and negative weight parameters in the network is roughly equal), a differential-pair design is typically used in memristor neural networks. Two memristors (denoted as w1 and w2) are bound together to form a new set of weights w = w1 - w2 for network calculation. Specifically, as follows: Figure 7 As shown in (a), the two memristor arrays respectively use current-to-voltage conversion circuits to calculate their respective voltage output data. and The difference between the two values ​​is then calculated using a computing circuit (specifically a differential circuit), yielding the final output of the neuron.

[0104] The conventional approach (which we'll call "Scheme 0") updates weights in the following way: If the weight *w* is to be increased, the standard stochastic gradient descent (SGD) algorithm is used to calculate the weight change and the corresponding number of electrical pulses required. Several positive electrical pulses are then applied to device *w1*, while *w2* remains idle. If the weight *w* is to be decreased, the same SGD algorithm is used to calculate the weight change and the corresponding number of electrical pulses required. Several negative electrical pulses are then applied to device *w1*, while *w2* remains idle. This complex calculation significantly reduces the system's energy efficiency. Secondly, the second memristor (*w2*) is never operated during the update process, failing to meet the design principle of maximizing resource utilization. Furthermore, there is no compensation for the performance loss caused by the nonlinear characteristics of the devices.

[0105] Based on the above considerations, this application proposes a novel memristor differential circuit architecture, such as... Figure 7 As shown in (b), memristor devices w1 and w2 with different parameters are introduced and then combined into a new differential architecture unit.

[0106] like Figure 6 As shown, the method includes:

[0107] S10, the two memristors that make up each differential circuit are bound together to form a neuron;

[0108] In software systems, specifically neural networks, half of the weight parameters are initially set to positive values, while the remainder are set to negative values. This is because the online learning process of a neural network is essentially a process of solving a system of multivariable equations, where the number of unknowns far exceeds the number of equations. Therefore, there are infinitely many sets of solutions that satisfy the conditions, which is the source of the robustness of neural networks. Thus, setting half of the weight parameters to positive values ​​and the other half to negative values ​​does not affect the final performance of the network. Each differential circuit is bound to form a neuron, and the conductance values ​​of two memristors represent the weights of the neurons they form. For the preset positive definite weights, we use a circuit scheme of memristor w1-memristor w2 (i.e., w1-w2); for the preset negative definite weights, we use a circuit scheme of memristor w2-memristor w1 (i.e., w2-w1). The device parameters (such as the cross-sectional area of ​​the devices) of memristor w1 and memristor w2 are different.

[0109] S20, the weights of neurons are mapped based on the difference in conductance between the two memristors;

[0110] In the conventional scheme of related technologies ("Scheme 0"), the two memristors (w1 and w2) in each group of differential devices are physically indistinguishable. However, the scheme provided in this application specifies that the two memristors have different area sizes from the initial circuit design stage (for example, the area of ​​w1 can be designed to be twice the area of ​​w2). Therefore, the intrinsic conductivity of device w1 will (with a high probability) be greater than the conductivity of w2. In this case, device w = w1 - w2 is set as the positive weight mapping of the network, and device w = w2 - w1 is set as the negative weight mapping of the network. It is worth noting that in the above settings, due to the random initialization of the weights, the weight values ​​in the positive weight mapping are only likely to be positive, and cannot be guaranteed to be positive. The same applies to the negative weight mapping.

[0111] This introduces a certain degree of probabilistic update behavior, which can serve as a compensation mechanism for network weight updates. This probability is named Probability 1.

[0112] This probability value can be calculated as follows: Assuming w1 and w2 are randomly initialized according to the functional distribution of f(x), the range of w1 is [Gmin1, Gmax1], and the range of w2 is [Gmin2, Gmax2], then the probability of w = w1 - w2 > 0 can be expressed as:

[0113]

[0114] Specifically, if f(x) represents a uniform distribution and Gmin1 = Gmin2 = 0, Gmax1 = 2, Gmax2 = 1, the probability value is 75%.

[0115] It should be noted that the two memristors constituting the differential circuit here can be composed of two independent memristors, or one independent memristor and one shared memristor. When using the former differential circuit architecture, step S20 includes S21: half of the neurons' weights are mapped to positive weights; the other half's weights are mapped to negative weights. This facilitates accurate prediction of the weight update direction with a high probability. This scheme is suitable for 1T1R-1T1R circuit structures. When using the latter differential circuit architecture, step S200 is included before step S20: the conductance R of the shared memristor in all neurons is initialized. There are two initialization methods: S201 randomly initializes the conductance R of the shared memristor in all neurons, where R is a constant value; S202 predetermines the conductance R of the shared memristor in all neurons, where R is half the sum of the maximum and minimum conductance values ​​of the independent memristor, R = (G... max / 2+G min / 2); This scheme is applicable to the 1T1R-1R circuit structure.

[0116] S30, update the neuron weights by applying a single-step pulse to obtain the update direction of the neuron weights;

[0117] After the above settings, the memristor network uses our learning algorithm for parameter training. Specifically, in each update process, the two devices constituting the difference pair change in the same direction, but their change magnitudes are unequal. This means that it cannot be guaranteed that the new weight change obtained after each pulse application will be consistent with the expected weight change. However, appropriate parameter design (mainly the area parameters of devices w1 and w2) can ensure that the above requirements are met with a high probability. For example, when the device area of ​​device w1 is twice that of device w2, the change magnitude of device w1 is likely to be greater than that of device w2. Therefore, the weight values ​​in the network are likely to be updated according to the requirements given by the algorithm, that is, the update of the weights in the network is likely to be successful. This setting introduces another probabilistic update behavior, which can serve as another compensation mechanism for network weight updates. This probability is named Probability 2.

[0118] In the single-pulse update scheme, we cannot guarantee the amount of new weight change (Δw) obtained after each pulse application. actual ) and the expected change in weight (Δw) predict While the direction of updates may not be consistent, proper parameter design (primarily the area parameters of devices w1 and w2) can ensure that the above requirements are met under high probability conditions. Below, we will derive the solution method for this probability using a positive update of a positive weight as an example.

[0119] above Figure 8(a) Taking the update process of the "positive weights" in the network when Δw>0 as an example, the conductivity value can be expressed as:

[0120] W old =w 1_old -w 2_old =G(x1)-G(x2) (3)

[0121] W new =w 1_new -w 2_new =G(x1+1)-G(x2+1) (4)

[0122] ΔW=W new -W old >0 (5)

[0123] Where G(x) is the memristor conductance-pulse state function in formula (1), and x represents the pulse state of the memristor. Formula (5) represents that the weight parameter has been successfully updated in this round of weight update. It can be approximated by the low-order expansion of the Taylor function. For the first-order linear approximation, it can be equivalent to:

[0124] ΔW=[G(x1+1)-G(x1)]-[G(x2+1)-G(x2)]≈G'(x1)-G'(x2)>0 (6)

[0125] Where G'(x) is the derivative of formula (1). By iterating through x1 and x2, the probability of a successful update in one step can be obtained by integration:

[0126]

[0127] Where Pmax represents the number of conductance states of devices w1 and w2, which can be understood as the final number of single-step pulses applied. In this scheme, the number of conductance states of devices w1 and w2 is the same, meaning that a single pulse is applied to both devices each time. A1 is the nonlinear coefficient of device w1 in the weight update process, Gmax1 is the maximum conductance value of device w1, and Gmin1 is the minimum conductance value of device w1; A2 is the nonlinear coefficient of device w2 in the weight update process, Gmax2 is the maximum conductance value of device w2, and Gmin2 is the minimum conductance value of device w2. Here, device w1 can be understood as the first memristor, and device w2 can be understood as the second memristor. Temporary variables are defined as follows:

[0128]

[0129]

[0130] We found that the probability of G'(x1)-G'(x2)>0 is related to the ratio of Gmax1 and Gmax2, G. max 1 / G max 2. They have a close relationship. For example, when the ratio of Gmax1 to Gmax2 is G... max 1 / G max 2 = 1 (at the same time G) min1 =G min2 When = 0, the probability of G'(x1) - G'(x2) > 0 will reach 50%, which is consistent with our intuitive understanding. When the maximum conductance Gmax1 of device w1 is twice the maximum conductance Gmax2 of device w2, this probability value will reach Figure 9 The figure of 95.4% is shown. This means that after a single-pulse update, W actually... new The value is likely greater than W. old The value of , meaning the probability of a successful weight update, is quite high. This ensures that the neural network can complete the learning process well and ultimately achieve good network performance.

[0131] By calculation, the probability of G'(x1)-G'(x2)>0 is the ratio of Gmax1 and Gmax2, G. max 1 / G max The relationship between 2 is as follows: Figure 9 As shown, the greater the ratio of the maximum conductance values ​​of the two memristors constituting the same differential circuit, the higher the probability of a successful weight update in a single operation. Figure 8 When Δw < 0 in (b), the same conclusion can be obtained through the derivation of the above formula, which will not be repeated here.

[0132] S40, when the weight update direction is positive, apply a positive single-step pulse to the memristor with a large change in conductance to update the neuron weights again;

[0133] Based on the above derivation and calculation, this scheme uses a single-step pulse iterative training method to train the weights of each neuron in the network. When the probability of a neuron being updated to be positive is high, applying a positive single-step pulse to a memristor with a large change in conductance can make it move closer to the optimal value.

[0134] S50, when the weight update direction is negative, apply a negative single-step pulse to the memristor with a large change in conductance to update the neuron weights again;

[0135] Similarly, when the probability of a neuron updating negatively is high, applying a negative single-step pulse to a memristor with a large change in conductance can make it move closer to the optimal value.

[0136] When two independent memristors are used in S20, the memristor with a larger change in conductance value in step S40 is actually the memristor with a larger change in weight in the positive weighting mapping; the memristor with a larger change in conductance value in step S50 is the memristor with a larger change in weight in the negative weighting mapping.

[0137] When a combination of an independent memristor and a shared memristor is used in S20, a single pulse is applied to the independent memristor in both steps S40 and S50.

[0138] S60: When the weight update iterations of all neurons reach the convergence condition, the iteration ends and the training is completed.

[0139] Repeat steps S20-S50. When the preset number of iterations or the set convergence condition is reached, stop the iteration and complete the training.

[0140] In the algorithm of this application, the calculations of network weight update amounts in traditional memristor network algorithms are simplified. In the original "Scheme 0" update scheme, the update amount of each weight is determined by the backpropagation mechanism and the stochastic gradient descent algorithm, and the learning rate, as a hyperparameter of the network, is also introduced into the memristor neural network, which is consistent with the research ideas of neural networks in the conventional field of artificial intelligence. However, as a dedicated circuit architecture of artificial neural networks, if the forward and backward propagation mechanisms of memristor neural networks are completely copied from the original artificial neural networks, and the memristor device is only used as a variable resistor, the characteristics of the memristor itself cannot be fully utilized. In addition, due to the limitations of the memristor's own conductance fluctuation phenomenon and the limited number of conductance states, the error generated during the weight update process will accumulate layer by layer in the neural network, and the network performance achieved by this scheme will inevitably have an upper limit. Based on these considerations, we will eliminate the calculations of learning rate, weight update amount, and the number of pulses required in the update scheme, and only provide parameter suggestions on the update direction (whether the weight needs to be increased or decreased) based on the gradient descent algorithm. This will undoubtedly significantly reduce the computational complexity required for the update and is also more conducive to hardware implementation.

[0141] After obtaining the direction of weight updates using the gradient descent algorithm, we also propose a single-pulse scheme for weight updates. Knowing the direction of weight updates in each learning iteration, a pulse is applied to the corresponding memristor device to complete the weight update. During each weight update (learning) process, the weight parameters move towards the optimal value with a certain step size until they reach or are infinitely close to the optimal value. This system, which continuously and automatically adjusts the processing method according to the different characteristics of the processed data, thereby keeping it in or approaching the optimal operating state, is called a system, and the learning method based on the system is called a learning method. For memristor neural networks, a certain learning ability is required, and the network parameters must be continuously adjusted using the network's capabilities.

[0142] Example 3

[0143] As a second preferred embodiment of the first example, the hardware system for the memristor convolutional neural network weight training method is described in [reference needed]. Figure 10 The difference from Embodiment 2 is that, in the two devices constituting the differential circuit, one device includes a memristor, and the other device includes a shared resistor. Each row of memristors in the memristor array shares a common resistor, and each memristor in that row forms a differential circuit with the shared resistor. In each differential circuit, the memristor and the shared resistor are bound together to form a neuron. See [link to previous section]. Figure 11 The circuit architecture weight training method includes the following steps:

[0144] Step S100: Bind the memristors and shared resistors that make up each differential circuit to form a neuron;

[0145] For each row of memristors in the memristor array, each memristor in that row is bound to a shared resistor in that row to form a neuron. If there are N rows of memristors in the memristor array, the hardware system is configured with N shared resistors, one for each row of memristors. During the binding process, row-specific memristors are paired in an orderly manner. Compared to Embodiment 2, this scheme significantly reduces the number of devices, circuit area, and power consumption because the shared resistors change the connection relationship of the differential circuit.

[0146] Step S200: Map the weights of neurons based on the difference between the conductance of the memristor and the resistance of the shared resistor.

[0147] The conductance of a memristor exhibits a non-linear change when receiving electrical pulses. Utilizing this characteristic, each neuron maps its weights to the difference between the memristor's conductance and the resistance of the shared resistor in the differential circuit. The value of the shared resistor can be initialized, with the assigned value ranging between the memristor's minimum and maximum conductance.

[0148] Step S300: Update the neuron weights by applying a single-step pulse to obtain the update direction of the neuron weights;

[0149] This step is the same as that in Example 2, and will not be repeated here.

[0150] Step S400: When the weight update direction is positive, apply a positive single-step pulse to the memristor to update the neuron weights again.

[0151] Step S500: When the weight update direction is negative, apply a negative single-step pulse to the memristor to update the neuron weights again.

[0152] In step S600, when the weight update iterations of all neurons reach the convergence condition, the iteration ends and the training is completed.

[0153] After obtaining the weight update direction in step S300, the difference between steps S400 and S500 and those in Embodiment 2 is that in this scheme, only one single-step pulse needs to be applied to one memristor in the differential circuit, while in Embodiment 2, a single-step pulse can be applied to one memristor in the differential circuit, or simultaneously to both memristors in the differential circuit. The former requires selectively applying pulses to both ends of the memristor, that is, applying pulses to memristors with larger conductance changes is necessary to achieve convergence. Therefore, compared to Embodiment 2, this scheme simplifies the neural network hardware circuit and power consumption, but the update flexibility is relatively weaker.

[0154] As the first specific implementation scheme of Example 3:

[0155] Step S101: Bind the memristors and shared resistors that make up each differential circuit to form a neuron;

[0156] Step SA: Randomly initialize the resistance value R of the shared resistor, where Gmin≤R≤Gmax, where Gmax is the maximum conductance of the memristor and Gmin is the minimum conductance of the memristor. It should be noted that the resistance value of the shared resistor in each row is randomly set by the system. Once it is set for the first time, it will not change during subsequent training. Therefore, it can be considered a constant value. However, the resistance value of the shared resistor in each row may be different.

[0157] Step S201: Map the weight W of the neuron based on the difference between the conductance W' of the memristor and the resistance R of the shared resistor;

[0158] The weight W of the neuron is: W' – R.

[0159] Step S300: Update the neuron weights by applying a single-step pulse to obtain the update direction of the neuron weights;

[0160] Step S400: When the weight update direction is positive, apply a positive single-step pulse to the memristor to update the neuron weights again.

[0161] Step S500: When the weight update direction is negative, apply a negative single-step pulse to the memristor to update the neuron weights again.

[0162] In step S600, when the weight update iterations of all neurons reach the convergence condition, the iteration ends and the training is completed.

[0163] As a second specific implementation scheme of Example 3:

[0164] Step S101: Bind the memristors and shared resistors that make up each differential circuit to form a neuron;

[0165] Step SA, predetermine the conductance value R of the shared resistance in all neurons, where R is the maximum conductance value G of the memristor. max With minimum conductivity G min Half of the sum, R = (G max / 2+G min / 2);

[0166] It should be noted that the difference from the first implementation scheme is that the shared resistors in each row here have the same resistance value, and are all the maximum conductance value G of the memristor. max With minimum conductivity G min Half of the sum.

[0167] Step S201: Map the weight W of the neuron based on the difference between the conductance W' of the memristor and the resistance R of the shared resistor;

[0168] The neuron's weight W = W' – R = W' - (G) max / 2+G min / 2).

[0169] Step S300: Update the neuron weights by applying a single-step pulse to obtain the update direction of the neuron weights;

[0170] Step S400: When the weight update direction is positive, apply a positive single-step pulse to the memristor to update the neuron weights again.

[0171] Step S500: When the weight update direction is negative, apply a negative single-step pulse to the memristor to update the neuron weights again.

[0172] In step S600, when the weight update iterations of all neurons reach the convergence condition, the iteration ends and the training is completed.

[0173] Based on this algorithm, four different weight update schemes are proposed, using two different circuit structures. These are named Scheme 1, Scheme 2, Scheme 3, and Scheme 4, and will be described in detail below:

[0174] Scheme 1 and Scheme 2 based on 1T1R-1T1R circuit structure:

[0175] Schemes 1 and 2 are based on a differential circuit structure of 1T1R-1T1R (1-transistor-1-memristor+1-transistor-1-memristor). However, some modifications have been made. In the original "Scheme 0", the two memristors (w1 and w2) in each group of differential devices are physically identical. However, Schemes 1 and 2 will limit their area dimensions to be different from the beginning of the circuit design (for example, the area of ​​w1 can be designed to be twice the area of ​​w2). Therefore, the intrinsic conductance of device w1 will (with a high probability) be greater than the conductance of w2. In this case, device w = w1 - w2 is set as the positive weight mapping of the network, and device w = w2 - w1 is set as the negative weight mapping of the network. It is worth noting that in the above settings, the weight values ​​in the positive weight mapping are only likely to be positive, and are not guaranteed to be positive. The same applies to the negative weight mapping. This introduces a certain degree of probabilistic update behavior, which can be used as a compensation mechanism for network weight updates. This probability is named Probability 1.

[0176] After the above settings are established, the memristor network is trained using the learning algorithm proposed in this application. Although Scheme 1 and Scheme 2 have the same circuit architecture, they differ in that the conductance of the two devices constituting the differential pair changes differently during each weight update.

[0177] Specifically, in Scheme 1, the two devices constituting the difference pair change in the same direction during each update of the network weights, but their magnitudes are unequal. This means it cannot be guaranteed that the new weight change obtained after each pulse application will match the expected weight change. However, appropriate parameter design (mainly the area parameters of devices w1 and w2) can ensure that the above requirement is met with a high probability. For example, when the area of ​​device w1 is twice that of device w2, the magnitude of the change in device w1 is likely to be greater than that of device w2. Therefore, the weight values ​​in the network are likely updated according to the requirements given by the algorithm, meaning the weight update in the network is likely successful. This setting introduces another probabilistic update behavior, which can serve as another compensation mechanism for network weight updates. This probability is named Probability 2.

[0178] The update method for Scheme 1 is shown in Table 1 below:

[0179] Table 1. Weight update steps for Scheme 1

[0180]

[0181]

[0182] The weight update method for Scheme 1 is as follows:

[0183] Step 1: Randomly assign two weights to the network: "positive weights" and "negative weights". For "positive weights", the differential pair in the hardware circuit is designed as "w = w1 - w2". Conversely, for "negative weights", the differential pair in the hardware circuit is designed as "w = w2 - w1". Devices w1 and w2 have different device sizes, and the device size of w1 is larger than that of w2.

[0184] Step 2: Randomly initialize network weights. Because the device area of ​​device w1 is larger than that of device w2, the maximum conductance value Gmax1 of device w1 is greater than the maximum conductance value Gmax2 of device w2. Assuming the range of w1 is [Gmin1, Gmax1] and the range of w2 is [Gmin2, Gmax2], then Gmax1 > Gmax2.

[0185] Step 3: Update the network weights using a single-step pulse update method. The standard SGD algorithm is used to obtain the update direction of the weights, determining the sign of Δw. Specifically: If Δw > 0, a positive pulse is applied to the two devices w1 and w2 in the difference pair representing "positive weights" in the network, and a negative pulse is applied to the two devices w1 and w2 in the difference pair representing "negative weights" in the network. If Δw < 0, a negative pulse is applied to the two devices w1 and w2 in the difference pair representing "positive weights" in the network, and a positive pulse is applied to the two devices w1 and w2 in the difference pair representing "negative weights" in the network.

[0186] Step 4: Determine if all iterations have been completed. If so, stop the learning process.

[0187] Scheme 2 has the same circuit structure as Scheme 1. The difference between Scheme 2 and Scheme 1 lies in the "third step" mentioned above. The steps for updating the weights in Scheme 2 are as follows:

[0188] Step 1: Randomly assign two weights to the network: "positive weights" and "negative weights". "Positive weights" mean that the differential pairs in the hardware circuit are designed as "w = w1 - w2" and "negative weights" mean that the differential pairs in the hardware circuit are designed as "w = w2 - w1", where the device size of w1 is larger than the device size of w2.

[0189] Step 2: Randomly initialize network weights. The conductance of device w1 ranges from [Gmin1, Gmax1], and the conductance of w2 ranges from [Gmin2, Gmax2]. Similar to Scheme 1, Gmax1 > Gmax2.

[0190] Step 3: Update the network weights using a single-step pulse update method. The standard SGD algorithm is used to obtain the update direction of the weights, determining the sign of Δw. Specifically: If Δw > 0, a positive pulse is applied to device w1 in the difference pair representing "positive weights" in the network, and a negative pulse is applied to device w2. Conversely, a negative pulse is applied to device w1 in the difference pair representing "negative weights" in the network, and a positive pulse is applied to device w2. If Δw < 0, a negative pulse is applied to device w1 in the difference pair representing "positive weights" in the network, and a positive pulse is applied to device w2. Conversely, a positive pulse is applied to device w1 in the difference pair representing "negative weights" in the network, and a negative pulse is applied to device w2.

[0191] Step 4: Determine if all iterations have been completed. If so, stop the learning process.

[0192] Schemes 3 and 4 are based on the 1T1R-1R circuit structure:

[0193] Following Schemes 1 and 2, we subsequently proposed two more schemes (Schemes 3 and 4) based on different circuit structures, further reducing the circuit complexity and area of ​​the memristor neural network chip. Compared to the 1T1R-1T1R circuit structure of Schemes 1 and 2, Schemes 3 and 4 adopt a differential circuit structure based on 1T1R-1R (1-transistor-1-memristor+1-resistor), which significantly reduces the circuit area. This is because the 1T1R-1T1R structure requires twice the number of components, while different differential pairs in the 1T1R-1R structure can share the same resistor. Theoretically, this circuit structure can reduce the circuit area consumption by 50% for the same network. Meanwhile, the weight update schemes of Schemes 3 and 4 still use the learning methods from Schemes 1 and 2, which can improve the performance, area, and power consumption of the memristor neural network chip.

[0194] The difference between Scheme 3 and Scheme 4 lies in the initial method of the resistance values. Specifically, in Scheme 3, all resistance values ​​(corresponding to device w2) are randomly initialized at the beginning, and these values ​​remain unchanged (equal to the initial values) throughout the weight update process, which is very similar to the original "Scheme 0". For Scheme 4, all resistance values ​​are fixed throughout the weight update process, and their values ​​are fixed at (G max / 2+G min / 2).

[0195] The weight update schemes in the four different modes described above differ from the weight update method in "Scheme 0" primarily because they do not require calculating the specific number of voltage pulses. What we need to do is determine the sign of the weight update amount Δw during training. This simple learning rule enables the gimbal network to update its parameters efficiently, thereby achieving self-optimization capabilities.

[0196] Below, we will summarize the four weight update schemes. The implementation details of the four schemes are shown in Table 2 below:

[0197] Table 2 Four weight update schemes

[0198]

[0199] We trained the classic LeNet-5 convolutional neural network using the algorithm proposed in this application and other memristor-based algorithms, and tested the recognition accuracy on the MNIST handwritten digit dataset, comparing the network performance under different algorithms.

[0200] From above Figure 12 As can be seen, our algorithm achieves a recognition accuracy comparable to the standard linear stochastic gradient descent algorithm, with a similar learning rate. It also significantly surpasses the recognition accuracy obtained using the nonlinear SGD algorithm, and its convergence speed is faster.

[0201] Furthermore, we conducted a comprehensive and detailed comparison of the network performance using four different weight update schemes. We tested experimental results for 49 typical combinations of Ap (LTP process nonlinearity) and Ad (LTD process nonlinearity) values ​​under the four schemes. Ap ranged from 0 to 6, and Ad ranged from 0 to -6, as detailed below. Figure 13 As shown.

[0202] From above Figure 13 We can observe that Scheme 1 exhibits the best network performance, while Scheme 2 shows worse performance. This is because in Scheme 2's weight update method, the weight update amount is obtained from the difference between the two devices in the difference pair after updating in opposite directions. This will likely result in an actual update amount greater than the amount required by the algorithm. This situation is very similar to choosing an excessively large learning rate in the stochastic gradient descent algorithm. We know that in stochastic gradient descent, if the learning rate is too large, it will cause the network to oscillate around the extreme point, resulting in difficulty in convergence. Scheme 2 is similar; in each iteration, the excessively large weight update amount causes the weights to oscillate around the optimal value. However, unlike the stochastic gradient descent algorithm, due to the algorithm used in this paper, through continuous iterative training, Scheme 2 can gradually bring the weights closer to the optimal value and eventually stop near it, achieving near-convergence. Figure 12 In Scheme 2, with Ap = 1 and Ad = -1, the network's recognition accuracy exceeds 90%, which proves that the network trained using Scheme 2 is convergent.

[0203] Schemes 3 and 4 employ a 1T1R-1R differential pair architecture designed to reduce circuit area. In both weight update schemes, the second weight (w2) in the differential pair is never involved. Therefore, their performance is lower than Scheme 1. Nevertheless, they still exhibit good network performance. It should be noted that Scheme 4 has higher accuracy than Scheme 3 because it has a uniform and symmetrical weight distribution. Gmax = 1, Gmin = 0, and the number of weight states Pmax is 11. In Scheme 4, w2 = (Gmax + Gmin) / 2 = 0.5, therefore the possible weight values ​​for each differential pair (W = w1 - w2 = w1 - 0.5) are in the set W = -0.5, -0.4, -0.3, -0.2, -0.1, 0, 0.1, 0.2, 0.3, 0.4, 0.5. In Scheme 3, the distribution of W is typically asymmetrical. For example, if w2 is (randomly) initialized to w2 = Gmax (or Gmin), the weights of the difference pair (W = w1 - w2) will always be negative (or positive), which is detrimental to high-precision online learning.

[0204] Table 3 below shows the specific recognition accuracy values ​​for seven typical Ap and Ad combinations under four weight update schemes:

[0205] Table 3. Accuracy of LeNet-5 networks trained using different schemes at typical Ap and Ad values.

[0206]

[0207]

[0208] The table above demonstrates that our algorithm is well-suited for memristor networks, and all four weight update schemes based on this algorithm exhibit good network performance. Even scheme 2, with typical nonlinearity Ap and Ad values ​​of ±1, still achieves a network recognition accuracy as high as 92.42%. This proves that the algorithm largely suppresses the weight update error caused by the nonlinear weight update characteristics of memristors.

[0209] Simulation results demonstrate that, compared to conventional learning schemes based on memristor neural networks, our algorithm effectively reduces weight update errors caused by the nonlinear characteristics of memristors, exhibiting better network performance. Furthermore, since the algorithm adopted in this patent does not require calculating the specific number of pulses corresponding to weight changes, it avoids complex peripheral circuits caused by redundant calculations, thus making it more hardware-friendly.

[0210] The above description of the embodiments is only used to provide a detailed introduction to the technical solutions of this application. However, the description of the above embodiments is only for the purpose of helping to understand the methods and core ideas of this application, and should not be construed as a limitation of this application. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application.

Claims

1. A method for training weights in a memristor convolutional neural network, characterized in that, A memristor array is constructed using several differential circuits to form a neural network hardware system, with each pair of devices constituting a differential circuit, and at least one of the two devices including a memristor; the method includes: The two devices that make up each differential circuit are bound together to form a neuron; The weights of neurons are mapped based on the difference in electrical parameters between the two devices; The direction of neuron weight update is obtained by updating neuron weights by applying single-step pulses; When the weight update direction is positive, a positive single-step pulse is applied to the memristor to update the neuron weights again; When the weight update direction is negative, a negative single-step pulse is applied to the memristor to update the neuron weights again; The training process ends when the weight updates of all neurons converge.

2. The memristor convolutional neural network weight training method according to claim 1, characterized in that, Both devices constituting the differential circuit include memristors; the parameters of the two memristors forming the differential circuit are different to ensure that a predictable trend of change in conductance is obtained after an electrical pulse is applied; the method includes: The two memristors that make up each differential circuit are bound together to form a neuron; The weights of neurons are mapped based on the difference in conductance between two memristors; The direction of neuron weight update is obtained by updating neuron weights by applying single-step pulses; When the weight update direction is positive, a positive single-step pulse is applied to the memristor with a large change in conductance to update the neuron weights again; When the weight update direction is negative, a negative single-step pulse is applied to the memristor with a large change in conductance to update the neuron weights again; The training process ends when the weight updates of all neurons converge.

3. The memristor convolutional neural network weight training method according to claim 2, characterized in that, Each differential circuit consists of two independent memristors; the step of mapping the neuron weights based on the difference in conductance between the two memristors includes: Half of the neurons are mapped to positive weights; the other half are mapped to negative weights. In the step of applying a positive single-step pulse to the memristor with a large change in conductance to update the neuron weights again when the weight update direction is positive: Memristors with large changes in conductivity are memristors with large changes in weight in the positive weighting mapping. In the step of applying a negative single-step pulse to the memristor with a large change in conductance to update the neuron weights again when the weight update direction is negative: Memristors with large changes in conductivity are memristors with large changes in weight in negative weighting mapping.

4. The memristor convolutional neural network weight training method according to claim 3, characterized in that, The differential circuit is composed of two memristors with different cross-sectional areas; the weight change of the first memristor is relatively larger than that of the second memristor after a single pulse is applied; the conductance of the first memristor is W1, and the conductance of the second memristor is W2. The steps for mapping positive weights include: The difference between the conductance value W1 of the first memristor and the conductance value W2 of the second memristor is used to map the positive weight of the neuron: W1-W2; The steps for mapping negative weights include: The negative weight of the neuron is mapped by the difference between the conductance value W2 of the second memristor and the conductance value W1 of the first memristor: W2-W1.

5. The memristor convolutional neural network weight training method according to claim 4, characterized in that, The step of applying a positive single-step pulse to memristors with large weight changes in the positive weight mapping when the weight update direction is positive includes: Apply a positive single-step pulse to both the first and second memristors, or A positive single-step pulse is applied to the first memristor, and a negative single-step pulse is applied to the second memristor.

6. The memristor convolutional neural network weight training method according to claim 4, characterized in that, The step of applying a negative single-step pulse to memristors with large weight changes in the negative weight mapping when the weight update direction is negative includes: Apply a negative single-step pulse to both the first and second memristors, or A negative single-step pulse is applied to the first memristor, and a positive single-step pulse is applied to the second memristor.

7. The memristor convolutional neural network weight training method according to claim 1, characterized in that, One of the two components constituting the differential circuit includes a memristor, and the other includes a shared resistor; prior to the step of updating the neuron weights by applying a single-step pulse to obtain the update direction of the neuron weights, the circuit further includes: The resistance value of the shared resistance in all neurons is randomly initialized to R, where R is a constant value; The conductance value of the memristor is W', and the step of mapping the weights of the neuron based on the difference in electrical parameters between the two devices includes: The neuron's weight is mapped based on the difference between the memristor's conductance and the initial value of the shared resistance. The neuron's weight W is: W' – R.

8. The memristor convolutional neural network weight training method according to claim 1, characterized in that, Another component constituting the differential circuit includes a shared resistor; prior to the step of updating the neuron weights by applying a single-step pulse to obtain the update direction of the neuron weights, the circuit further includes: The resistance value R of the shared resistance in all neurons is predetermined, where R is the maximum conductance of the memristor. With minimum conductivity Half of the sum, ; The conductance value of the memristor is W', and the step of mapping the weights of the neuron based on the difference in electrical parameters between the two devices includes: The neuron's weight is mapped based on the difference between the memristor's conductance and the initial value of the shared resistance. The neuron's weight W is: .

9. A memristor convolutional neural network weight training device, characterized in that, include: The memory stores the weight training program. The processor performs the steps of the method according to any one of claims 1 to 8 when running the weight training program.