A neural network optimization circuit

CN119026653BActive Publication Date: 2026-08-28INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411117434.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-14
Publication Date
2026-08-28
Estimated Expiration
2044-08-14

AI Technical Summary

Technical Problem

[0004]有鉴于此,本发明提供了一种神经网络优化电路,以解决神经网络对批量训练数据进行处理过程中,无法保证神经网络数据的稳定性的问题

Benefits of technology

[0033]本发明涉及电路技术领域一种神经网络优化电路,应用于目标神经网络,包括:第一转换单元,用于将目标训练数据转换为目标训练数据对应的目标直流电平信号;第二转换单元,用于将目标训练数据对应的标签数据转换为标注结果电平信号;第一乘法器,用于基于目标训练数据对应的目标直流电平信号和加法器输出信号作乘,输出目标训练数据对应的预测结果信号;误差计算单元,用于计算预测结果信号与目标训练数据的标注结果信号之间的MSE误差信号;第一微分器,用于接收MSE误差信号以及输出第一微分器输出信号;第二微分器,用于接收加法器输出信号,以及输出第二微分器输出信号;第二乘法器,用于接收第一微分器输出信号、第二微分器输出信号,以及输出第二乘法器输出信号;积分器,用于接收第二乘法器输出信号以及输出积分器输出信号;加法器,用于接收多个权重变化数据的多个权重变化信号、积分器输出信号以及输出加法器输出信号;其中,每个权重变化信号对应不同的权重值,每个权重变化数据表示权重变化数据对应的权重值在时间上的变化量。本发明将第二微分器从扰动源处后置到加法器输出的位置,可以保持电路的原有收敛方向,减小对扰动功率的要求,提高网络的稳定性。同时,同时训练多条目标训练数据,并对多条目标训练数据同时进行误差优化,高度并行、计算速度快,具备片上学习能力,训练过程没有开关信号、时钟信号等非连续信号,在严格的数学理论保证下实现快速收敛,全部使用模拟电路,扩展成本低。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119026653B_ABST
    Figure CN119026653B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of circuits, and discloses a neural network optimization circuit, which can simultaneously train multiple target training data and simultaneously optimize errors of the multiple target training data, has high parallelism, fast calculation speed, on-chip learning ability, no non-continuous signals such as switch signals and clock signals in the training process, realizes fast convergence under the guarantee of strict mathematical theory, uses analog circuits, and has low expansion cost. Furthermore, the second differentiator is placed from the disturbance source to the position of the output of the adder, the original convergence direction of the circuit can be maintained, the requirement for disturbance power is reduced, and the stability of the network is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of circuit technology, and more specifically to a neural network optimization circuit. Background Technology

[0002] In recent years, with the development of AI technology, the demand for computing power has risen rapidly, and improving optimization speed has become an important development direction. As a result, neural networks are being used more and more widely.

[0003] In related technologies, the stability of neural network data cannot be guaranteed during the processing of batch training data due to the influence of perturbation power. Summary of the Invention

[0004] In view of this, the present invention provides a neural network optimization circuit to solve the problem that the stability of neural network data cannot be guaranteed during the processing of batch training data.

[0005] According to a first aspect, embodiments of this disclosure provide a neural network optimization circuit applied to a target neural network, comprising:

[0006] The first conversion unit is used to convert the target training data into the target DC level signal corresponding to the target training data;

[0007] The second conversion unit is used to convert the label data corresponding to the target training data into a labeling result level signal;

[0008] The first multiplier is used to output the prediction result signal corresponding to the target training data based on the target DC level signal corresponding to the target training data and the adder output signal;

[0009] The error calculation unit is used to calculate the MSE error signal between the prediction result signal and the labeled result signal of the target training data;

[0010] The first differentiator is used to receive the MSE error signal and output the first differentiator output signal.

[0011] The second differentiator is used to receive the output signal of the adder and to output the output signal of the second differentiator.

[0012] The second multiplier is used to receive the output signal of the first differentiator and the output signal of the second differentiator, and to output the output signal of the second multiplier.

[0013] An integrator is used to receive the output signal of the squared multiplier and output the integrator's output signal.

[0014] The adder is used to receive multiple weight change signals of multiple weight change data, the integrator output signal, and the output signal of the adder;

[0015] Each weight change signal corresponds to a different weight value, and each weight change data represents the amount of change in the weight value corresponding to the weight change data over time.

[0016] In one optional implementation, the adder is a unidirectional adder.

[0017] The co-current adder is used to receive multiple weight change signals of multiple weight change data and DC level signals of the weight value corresponding to each weight change data, and to transmit the updated weight value of each weight change data to the first multiplier and the second differentiator. The multiple weight change signals of multiple weight change data are disturbance signals output by introducing noise sources.

[0018] In one optional implementation, the second differentiator is a co-directional differentiator; the co-directional differentiator is used to transmit the change in weight value over time corresponding to each weight change data output to the second multiplier to form a closed loop.

[0019] In one alternative implementation, the first multiplier is used to output the product of the input, a first numerical DC level signal, and the updated weight values ​​for each weight change data.

[0020] In one optional implementation, the error calculation unit includes: a subtractor, and a first differentiator that is an inverse differentiator;

[0021] The first subtractor is used to subtract the product of multiple weight change signals from the second numerical DC level signal of the true label corresponding to the labeled result signal of the target training data. After the difference is calculated, it is squared by the third multiplier and then differentiated by the inverse differentiator and the opposite number is taken. The DC level signal after the inverse differentiator represents the opposite number of the MSE differential.

[0022] The second multiplier is used to receive the inverse of the MSE differential and the change in weight value over time corresponding to each weight change data, and to provide the product result of the second multiplier as the weight modification amount to the integrator; the integrator and the amplifier together form an integrator with proportional amplification.

[0023] In one alternative implementation, the target neural network includes multiple different subnetworks, each subnetwork establishing a control chip, where each pin on the input side of the control chip corresponds to an input neuron, and each pin on the output side of the control chip corresponds to an output neuron.

[0024] Both the input and output neurons are DC level signals;

[0025] In two adjacent subnets, the output neurons of the front subnet are connected to the input neurons of the rear subnet;

[0026] When training the target training data in the target neural network, the MSE error signal is fed back to all subnetworks of the target neural network through wires.

[0027] In one alternative implementation, the input side of the target neural network is connected to the first conversion unit, and the output side of the target neural network is connected to the error calculation unit.

[0028] In one alternative implementation, the neural network optimization circuit further includes:

[0029] The computer is used to send digital signals corresponding to the target training data to the first conversion unit, send digital signals corresponding to the label data of the target training data to the second conversion unit, and receive prediction result signals corresponding to the target training data.

[0030] In one alternative implementation, the target neural network modifies the weight values ​​corresponding to each weight change data while training the target training data.

[0031] In one alternative implementation, the target neural network includes: a Transformer neural network, a CNN convolutional neural network, an RNN recurrent neural network, or a fully connected neural network.

[0032] The technical solution of this invention has the following advantages:

[0033] This invention relates to a neural network optimization circuit in the field of circuit technology, applied to a target neural network, comprising: a first conversion unit for converting target training data into a target DC level signal corresponding to the target training data; a second conversion unit for converting label data corresponding to the target training data into a labeled result level signal; a first multiplier for multiplying the target DC level signal corresponding to the target training data and the output signal of the adder, and outputting a prediction result signal corresponding to the target training data; an error calculation unit for calculating the MSE error signal between the prediction result signal and the labeled result signal of the target training data; and a first differentiator for receiving the MSE. The system comprises: an error signal and an output signal from a first differentiator; a second differentiator, used to receive the output signal from an adder and output its own second differentiator signal; a second multiplier, used to receive the output signals from the first and second differentiators and output its own second multiplier signal; an integrator, used to receive the output signal from the second multiplier and output its own integrator signal; and an adder, used to receive multiple weight change signals, the integrator's output signal, and output its own adder signal for multiple weight change data. Each weight change signal corresponds to a different weight value, and each weight change data represents the change in the weight value corresponding to the weight change data over time. This invention places the second differentiator after the disturbance source at the adder output position, which can maintain the original convergence direction of the circuit, reduce the requirement for disturbance power, and improve network stability. Simultaneously, it trains multiple target training data simultaneously and performs error optimization on multiple target training data at the same time. It is highly parallel, has a fast computing speed, and has on-chip learning capabilities. The training process does not have discontinuous signals such as switching signals and clock signals. It achieves fast convergence under the guarantee of strict mathematical theory. It uses all analog circuits, resulting in low expansion costs. Attached Figure Description

[0034] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0035] Figure 1 This is a schematic diagram of the neural network optimization circuit training process according to an embodiment of the present invention;

[0036] Figure 2 This is a schematic diagram of a scenario where the target neural network is a CNN convolutional neural network according to an embodiment of the present invention;

[0037] Figure 3 This is a schematic diagram of the training process and structural framework of a neural network optimization circuit according to an embodiment of the present invention;

[0038] Figure 4 This is a circuit diagram of the target neural network during training with batch data according to an embodiment of the present invention;

[0039] Figure 5 These are schematic diagrams of the reasoning process and structural framework of the neural network optimization circuit according to embodiments of the present invention, and schematic diagrams of the modular layout of the analog circuit neural network according to embodiments of the present invention.

[0040] Figure 6 This is a schematic diagram of the specific circuit structure of the weighting circuit according to an embodiment of the present invention;

[0041] Figure 7 This is a schematic diagram illustrating a three-layer 2*2*2 neural network with activation functions built in Simulink according to an embodiment of the present invention;

[0042] Figure 8 This is a schematic diagram of the loss curves for training two datasets according to an embodiment of the present invention;

[0043] Figure 9 This is a schematic diagram illustrating the changes in the solution during the training process according to an embodiment of the present invention;

[0044] Figure 10 This is a schematic diagram illustrating how the circuit according to an embodiment of the present invention completes the XOR task within 10 ns. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0046] In related technologies, some neural networks, while capable of training batches of training data, may also experience weight changes due to batch weight variations. These weight changes are affected by disturbance power, making it impossible to guarantee the stability of the neural network data.

[0047] In view of this, according to embodiments of the present invention, such as Figure 1 The diagram illustrates a neural network optimization circuit applied to a target neural network. This circuit includes: a first conversion unit 11, a second conversion unit 12, a first multiplier 13, a first differentiator 14, a second differentiator 15, a second multiplier 16, an adder 17, an integrator 18, and an error calculation unit 19. Figure 1The neural network optimization circuit shown is the basic training unit in this embodiment of the present disclosure, namely the weights of the neural network. Figure 1 In the diagram, the brown dashed line represents any target neural network, which can be a Transformer neural network, a CNN convolutional neural network, an RNN recurrent neural network, or a fully connected neural network.

[0048] In practical design, depending on the type of disturbance and the way the disturbance signal is introduced, the basic units of analog circuit neural networks can be divided into several different basic structures. The neural network optimization circuit in this embodiment is suitable for scenarios involving the processing of batch training data, and has the widest applicability, allowing for the expansion of networks sharing weights, such as batch processing and attention mechanisms. Figure 1 The neural network optimization circuit shown is suitable for scenarios where the target neural network is a Transformer neural network, a CNN convolutional neural network, or an RNN recurrent neural network. For example... Figure 2 As shown, the target neural network is a CNN convolutional neural network, where each weight participates in the calculation of multiple data points simultaneously.

[0049] In this embodiment of the disclosure, in Figure 1 In the example, a two-layer 2x2 fully connected neural network is used. The two circles on the left are the input neurons, with DC level as the input signal and DC level as the output signal. The input and output neurons are fully connected by four weights. (Solid lines connect the neurons on the left and right sides). Figure 1 The analog circuitry on the lower middle side illustrates the internal structure of the weights.

[0050] The first conversion unit 11 is used to convert the target training data into the target DC level signal corresponding to the target training data.

[0051] Specifically, such as Figure 3 As shown, the first conversion unit 11 in this embodiment can be a D / A digital-to-analog converter, which converts the target training data into a target DC level signal corresponding to the target training data. The target training data can be different types of data sample sets, and each data point in the data sample set is converted into a target DC level signal. For example, the target training data is represented by M, where M = {a1, a2, a3, ... a...}. n The method converts each of the n data points in the data sample set into a target DC level signal. In this embodiment, the target training data is also the input training data. Converting the target training data into the corresponding target DC level signal and using it as the input signal for the neural network helps ensure the continuity of the analog system signal, thereby enabling the circuit to train the target neural network at high speed.

[0052] The second conversion unit 12 is used to convert the label data corresponding to the target training data into a labeling result level signal.

[0053] Specifically, in Figure 3 In this embodiment, the second conversion unit 12 can be a D / A digital-to-analog converter, which converts the label data corresponding to the target training data into a labeling result level signal. The labeling result signal of the target training is a label data level signal. Figure 1 In this context, both the input data level signal x and the label data level signal y are provided. For example, the target training data is represented by M, where M = {a1, a2, a3, ... a...}. n}, then all n data points in the data sample set will be converted into target data level signal x and tag data level signal y.

[0054] Among them, Figure 1 In the first multiplier 13, the prediction result signal corresponding to the target training data is output based on the target DC level signal corresponding to the target training data and the adder output signal.

[0055] Specifically, in Figure 1 In the process of reasoning the target neural network, the output value is obtained by multiplying it with the weight level signal at the output of adder 17. Inference can then be completed. For example, the target DC level signal x corresponding to the target training data is multiplied by the adder output signal (weight value level signal) to output the prediction result signal corresponding to the target training data.

[0056] Among them, Figure 1 In the middle, the error calculation unit 19 is used to calculate the MSE error signal between the prediction result signal and the labeled result signal of the target training data.

[0057] Specifically, in Figure 1 In this circuit, the error calculation unit 19 and the first differentiator 14 can be MSE differentiating units. These MSE differentiating units receive the DC level of the tag data output from the PC via a D / A digital-to-analog converter. Internally, the MSE differentiating unit calculates the inverse of the MSE loss differential, outputs it as a wire, and then feeds it back to the analog circuit neural network. This forms a closed loop, and the weights of the analog circuit neural network converge in the direction of minimizing the error, thus completing the training.

[0058] Among them, Figure 1 In the middle, the first differentiator 14 is used to receive the MSE error signal and output the first differentiator output signal.

[0059] Specifically, in Figure 1In the training of the target neural network, both the input data level signal x and the label data level signal y are given simultaneously, and the circuit's output signal... The MSE loss is calculated with y and inverted. Then, after passing through the first differentiator 14, it is connected to the second multiplier 16 to represent the negative number of the change in MSE loss.

[0060] Among them, Figure 1 In the first part, the second differentiator 15 is used to receive the adder output signal and output the second differentiator output signal. The adder output signal is generated based on multiple weight change signals and weight DC level signals derived from multiple weight change data.

[0061] Specifically, in the embodiments of this disclosure, the multiple weight change signals of multiple weight change data represent weight perturbation noise signals. For example, when the target neural network is training with batch data, there are N noise signals, which are described as multiple weight change signals. If the target neural network has N weight change data, it means that there are N corresponding noise signals, and each noise signal represents the amount of change of the weight value over time.

[0062] like Figure 4 The diagram shows a circuit diagram of the target neural network during training with a batch of data.

[0063] Among them, Figure 1 In this circuit, the second multiplier 16 receives the output signals of the first differentiator and the second differentiator, and outputs the second multiplier output signal. The integrator 18 receives the second multiplier output signal and outputs the integrator output signal. The adder 17 receives multiple weight change signals of multiple weight change data, the integrator output signal, and outputs the adder output signal. Each weight change signal corresponds to a different weight value, and each weight change data represents the change in the weight value corresponding to the weight change data over time.

[0064] Specifically, in Figure 1 In the process, during inference, the target neural network multiplies the given input data level signal x with the weight value level signal at the output of adder 17 to obtain the output value. Inference can then be completed. During training, both the input data level signal x and the tag data level signal y are given simultaneously. The circuit's output signal... The MSE loss is calculated by inverting the derivative of y, and then passed through the first differentiator 14 and connected to the second multiplier 16, representing the negative of the change in MSE loss. The other end of the second multiplier 16 is connected to the output of the adder 17 via the second differentiator 15. Simultaneously, noise is connected to the input of the adder 17, and the other input of the adder 17 is the DC level signal corresponding to the weight of each weight change data. Therefore, the output of the adder 17 is the weight value (W+ΔW) of each weight change data after perturbation and updating. The weight DC level signal is also the output of the integrator 18, meaning the change in weight value is determined by the product of the two inputs of the bottom-most second multiplier 16. One quantity, the negative of the differential of the MSE loss, represents the trend of loss change under the current perturbation; only the positive or negative sign matters, indicating whether the error has decreased or increased. The other quantity represents the amount of weight change. The product of the two is the amount of weight change that reduces the error. For example, when the weights change, if the negative of the differential value of the MSE loss is positive, it indicates that the error is decreasing, and the weights should shift in that direction. That is, the weight change should be a positive number multiplied by the original weight change. If the negative of the differential value of the MSE loss is negative, it indicates that the error is increasing, and the weights should shift in the opposite direction. That is, the weight change should be a negative number multiplied by the original weight change. This circuit ensures that the weights always shift in the direction that reduces the error, thus completing the training of the neural network.

[0065] The neural network optimization circuit in this embodiment improves the optimization speed of neural network training equipment while simultaneously possessing the capabilities of fast inference training, low expansion cost, high parallelism, in-memory computing, and on-chip learning. Furthermore, the neural network optimization circuit in this embodiment is entirely implemented using analog circuitry, eliminating the need for multiple computer calculations of weight error values. It also eliminates discontinuous signals such as switching signals and clock signals during the entire training process, achieving rapid convergence under strict mathematical guarantees. The convergence speed is faster than existing optimizers (Simulink's simulation training speed is 147,900,000 times that of GPUs).

[0066] In one optional implementation, the target neural network in this embodiment includes multiple different subnetworks. Each subnetwork establishes a control chip, with each pin on the input side of the control chip corresponding to an input neuron and each pin on the output side corresponding to an output neuron. Both the input and output neurons are DC level signals. In adjacent subnetworks, the output neuron of the preceding subnetwork is connected to the input neuron of the following subnetwork. When training the target neural network with target training data, the MSE error signal is fed back to all subnetworks of the target neural network via wires.

[0067] Specifically, such as Figure 5Part B shown illustrates the structural framework of the circuit (Analog neural network) for inference in this embodiment. The left side of the circuit represents the input terminal, with each pin corresponding to an input neuron, using a DC level as the input signal. The right side of the circuit represents the output terminal, with each pin corresponding to an output neuron, using a DC level as the output signal. The input signal is provided by a computer-controlled D / A digital-to-analog converter. The output signal is converted into a digital signal by an A / D analog-to-digital converter and transmitted to the computer.

[0068] exist Figure 5 Part D of the diagram illustrates the modular layout within an analog circuit neural network. When the network is large-scale or heterogeneous, it can be divided into different subnetworks, each with its own dedicated chip. In the forward propagation section, the output neurons of the preceding subnetworks are connected to the input neurons of the following subnetworks. The training section then feeds back the final differential error to all subnetworks to perform the training.

[0069] Therefore, the neural network optimization circuit in this embodiment simultaneously trains multiple target training data and performs error optimization on multiple target training data simultaneously. It is highly parallel, has a fast computation speed, and possesses on-chip learning capabilities. The training process does not involve discontinuous signals such as switching signals or clock signals. It achieves rapid convergence under the guarantee of strict mathematical theory, uses all analog circuits, and has low expansion costs. Furthermore, by moving the second differentiator from the disturbance source to the adder output position, the original convergence direction of the circuit can be maintained, the requirement for disturbance power can be reduced, and the stability of the network can be improved.

[0070] In this embodiment, both the input and output neurons are DC level signals. Circuit storage and computation are implemented using weight levels, and since they share the same DC level signal, it belongs to an in-memory computing architecture. Furthermore, the DC level signal helps ensure the continuity of the analog system signal, thereby enabling the circuit to train the target neural network at high speed.

[0071] In one alternative implementation, in Figure 1 In the process, the input side of the target neural network is connected to the first conversion unit 11, and the output side of the target neural network is connected to the error calculation unit 19.

[0072] In one alternative implementation, such as Figure 3 and Figure 5 As shown, it also includes a computer 20 and a third conversion unit 130. The computer 20 is used to send digital signals corresponding to the target training data to the first conversion unit 11, send digital signals corresponding to the label data of the target training data to the second conversion unit 12, and receive prediction result signals corresponding to the target training data. The third conversion unit 130 is used to convert the prediction result signals corresponding to the target training data into digital signals and transmit them to the computer 20.

[0073] Specifically, the first conversion unit 11 and the second conversion unit 12 mentioned above are both D / A digital-to-analog converters, and the third conversion unit 130 is an A / D analog-to-digital converter. Figure 3 The computer 20 in this embodiment simultaneously provides input data and label data. The input data is provided to the first conversion unit 11, and the label data is provided to the second conversion unit 12. In this embodiment, the computer only provides data and does not participate in the weight error calculation. During the weight error calculation process, the data is processed by... Figure 1 The circuit shown is implemented using analog circuits. Therefore, the neural network optimization circuit in this embodiment is implemented entirely using analog circuits, without the participation of computers or other digital circuits, which can take advantage of the high speed of analog computing.

[0074] In one optional implementation, the target neural network simultaneously modifies the weight values ​​corresponding to each weight change data while training the target training data. Therefore, the neural network optimization circuit in this embodiment has a high degree of parallelism, and all weights are modified simultaneously during neural network training, resulting in low time complexity, which is only 1. This time complexity of only 1 means that all weight units are optimized simultaneously, and for each weight, the calculation process completes one optimization of the neural network in only one operation.

[0075] In one alternative implementation, in Figure 1In this circuit, adder 17 is a co-directional adder, used to receive multiple weight change signals of multiple weight change data and the DC level signal of the corresponding weight value of each weight change data, and to transmit the different weight values ​​corresponding to each weight change data to the first multiplier 13 and the second differentiator 15. The multiple weight change signals of multiple weight change data are disturbance signals output by introducing noise sources. The second differentiator 15 is a co-directional differentiator; it is used to transmit the time-varying amount of the weight value corresponding to each weight change data to the second multiplier 16 to form a closed loop. The first multiplier 13 is used to output the product result of the input being a first value (10V) DC level signal and the time-varying amount of the weight value corresponding to each weight change data. Error calculation unit 19 includes: a subtractor; a first differentiator 14 is an inverse differentiator; the first subtractor is used to subtract the product result of multiple weight change signals from the second value (6V) DC level signal of the real label corresponding to the labeled result signal of the target training data, and then perform square calculation by a third multiplier, and then perform differentiation and take the opposite number by the inverse differentiator. The DC level signal after the inverse differentiator represents the opposite number of the MSE differential; the second multiplier 16 is used to receive the opposite number of the MSE differential and the change in weight value corresponding to each weight change data over time, and to provide the product result output by the second multiplier 16 as the weight modification amount to the integrator 18; the integrator 18 and the amplifier together form an integrator with proportional amplification.

[0076] Further use of multisim Figure 1 The neural network optimization circuit is constructed using a single-weight circuit. Figure 6 As shown, the first element in the first row is a multiplier, that is... Figure 1 The first multiplier in the first row has a fixed 10V DC input representing the input signal x. The other input comes from the device in the second row, representing the weighted result after perturbation. The output is the product of the two signals. The second element in the first row is a subtractor. One input to this subtractor comes from the first multiplier, and the other input comes from the second conversion unit 12. It subtracts the 6V level signal of the actual label y, and then performs a square calculation via the third multiplier. The second and third elements in the first row form the circuit for calculating the MSE. The fourth element in the first row is an inverting differentiator, i.e. Figure 1The first differentiator corresponds to the operation of differentiating and then taking the opposite. Therefore, after the signal passes through the fourth element, the DC level signal represents the opposite of the derivative of the MSE. The fifth element in the first row is the second multiplier, receiving the opposite of the derivative of the MSE at one end and the weight change at the other. The product is provided as the weight modification to the integrator in the sixth element in the first row. The sixth element in the first row and the first element in the second row together form an integrator with proportional amplification, and the first element in the second row is an amplifier. Therefore, the output of the first element in the second row is the DC level of the weight signal. The second element in the second row and the third element in the second row form a co-inverting adder, with the weight DC level signal at one input and the added noise source at the other. After passing through the co-inverting adder, the signal represents the weight after introducing noise. This weight is passed to the first multiplier in the first row (mentioned above) and to the fourth element in the second row (the second differentiator). The fourth and fifth elements in the second row form a co-inverting differentiator, which is passed as the weight change to the second multiplier in the fifth element in the first row, thus forming a closed loop. This circuit has been verified through simulation to achieve weight convergence.

[0077] Furthermore, there are several implementations in the circuit diagram that are not unique. These will be listed below. For the non-integral circuit with proportional amplification formed by the sixth element in the first row and the first element in the second row, this part can achieve both integration and proportional amplification by modifying the parameters of the surrounding circuit components of the sixth element in the first row and connecting the signal to the non-inverting input of the op-amp. This eliminates the need for the first element in the second row. For the non-integral adder formed by the second and third elements in the second row, the third element in the second row can also be eliminated by connecting the signal to the non-inverting input of the adder. For the non-integral differentiator formed by the fourth and fifth elements in the second row, the fifth element in the second row can also be eliminated by connecting the signal to the non-inverting input of the differentiator. Therefore, this circuit is... Figure 1 The conceptual structure shown is a specialized form, and its specific implementation is not unique.

[0078] Furthermore, in Figure 6The circuit functions are described in order of core components. The 10V signal to the left of the first component in the first row is the input data. This data is multiplied by the weights in the first multiplier to obtain the neural network output y. The difference between the neural network output and the 6V monitoring signal is calculated, then processed by a subtractor and a third multiplier to calculate the MSE loss. The result is then processed by a first differentiator (inverting differentiator) to obtain the negative of the differential MSE, which serves as the input to the second multiplier. This second multiplier also receives the input of the weight change. The product of these two values ​​represents the weight change that minimizes the error loss. For example, when the weights change, if the negative of the differential MSE loss is positive, it indicates that the error is decreasing, and the weights should shift in that direction. That is, the weight change should be a positive number multiplied by the original weight change. If the negative of the differential MSE loss is negative, it indicates that the error is increasing, and the weights should shift in the opposite direction. That is, the weight change should be a negative number multiplied by the original weight change. This circuit ensures that the weights always shift in the direction that reduces the error, thus completing the training of the neural network. The weight modification amount is integratored to obtain the final weight DC level signal. Then, a noise source perturbation signal is added to the weight DC level signal to become the perturbed weight signal used in the calculation. Simultaneously, the perturbed weight signal is differentiated by a second differentiator to provide the position of the required weight change amount in the preceding calculation.

[0079] Therefore, after the circuit is powered on, the weighted DC level signal is immediately modified until the circuit's output value y is always equal to the monitoring signal, at which point the circuit reaches a steady state and the training is completed.

[0080] In a specific example, the neural network optimization circuit in this disclosure embodiment can achieve fast inference training. Fast inference refers to how long it takes for the neural network to obtain a stable output after being given an input. Since the weights do not change during the inference phase, it is equivalent to each weight unit only performing multiplication calculations. Under ideal device conditions, the entire neural network inference time is only constrained by the speed of light; under non-ideal device conditions, it is affected by device latency. It can generally be completed in nanoseconds or sub-nanosecond time. The theoretical guarantee of fast training is that the weight training effect is equivalent to backpropagation (BP), and the direction of weight modification is the unbiased estimate of BP. When external input data and label data are given, the weight levels of the neural network will converge, and the convergence direction is the negative direction of the gradient of the weights relative to the loss function, that is, the weight modification direction of BP (backpropagation). Therefore, as long as a neural network can be trained using BP, it can theoretically be trained using the circuit proposed in this disclosure embodiment, and the training results are consistent. If the training problem is a convex problem, the circuit will definitely converge to the optimal solution. The unbiased theoretical proof process is as follows:

[0081] The neural network optimization circuit in this embodiment of the disclosure, in Figure 1In this case, the second differentiator 15 is moved from the disturbance source to the output position of the adder 17.

[0082]

[0083] Its convergence is proven as follows:

[0084]

[0085] make

[0086]

[0087] The significance of this operation is that it maintains the original convergence direction of the circuit, reduces the requirement for disturbance power, and improves the stability of the network. Regarding rapid training capability, since the system in an analog circuit is a continuous system, the learning rate can be considered infinitely small. According to the convergence theorem, the continuous convergence of the circuit can be guaranteed. Furthermore, because the electric field is established at the speed of light, the iteration speed (number of iterations per unit time) can be considered infinitely large when the circuit size is sufficiently small. Therefore, under ideal device conditions, the circuit can theoretically achieve extremely fast training speeds.

[0088] In the embodiments disclosed herein, such as Figure 7 As shown, this is a three-layer 2x2x2 neural network with activation functions built in Simulink. Figure 8 The diagram shows the loss curves for training two datasets. The experimental results show that training on both datasets was completed within 1 ns. The data was switched 4-9 times. Under the same data and network, GPU training required an average of 12 steps and took an average of 13.6 ms. The circuit involved in this embodiment is 13,600,000 times faster than the digital circuit GPU. To ensure experimental rigor, the GPU time only includes forward and backward propagation times, excluding GPU initialization, data initialization, and optimizer initialization times. Furthermore, the Simulink simulation uses ideal devices and does not consider latency, error, or noise. Moreover, GPUs are versatile and can be applied to networks with different structures, so GPUs may sacrifice some speed for versatility. Dedicated digital circuits may be faster. However, even dedicated digital circuits will not be this many orders of magnitude faster. Therefore, this experimental result proves that analog circuits are much faster than GPUs, as shown in Table 1 below.

[0089] Table 1

[0090]

[0091] To further ensure rigor, the experiment compared different CUDA versions, different GPUs, and different operating systems, and obtained similar experimental results.

[0092] like Figure 9 The figure shows the changes in the solution during training. It can be seen from the graph that the solution slides to the lowest point along the steepest descent direction. This is consistent with the proof of unbiased theory. Figure 10 As shown, the circuit completed the training of the XOR task within 10 ns. This also proves that the circuit has a fast training speed. The XOR task is a classic task in neural networks, which requires multiple layers of nonlinear networks to learn. The circuit involved in this embodiment also learned the XOR task within 10 ns. This shows that the circuit has the ability to be extended to other neural network structures, such as convolutional neural networks, ResNet, Transformer, GPT, etc.

[0093] Low expansion cost primarily considers the complexity of the optimization unit and the impact of networking on efficiency. Regarding the complexity of the optimization unit, the circuits involved in the embodiments of this disclosure are all... Figure 1 The circuit structure shown is an infinite repetition, and its structure is simple. In contrast, the complexity of a GPU is mainly reflected in the following aspects: 1. Complex programming and control logic. GPUs contain complex control units for managing task scheduling, memory management, error correction, and other system-level functions. 2. Complex data paths: GPUs are designed with complex data paths capable of handling large amounts of data movement, including data caching, prefetching, and efficient data transfer mechanisms. 3. Complex thermal and power management: High-performance GPUs require complex thermal and power management systems to cope with heat and power consumption under high loads.

[0094] Regarding networking efficiency, the circuits in the embodiments of this disclosure can use... Figure 5The structure shown is extended. The output neurons of the front subnet are connected to the input neurons of the back subnet during forward propagation, and the final error differential wire is fed back to all subnets. In contrast, GPU networking is much more complex. The main issues involved are: 1. Data transfer latency. Cross-GPU communication: In model-parallel settings, different parts of the model are assigned to different GPUs. This means that data generated by each GPU during computation (such as activation values ​​and gradients) needs to be passed to other GPUs. Data transfer can cause significant latency when GPUs are connected via PCIe bus or network (in the case of cross-nodes). Bandwidth limitations: Data transfer speed is limited by hardware bandwidth. Even high-speed connections such as InfiniBand or NVLink can become bottlenecks when the data volume is large. 2. Synchronization overhead. Barrier synchronization: Parallel processing requires ensuring that all processing units have completed their current step before proceeding to the next computation. This synchronization typically requires a barrier synchronization mechanism, which ensures that all GPUs reach the same state before starting the next computation cycle. Each synchronization causes all GPUs to wait for the slowest one, introducing latency. Consistency and State Management: Maintaining data consistency is essential during parallel processing, which may involve additional communication and processing overhead to ensure that the data state of all processing units is consistent. 3. Communication to Computation Ratio. Computation / Communication Ratio: This is an important metric for measuring parallel efficiency, describing the ratio of the time required to perform computation to the time required for communication. Ideally, we want this ratio to be as high as possible, meaning that most of the time is spent on actual computation rather than data transfer. However, in many model parallel scenarios, especially when the model is large or distributed across multiple computing nodes, this ratio may be too low.

[0095] High parallelism refers to the proportion of parameters being optimized simultaneously out of all parameters. A higher proportion indicates a higher degree of parallelism in the computing hardware. Higher parallelism leads to higher computational efficiency. According to Amdahl's Law, if the proportion of tasks that can be executed in parallel is 1, and the number of computing resources is the weight number N, then the Amdahl speedup is N. The main advantage of high parallelism is reflected in time complexity, as shown in Table 2.

[0096] Table 2 Time Complexity

[0097] Digital circuits ACO Forward propagation <![CDATA[O(LN 2 )]]> O(1) Backpropagation <![CDATA[O(N 2 )]]> O(1)

[0098] L is the number of layers in the neural network.

[0099] N is the number of neurons in the hidden layer.

[0100] In this embodiment, all weights of the analog circuit are optimized simultaneously, therefore the time complexity is not affected by the network size. The optimization speed of the digital circuit slows down as the network size increases.

[0101] Processing-in-Memory (PIM) computing is a technology that performs computational tasks directly within storage devices. Since the computation and storage of each weight are the same physical quantity, determined by the output level of the integrator, the circuit architecture of this invention belongs to the PIM architecture. This architecture breaks through the boundary between the processor and storage in traditional architectures, allowing data processing during the storage stage and reducing the need for data transfer between the processor and memory. This invention belongs to the PIM architecture and has its advantages. PIM has several significant advantages, especially suitable for data-intensive applications and large-scale data processing scenarios. 1. Reduced data transfer requirements: In traditional computing architectures, the CPU needs to frequently read and write data from memory, which leads to significant energy consumption and latency. PIM design significantly reduces data movement by processing data directly in memory, thereby reducing energy consumption and increasing data processing speed. 2. Reduced energy consumption: Data transfer is one of the main sources of energy consumption in modern computing systems. PIM technology effectively reduces the overall system energy consumption by reducing data transfer on the system bus and other interfaces. 3. Improved Performance: Since data processing speed is limited by data transfer speed, CPUs in traditional architectures often remain idle while waiting for data. In-memory computing reduces processor wait time, allowing data processing tasks to be completed faster, thus improving overall computing performance. 4. Scalability and Flexibility: In-memory computing technology offers greater flexibility, allowing for customized configuration of memory and computing resources based on application needs. This enables optimization for specific applications, such as big data analytics and machine learning. 5. Simplified System Design: By moving computational tasks to memory, the overall hardware design of the system can be simplified, reducing reliance on high-speed data buses and other high-cost hardware, and potentially reducing the physical size of the system. 6. Improved Parallel Processing Capabilities: In-memory computing architecture allows computational tasks to be executed in parallel across multiple memory modules, significantly enhancing parallel processing capabilities. This is particularly important for applications requiring high parallelism. 7. Enhanced Data Processing Security: By processing data directly in memory, data transfer within the system is reduced, thereby lowering the risk of potential data leaks and enhancing data security.

[0102] On-chip learning is a concept involving implementing learning and inference tasks directly on a microchip (typically a dedicated integrated circuit or processor). This invention, due to its rapid training capabilities, enables real-time training on edge devices, achieving on-chip learning. This approach is particularly important in the field of artificial intelligence, focusing on integrating the training and execution of machine learning models onto the hardware device itself, rather than relying on external large-scale computing systems or cloud infrastructure. On-chip learning is a crucial component of edge computing and smart hardware design, significantly improving the speed and efficiency of data processing while reducing data transmission requirements and latency.

[0103] In summary, the neural network optimization circuit in this embodiment has the following advantages:

[0104] 1. Fast inference and training: The circuit uses analog circuits for inference and training of neural networks while preserving the continuity of analog system signals, enabling the circuit to train neural networks at high speed.

[0105] 2. Low expansion cost: Neural networks can be expanded through direct connections, eliminating the need for signal conversion (such as NVLink), encoding / decoding, and other operations between circuits. Furthermore, analog circuit neural networks integrate in-memory computation and have a simple weight structure, resulting in low cost (compared to the in-memory computation separation structure of GPUs).

[0106] 3. Highly parallel: During neural network training, all weights are modified simultaneously, resulting in low time complexity. The time complexity is only 1.

[0107] 4. In-memory computing: Both circuit storage and computation are implemented through weighted levels and are the same DC level signal, thus belonging to the in-memory computing architecture.

[0108] 5. On-chip learning: Currently, edge computing devices capable of on-chip learning are scarce. Current on-chip learning devices generally suffer from low training speeds and can only train a subset of weights. The embodiments disclosed in this disclosure can achieve on-chip learning of all weights.

[0109] 6. Fast convergence speed, no clock-driven circuit required.

[0110] 7. Moving the second differentiator from the disturbance source to the adder output position can maintain the original convergence direction of the circuit, reduce the requirement for disturbance power, and improve the stability of the network.

[0111] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A neural network optimization circuit, characterized in that, Applied to target neural networks, including: The first conversion unit is used to convert the target training data into the target DC level signal corresponding to the target training data; The second conversion unit is used to convert the label data corresponding to the target training data into a labeling result level signal; The first multiplier is used to output the prediction result signal corresponding to the target training data based on the target DC level signal and the adder output signal corresponding to the target training data. An error calculation unit is used to calculate the MSE error signal between the prediction result signal and the labeled result signal of the target training data; The first differentiator is used to receive the MSE error signal and output the first differentiator output signal; The second differentiator is used to receive the output signal of the adder and to output the output signal of the second differentiator. The second multiplier is used to receive the output signal of the first differentiator and the output signal of the second differentiator, and to output the output signal of the second multiplier. An integrator is used to receive the output signal of the second multiplier and output the integrator output signal; An adder is used to receive multiple weight change signals of multiple weight change data, the integrator output signal, and output the adder output signal; Each weight change signal corresponds to a different weight value, and each weight change data represents the amount of change in the weight value corresponding to the weight change data over time.

2. The neural network optimization circuit according to claim 1, characterized in that, The adder is a unidirectional adder. The co-current adder is used to receive multiple weight change signals of the multiple weight change data and the DC level signal of the weight value corresponding to each weight change data, and to transmit the updated weight value of each weight change data to the first multiplier and the second differentiator. The multiple weight change signals of the multiple weight change data are disturbance signals output by introducing noise sources.

3. The neural network optimization circuit according to claim 2, characterized in that, The second differentiator is a co-directional differentiator; the co-directional differentiator is used to transmit the change in weight value over time corresponding to each weight change data output to the second multiplier to form a closed loop.

4. The neural network optimization circuit according to claim 3, characterized in that, The first multiplier is used to output the product of the input, a first numerical DC level signal, and the updated weight value of each weight change data.

5. The neural network optimization circuit according to claim 4, characterized in that, The error calculation unit includes: a subtractor, wherein the first differentiator is an inverse differentiator; The subtractor is used to subtract the product of the multiple weight change signals from the second numerical DC level signal of the true label corresponding to the labeling result signal of the target training data, and then perform square calculation through the third multiplier. After that, the subtractor performs the corresponding differentiation and takes the opposite number through the inverse differentiator. The DC level signal after the inverse differentiator represents the opposite number of the MSE differential. The second multiplier is used to receive the inverse of the MSE differential and the change in weight value over time corresponding to each weight change data, and to provide the product result output by the second multiplier as the weight modification amount to the integrator; the integrator and the amplifier together form an integrator with proportional amplification.

6. The neural network optimization circuit according to claim 1, characterized in that, The target neural network includes multiple different subnetworks, each subnetwork establishing a control chip. Each pin on the input side of the control chip corresponds to an input neuron, and each pin on the output side of the control chip corresponds to an output neuron. Both the input neuron and the output neuron are DC level signals; In two adjacent subnets, the output neurons of the front subnet are connected to the input neurons of the rear subnet; When training the target training data in the target neural network, the MSE error signal is fed back to all subnetworks of the target neural network through wires.

7. The neural network optimization circuit according to claim 1, characterized in that, The input side of the target neural network is connected to the first conversion unit, and the output side of the target neural network is connected to the error calculation unit.

8. The neural network optimization circuit according to claim 1, characterized in that, Also includes: The computer is used to send digital signals corresponding to the target training data to the first conversion unit, send digital signals corresponding to the label data of the target training data to the second conversion unit, and receive prediction result signals corresponding to the target training data.

9. The neural network optimization circuit according to claim 1, characterized in that, When training the target training data, the target neural network simultaneously modifies the different weight values ​​corresponding to each weight change data.

10. The neural network optimization circuit according to any one of claims 1 to 9, characterized in that, The target neural network includes: Transformer neural network, CNN convolutional neural network, RNN recurrent neural network, or fully connected neural network.

Citation Information

Patent Citations

  • A design method of a linear equation solver based on a variable parameter convergent neural network

    CN109033021A

  • Group Lasso-based neural network cutting method for power amplifier

    CN110414565A