Probability calculation accelerator
By converting the multiplication operation in the derivative process of the sigmoid function into random sampling operation of magnetic tunnel junctions, the problem of excessive multiplication operation in deep artificial neural network training is solved, and energy consumption reduction and training efficiency improvement is achieved.
Patent Information
- Application Number
- CN202510266964.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
When the prior art uses backward propagation algorithms to train deep artificial neural networks, there are a large number of multiplication operations, resulting in high energy consumption and increased training costs.
By converting the multiplication operation in the derivative process of the sigmoid function into a random sampling operation of the magnetic tunnel junction, the relationship between the flip probability of the magnetic tunnel junction and the driving voltage is used to calculate the derivative value of the sigmoid function, thereby reducing the number of multiplication operations.
It effectively reduces the proportion of multiplication operations during training, reduces energy consumption, and improves training efficiency, which is especially suitable for training large-scale models.
Smart Images

Figure CN120197655A_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of artificial intelligence, and more particularly, to a probability calculation accelerator that can be used for, for example, backpropagation algorithms. Background Art
[0002] Deep artificial neural networks are effective ways to achieve artificial intelligence. For example, convolutional neural networks, recurrent neural networks, long short-term memory networks, etc. have been widely used in many scenarios such as pattern recognition, computer vision, and large language models. In order to improve the training and operation efficiency of artificial neural network models, the academic community is on the one hand developing more efficient algorithms, and on the other hand, continuously researching and developing suitable hardware systems to give full play to the physical characteristics advantages of microelectronic devices themselves to increase the efficiency and speed of artificial neural network algorithms. Summary of the Invention
[0003] One aspect of the present invention provides a probability calculation accelerator, including: an operation unit for performing operations of a network model using sample data to obtain probability values of node outputs of the network model, where the nodes of the network model use a sigmoid function as an activation function; a sampling unit for sampling a magnetic tunnel junction unit according to the probability values of node outputs to obtain sampling values; and a parameter update device, including: a derivative function determination unit for determining the value of the derivative function of the sigmoid function of a node according to the sampling values obtained by the sampling unit; a gradient calculation unit for calculating the gradient value of the parameters of the network model according to the value of the derivative function of the sigmoid function of the node; and a parameter update unit for updating the parameters of the network model based on the gradient value of the parameters of the network model.
[0004] According to an embodiment, the probability calculation accelerator further includes a convergence determination device, and the convergence determination device includes: an error calculation unit for calculating the error between the output value of the network model corresponding to the sample data and the true value; and a convergence determination unit for determining whether the error has satisfied the convergence condition. Wherein, if the error has satisfied the convergence condition, the parameter update device is stopped from updating the parameters of the network model; if the error has not satisfied the convergence condition, the parameter update device is continued to be used to update the parameters of the network model.
[0005] According to an embodiment, the sampling unit samples the magnetic tunnel junction unit twice continuously according to the probability values of node outputs in the network model to obtain two sampling values. When the two sampling values are (1, 0), the derivative function determination unit determines that the value of the derivative function of the sigmoid function of the node is 1, and when the two sampling values are (0, 0), (0, 1), or (1, 1), the derivative function determination unit determines that the value of the derivative function of the sigmoid function of the node is 0.
[0006] According to an embodiment, the sampling unit performs 2n consecutive samplings on the magnetic tunnel junction unit according to the probability value output by the node in the network model to obtain n pairs of sampling values, where n is an integer greater than or equal to 1. The derivative determination unit determines the ratio of the sampling values of (1, 0) in the n pairs of sampling values and determines the ratio as the value of the derivative of the sigmoid function of the node.
[0007] According to an embodiment, the gradient calculation unit calculates the gradient value of the parameter of the network model as the partial derivative of the error function of the network model with respect to the parameter, and the parameter update unit updates the parameter of the network model using the product of the gradient value of the parameter of the network model and the learning rate.
[0008] Another aspect of the present invention provides a method for training a network model, where the node of the network model uses the sigmoid function as an activation function. The training method includes: performing operations on the network model using sample data to obtain the probability value output by the node of the network model; sampling the magnetic tunnel junction unit according to the probability value output by the node to obtain sampling values; determining the value of the derivative of the sigmoid function of the node according to the sampling values; calculating the gradient value of the parameter of the network model according to the value of the derivative of the sigmoid function of the node; and updating the parameter of the network model based on the gradient value of the parameter of the network model.
[0009] According to an embodiment, the training method further includes: calculating the error between the output value of the network model corresponding to the sample data and the true value; and determining whether the error has satisfied the convergence condition. If the error has satisfied the convergence condition, stop updating the parameter of the network model and complete the training of the network model; if the error has not satisfied the convergence condition, continue to execute the steps from sampling the magnetic tunnel junction unit to updating the parameter of the network model.
[0010] According to an embodiment, sampling the magnetic tunnel junction unit according to the probability value output by the node includes: performing two consecutive samplings on the magnetic tunnel junction unit according to the probability value output by the node in the network model to obtain two sampling values. When the two sampling values are (1, 0), determine the value of the derivative of the sigmoid function of the node as 1; when the two sampling values are (0, 0), (0, 1), or (1, 1), determine the value of the derivative of the sigmoid function of the node as 0.
[0011] According to one embodiment, sampling the magnetic tunnel junction unit according to the probability value output by a node includes: continuously sampling the magnetic tunnel junction unit 2n times according to the probability value output by the node in the network model to obtain n pairs of sampling values, where n is an integer greater than or equal to 1, determining the ratio of the sampling values of (1, 0) in the n pairs of sampling values, and determining the ratio as the value of the derivative function of the sigmoid function of the node.
[0012] According to one embodiment, the gradient value of the parameter of the network model is the partial derivative of the error function of the network model with respect to the parameter, and the product of the partial derivative and the learning rate is used to update the parameter.
[0013] The above and other features and advantages of the present invention will become apparent from the following description of exemplary embodiments with reference to the accompanying drawings. Description of the Drawings
[0014] Figure 1 is a graph showing the relationship between the switching probability of the magnetic tunnel junction and the driving voltage.
[0015] Figure 2A 、 2B and 2C are the theoretical curve, simulation curve, and experimental results of the probability distributions where the results of continuously sampling the magnetic tunnel junction twice are 1 and 0, respectively.
[0016] Figure 3 is a schematic diagram of an XOR gate network.
[0017] Figure 4 is a simulation curve of the change in the loss function for training the XOR gate network by using random sampling of the magnetic tunnel junction instead of taking the derivative of the sigmoid function.
[0018] Figure 5 is a block diagram of a probability calculation accelerator according to an embodiment of the present invention.
[0019] Figure 6 is a schematic diagram of a fully connected network.
[0020] Figure 7 is a flowchart of a network model training method according to an embodiment of the present invention. Detailed Embodiments
[0021] Exemplary embodiments of the present invention will be described below with reference to the accompanying drawings. Note that the drawings may not be drawn to scale.
[0022] A fundamental training algorithm widely used in many artificial neural networks is the BackPropagation algorithm, or simply the BP algorithm for short. Its working principle is as follows: Given the current network state or network parameter settings, a data point X from the training set is input at the network input end, and through the forward propagation of the network F, the output value Y = F w (X) is obtained, where W is a vector describing the weights of each edge or the bias parameters of each node in the artificial neural network, and F is the artificial neural network function characterized by the network parameter W, also known as the algorithm or model, which contains a large number of non-linear activation functions f. This process is the forward propagation process of the network. Subsequently, by comparing the output value Y of the network at this time with the ideal value Y ideal corresponding to X, the error function E of the network can be obtained as E = |Y - Y ideal | 2 . Obviously, given (X, Y ideal ), the error value is also a function of the network parameter W. Given the network architecture and function f, the partial derivatives of the error function E with respect to each component of the weight W can be calculated using the chain rule and based on this, the parameter vector W is updated to obtain a new state of the artificial neural network. This cycle continues until the error function E converges to the minimum value, thus completing the training process.
[0023] Decomposing the training process of the above BP algorithm, it can be seen that this process involves a huge number of multiplication, addition, and derivative operations. Among them, thanks to the multi-layer network architecture design of the artificial neural network, some trivial multiplication and addition operations can be effectively reduced to matrix multiplication operations. The reduced matrix multiplication operations can be simulated by leveraging the transport properties constrained by Kirchhoff's law in the crossbar structure of non-volatile random access memories such as magnetic random access memory (MRAM), resistive random access memory (RRAM), phase change random access memory (PCRAM), etc., so that corresponding hardware can be developed to accelerate the matrix multiplication operation, and thus accelerate the training or operation process of the artificial neural network. Nevertheless, the operation process of the artificial neural network still contains a relatively high proportion of multiplication operations, including the multiplication operations required for differentiating the function f and the multiplication operations between matrix elements, etc. And these complex multiplication operations, even with the assistance of corresponding matrix operation hardware accelerators, are still a major source of energy consumption and need to be further compressed and saved from the hardware aspect.
[0024] A large part of the multiplication operations in the BP algorithm come from the multiplication operations for the derivative of the node activation function f (also known as the triggering function) and its derivative f'. Commonly used activation functions include the sigmoid function, Tanh function, ReLU function, ELU function, softplus function, softmax function, and swish function, etc. Among them, the sigmoid function has advantages such as a limited value range, continuous differentiability, and monotonic increase, and has been widely used in the field of artificial neural networks. The following Formula 1 and Formula 2 respectively show the sigmoid function and its first-order derivative function. It can be seen that the first-order derivative function of the sigmoid function can be represented by itself:
[0025]
[0026] If the f function is selected as the sigmoid function when constructing an artificial neural network, then as shown in Formula 2 above, f' = f × (1 - f). Then, for each additional network node, that is, one triggering of the sigmoid function, the single training process related to this node will increase by 2 multiplication operations, namely the operations of ×f and ×(1 - f). Currently, large-scale models may have billions to hundreds of billions of parameters, and a huge number of multiplication operations are required during the training process, consuming a large amount of energy, which occupies a large part of the model training cost.
[0027] Magnetic tunnel junction (MTJ) is a magnetic multi-layer film structure developed with the development of spintronics, mainly including a non-magnetic insulating barrier layer sandwiched between two ferromagnetic layers, and electrons can tunnel through this barrier layer. When the magnetic moment directions of the two ferromagnetic layers are parallel to each other, the magnetic tunnel junction has the minimum resistance, called the low-resistance state, for example, corresponding to the data "0"; when the magnetic moment directions of the two ferromagnetic layers are anti-parallel to each other, the magnetic tunnel junction has the maximum resistance, called the high-resistance state, for example, corresponding to the data "1". It has been proposed to use magnetic tunnel junctions as probability-controllable true random number generators. For example, see the applicant's prior invention patent applications CN202310508761.7, CN202111072657.5, and CN202410527496.1, etc. The present inventor noticed that when using the magnetic tunnel junction as a random number generator, the flip probability of generating a high-resistance state or its probability P of having a value of 1 is basically satisfied with a dependence relationship similar to the sigmoid function, P = S(V), as Figure 1 shown. In Figure 1 , the solid line represents the sigmoid function curve, and the circles represent the points defined by the driving voltage V of the magnetic tunnel junction and the corresponding flip probability P. From Figure 1 it can be seen that the mutual dependence relationship between the flip probability P and the driving voltage V basically satisfies the sigmoid function curve.
[0028] Based on the above characteristics of the magnetic tunnel junction, the present application proposes to convert the multiplication operations ×f and ×(1 - f) related to the derivative calculation process of the sigmoid function in the BP algorithm into sampling operations of random numbers, thereby reducing the amount of multiplication operations that need to be performed during the training process and effectively reducing the proportion of multiplication operations. As referred to above Figure 1 As described, the function value of the derivative function S(1 - S) of the sigmoid function is equivalent to the joint probability of sampling 1 and 0 in two consecutive sampling processes using the magnetic tunnel junction as a random number generator under the same trigger condition V. In the present application, consecutive sampling refers to multiple samplings performed on the magnetic tunnel junction successively, without inserting other samplings during these multiple samplings. For example, two consecutive samplings refer to the results of two directly adjacent samplings, without inserting other sampling results between them. In other words, consecutive sampling does not impose any restrictions on the sampling time, and the time interval between two samplings can vary, but the sampling conditions are basically the same. Figures 2A - 2C Shows the probability P of the results of two consecutive samplings being 1 and 0 10 distribution, where Figure 2A is the theoretical distribution curve, Figure 2B is the simulated sampling result, Figure 2C is the experimental demonstration result using the magnetic tunnel junction, where the circles represent the experimental values and the solid line is the theoretical distribution curve. From Figure 2A 、 2B and 2C, it can be seen that both the simulation results and the experimental results are consistent with the theoretical distribution curve. Or described from another perspective, using the magnetic tunnel junction as a Bernoulli random number generator, under the same trigger voltage V, only when the results of two consecutive samplings are 1 and 0, the derivative function S(1 - S) is regarded as 1; otherwise, in all other cases, the derivative function S(1 - S) is regarded as 0. In this case, it is equivalent to converting 2 multiplication operations into 2 random number sampling operations, and the energy consumption of the latter can be extremely low because the energy consumption of the thermally assisted random number generation process in the magnetic tunnel junction can be as low as the order of 10 fJ.
[0029] Taking the XOR gate as an example below, the process and effect of using magnetic tunnel junction sampling to replace the derivative calculation operation of the sigmoid function are described. Figure 3 Shows the network structure corresponding to the XOR gate, which includes three layers. The first layer contains two input nodes A, B and a bias node I; the second layer contains two hidden layer nodes D, E and a bias node I; the third layer contains an output node C. Nodes A and B correspond to the two input values of the XOR operation, and node C represents the output value obtained by propagating the input values of A and B forward along the network. Under the current network weight state, the output value O of node C under the current input can be predicted using the forward propagation process from the input values of nodes A and B c, and the process is shown in Equation 3 below, where f is the activation function of the corresponding node, and the sigmoid function is adopted here.
[0030] I D = w AD A + w BD B + b D
[0031] I E = w AE A + w BE B + b E
[0032] O D = f D (I D )
[0033] O E = f E (I E )
[0034] I C = w DC O D + W EC O E + b C
[0035] O C = f C (I C ) = P C Equation 3
[0036] Through the forward propagation process, it can be judged whether the current network weights can implement the function of the XOR gate. In order to make the network weights achieve the function of the XOR gate, the error backpropagation and gradient descent algorithms can be used to train the network. During the training process, the energy (i.e., the error function) is set to where O C is the actual output value of node C, is the true value, and the training set is ABC = (000, 011, 101, 110), with a total of four samples. The first two digits of each sample data represent the inputs of nodes A and B, and the last digit represents the ideal output value of node C, that is, the true value. The conventional error backpropagation and gradient descent process is as follows: for each sample input in the training set, the update step size of each weight value can be obtained according to formula 4. The four samples in the training set are input in sequence, and the four obtained update step sizes are averaged and used for the update of the network weights. Repeating this process continuously can obtain the training result. As shown in formula 4, this conventional process requires a large number of derivative calculations of activation functions and multiplication operations, consuming a large amount of resources. Therefore, the probability calculation accelerator proposed in this application can be used to perform the derivative calculation of the sigmoid activation function and related multiplication operations based on the probability sampling operation of magnetic tunnel junctions, which will be further described in detail later.
[0037]
[0038] For example, when training with a single sample in the training set, through the forward propagation process of formula 3, the probabilities O D 、O E 、O C that nodes D, E, and C are 1 can be obtained, which can be expressed as O i or P i , where i = D, E, C and O i = P i . Using these probabilities, 6 probability samplings are performed on the magnetic tunnel junction - two sampling operations are continuously performed under each probability O i (i = D, E, C), and a total of 3 × 2 = 6 random numbers are obtained. Here, sampling the magnetic tunnel junction with probability O i means that according to the relationship curve shown in Figure 1 , the magnetic tunnel junction is sampled at the driving voltage V corresponding to the probability O i , that is, the output of the magnetic tunnel junction is measured as "0" (for example, low resistance state) or "1" (for example, high resistance state). The output values of these 6 random numbers can help simplify the operations in the weight update process: for example, for the weight b C , when and only when the results of two consecutive samplings according to the probability O C are (1, 0), the weight b C is fine-tuned by 2(O C - O C ideal ) × α, and in other cases, the update step size of b C is 0, where α is the set learning rate; for the weight w AD , when and only when the results of two consecutive O C probability samplings are (1, 0), two consecutive OD When the probability sampling result is also (1, 0), the weight w AD Fine-tune 2 (O C -O C ideal ) × w DC × A × α, in other cases w AD The update step size is 0; for the weight w BE , when and only when the O C probability sampling result is (1, 0) twice in a row, and the O E probability sampling result is also (1, 0) twice in a row, the weight w BE Fine-tune 2 (O C -O C ideal ) × w EC × B × α, in other cases w BE The update step size is 0. Where A, B, O C ideal correspond to the values in the training samples. Other weights are updated according to the same rule. As described above, by sampling 6 times, the update step sizes of all weights can be obtained. Using 4 samples in the training set to train in sequence, the update step sizes of all weights will be obtained 4 times. Take the average of the 4 obtained update step sizes and use the average of the update step sizes to update the corresponding network weight parameters, as shown in formula 5 below.
[0039]
[0040] After updating the weight parameters of the network, use the sample data in the training set to run the updated network and calculate and record the loss function E = [O C (AB = 00) - 0] 2 + [O C (AB = 01) - 1] 2 + [O C (AB = 10) - 1] 2 + [O C (AB = 11) - 0] 2 . It can be evaluated whether the loss function E converges to the minimum value. If not, continuously repeat the above training process until the loss function E converges to the minimum value, and the training process is completed.
[0041] Figure 4 Shows the simulation results of training the XOR gate network by using random sampling of magnetic tunnel junctions instead of differentiating the sigmoid function. Approximately 10 5After the training, the loss function is basically stable near the minimum value, and this result shows the convergence of the training process. Table 1 below shows the final training results of the XOR gate network, and it can be seen that the obtained network model has a high accuracy rate. This method can not only reduce the proportion of multiplication operations, but also introduce a certain degree of randomness during the training process, which helps the system escape from local minima during the training process and helps to achieve global optimization. In addition, this training process is particularly beneficial for large models with a huge number of parameters. Conventional training methods need to perform a large number of operations for the parameter update process of each node because the derivative value of the sigmoid function corresponding to the node depends on a certain value of the node output value; while in the training process of the present invention, according to the random sampling process of the magnetic tunnel junction, the derivative value of the sigmoid function is determined to be 1 or 0, and only the parameters with the derivative value of the sigmoid function being 1 need to be updated, and the parameters with the derivative value of the sigmoid function being 0 do not need to be updated, which is equivalent to selecting some important parameters to calculate the update step and perform the update, while ignoring the unimportant parameters, so the amount of calculation during the training process is also greatly reduced, and the training process is accelerated.
[0042] Table 1: Training Results of XOR Network
[0043] Input A Input B Probability that Output C is 1 0 0 0.007 0 1 0.994 1 0 0.991 1 1 0.005
[0044] Figure 5 is a block diagram of a probability calculation accelerator according to an embodiment of the present invention. The probability calculation accelerator can be used for the training process of a network model using the sigmoid function as an activation function. As Figure 5 shown, the probability calculation accelerator may include an operation unit 110, a convergence judgment device 120, a sampling unit 130, and a parameter update device 140. The principle of the probability calculation accelerator shown below will be described by taking the training process of the XOR network shown Figure 3 as an example. Figure 5 shown.
[0045] The operation unit 110 can receive sample data as a training set. For example, the training set is ABC=(000, 011, 101, 110), and there are four samples in total. The first two bits of each sample data represent the inputs of nodes A and B, and the last bit represents the ideal output value of node C, that is, the true value. The operation unit 110 can perform the operation of the XOR network model using the received sample data and the current weight value of the XOR network model. As shown in the above formula 3, the actual output values of each node, which can also be called probability values, can be calculated. In this example, the probability values O D 、O E 、O C of nodes D, E, and C being 1 can be obtained.
[0046] The convergence judgment device 120 can calculate an error function (i.e., a loss function) based on the model network output value provided by the operation unit 110, and determine whether the error converges to a predetermined range. Specifically, the convergence judgment device 120 can include an error calculation unit 122 and a convergence judgment unit 124. The error calculation unit 122 can calculate the error between the current output of the network model and the ideal value (i.e., the true value). For example, for Figure 3 the XOR network shown, the error function E = [O C (AB = 00) - 0] 2 + [O C (AB = 01) - 1] 2 + [O C (AB = 10) - 1] 2 + [O C (AB = 11) - 0] 2 can be used to calculate the error E between the current output of the XOR network model and the ideal value (i.e., the true value). The convergence judgment unit 124 can determine whether the error value E calculated by the error calculation unit 122 is small enough, for example, lower than a predetermined value. For example, when the error values E obtained from multiple calculations are all lower than the predetermined value, the convergence judgment unit 124 can determine that the error has converged to the predetermined range, and thus the training process of the model can be completed. Otherwise, the convergence judgment unit 124 can determine that the error has not converged to the predetermined range, and more training set data still needs to be used to train the model.
[0047] If the convergence judgment device 120 determines that the model error has not converged to the predetermined range and the model still needs to be trained, the sampling unit 130 and the parameter update device 140 can be used to continue the training process of the network model. Specifically, the sampling unit 130 can perform probability sampling on the magnetic tunnel junction (MTJ) 132 according to the output probability values of each node in the network model calculated by the operation unit 110. Specifically, 2n sampling operations can be continuously performed on the magnetic tunnel junction 132 according to each probability value O i (i = D, E, C), where n is an integer greater than or equal to 1. Here, sampling the magnetic tunnel junction according to the probability value O i means that according to the relationship between the switching probability P of the magnetic tunnel junction, or rather the probability P of having a value of 1, and the drive voltage V (i.e., the voltage that drives the magnetic tunnel junction to perform a switching operation, also called the switching voltage), the magnetic tunnel junction is sampled at the drive voltage V corresponding to the probability O i , that is, the output of the magnetic tunnel junction is measured as "0" (e.g., low resistance state) or "1" (e.g., high resistance state). In some embodiments, each probability value O i(i = D, E, C) performs two consecutive sampling operations on the magnetic tunnel junction 132, that is, the case where the above n value is equal to 1. In this example, for the three probability values O D 、O E 、O C six random numbers can be obtained, or rather, three pairs of random numbers. In some other embodiments, according to each probability value O i (i = D, E, C), the magnetic tunnel junction 132 can be sampled 2n times continuously, where n is an integer greater than 1. For example, when n is 100, for each probability value O i (i = D, E, C), 100 pairs (i.e., 200) of sampling data are obtained.
[0048] The parameter update device 140 can update the parameters of each node of the network model based on the sampling results of the sampling unit 130. In one embodiment, the parameter update device 140 may include a derivative function determination unit 142, a gradient calculation unit 144, and a parameter update unit 146. For example, when the sampling unit 130 performs two (i.e., n = 1) consecutive sampling operations on the magnetic tunnel junction 132 according to each probability value O i (i = D, E, C), if the two sampling values corresponding to O i are 1 and 0, the derivative function determination unit 142 can determine that the derivative function P i (1 - P i ) of the sigmoid activation function in formula 4 is 1, where O i = P i ; if the two sampling values corresponding to O i are 0 and 0, 0 and 1, or 1 and 1, the derivative function determination unit 142 can determine that the derivative function P i (1 - P i ) of the sigmoid activation function in formula 4 is 0. In some other embodiments, if the sampling unit 130 performs 2n (where n is greater than or equal to 1) consecutive sampling operations on the magnetic tunnel junction 132 according to each probability value O i (i = D, E, C), the derivative function determination unit 142 can calculate / statistic the ratio of the cases where the sampling values of every two consecutive samplings are 1 and 0 among all the sampling values, and use it as the value of the derivative function P i (1 - P i ) in formula 4. For example, if according to the probability value O iThe magnetic tunnel junction 132 was sampled 200 times to obtain 100 pairs of sampling values. Each pair of sampling values includes the sampling results of two consecutive samplings. Then, the ratio of the sampling results of (1, 0) in the 100 pairs of sampling values can be calculated / statistically analyzed. For example, if there are 21 pairs of sampling values of (1, 0), then the derivative function P of the sigmoid activation function in Formula 4 can be determined i (1 - P i ) is 0.21.
[0049] The gradient calculation unit 144 can determine the derivative function P of the sigmoid activation function determined by the derivative function determination unit 142 i (1 - P i ) to calculate the gradients of each network parameter. For example, for the Figure 3 XOR network shown, the gradient calculation unit 144 can substitute the value of the derivative function P i (1 - P i ) into Formula 4 to calculate the gradients of each network parameter. For the case of continuously sampling the magnetic tunnel junction 132 twice according to the probability value O i , the value of the derivative function P i (1 - P i ) is 1 or 0. The gradient calculation unit 144 can calculate the gradient value only for the nodes where the value of the derivative function P i (1 - P i ) is 1, and ignore the nodes where the value of the derivative function P i (1 - P i ) is 0. For example, if the value of the derivative function P C (1 - P C ) corresponding to node C is 1, the gradient value of the weight b C is calculated as 2(O C - O C ideal ); if the value of the derivative function P C (1 - P C ) corresponding to node C is 1 and the value of the derivative function P D (1 - P D ) corresponding to node D is also 1, the gradient value of the weight w AD is calculated as 2(O C - O C ideal ) × w DC × A; if the value of the derivative function P C (1 - P C ) corresponding to node C is 1 and the value of the derivative function P E (1 - P E ) corresponding to node E is also 1, the gradient value of the weight w BE is calculated as 2(O C - OC ideal ) × w EC × B. Where A, B, and O C ideal correspond to the values in the training samples. The gradients of other weights are calculated according to the same rule. In some other embodiments, if according to the probability value O i the magnetic tunnel junction 132 is sampled continuously 2n times (n is greater than or equal to 1), and the derivative P of the sigmoid activation function determined by the derivative determination unit 142 i (1 - P i ) is the ratio of the number of times the sampling values are (1, 0) in two consecutive samplings to the total number of samplings, then the gradient calculation unit 144 can substitute the value of the derivative P i (1 - P i ) into Formula 4 to calculate the gradients of each network parameter. The calculation process is similar to that described above, with the only difference being that above, the derivative P i (1 - P i ) is determined to be 1 or 0, while here the derivative P i (1 - P i ) is determined to be a certain ratio, which may be a decimal greater than 0 and less than 1. For the specific calculation, refer to Formula 4, which will not be repeated here.
[0050] It should be noted that when the derivative P i (1 - P i ) is determined to be 1 or 0, the ratio of multiplication operations is greatly saved because only the sampling results are needed to determine the value of the derivative P i (1 - P i ), and when determining the gradient value, there is no need to calculate the multiplication step of multiplying by the derivative P i (1 - P i ) (when the derivative P i (1 - P i ) is 0, the gradient calculation of relevant parameters can be ignored; when the derivative P i (1 - P i ) is 1, only the other terms in the gradient value need to be calculated, omitting the step of multiplying by 1), which can further save the energy consumption required for training and thus save costs. On the other hand, when the derivative P i (1 - P i ) is determined to be the ratio of the number of times the sampling values are (1, 0) in two consecutive samplings to the total number of samplings, the step of determining this ratio requires performing a statistical / calculation process, and later a multiplication operation with this ratio needs to be performed for each parameter (unless this ratio is 0), so this hardly substantially reduces the amount of multiplication operations, but this scheme still provides a method to calculate the derivative P i (1 - P iA novel method, which is different from the prior art that calculates the derivative function P based on the above formula 2, i.e., S’(x) = S(x)(1 - S(x)). i (1 - P i ).
[0051] Finally, the parameter update unit 146 can use the gradient value calculated by the gradient calculation unit 144 to update the corresponding network model parameters, thus realizing one round of training process. For example, for Figure 3 the XOR network model shown, each network parameter can be updated according to the formula 5 given above. It should be noted that as shown in formula 5, the gradient value can be multiplied by a predetermined learning rate parameter α and then used to update the corresponding network parameter, where the value of the learning rate parameter α can be preset in advance. The probability calculation accelerator shown in Figure 5 can be used to continuously and iteratively repeat the above process until the error function E converges to a predetermined range, and the convergence judgment device 120 determines that the error function E satisfies the predetermined convergence condition, thus completing the training process.
[0052] To better understand the solution of the present invention, the training process performed using the probability calculation accelerator of the present invention will be described below in combination with another network model. Figure 6 is a schematic diagram of a fully connected network, which includes two layers, namely an input layer X and an output layer Y, where the input layer X includes nodes x1 to x m , a total of m nodes, and the output layer Y includes nodes y1 to y n , a total of n nodes. Figure 6 On the connection line between the input node x i and the output node y j , the weight value is w i,j . Figure 6 The fully connected network shown can be represented by the following formula 6.
[0053]
[0054] Among them, I yj is the input value of the node y j , O yj is the output value of the node y j , which can also be called the probability value P yj , and f is an activation function, which is a sigmoid function.
[0055] The following describes an example of using the Figure 5 shown probability calculation accelerator to perform the training process of this network. As described above, the operation unit 110 can receive sample data as the training set. For example, the training set is (X, Y ideal ), where X includes (x1, x2,..., x m) vector, Y ideal is a vector including (y1 ideal , y2 ideal , …, y n ideal ). The operation unit 110 receives a set of sample data (X, Y ideal ) each time, and calculates the output probability value O j of each node y yj according to Formula 6.
[0056] The error calculation unit 122 can calculate the error between the current output of the network model and the ideal value (i.e., the true value), for example, calculate the value of the error function , where is the true value of the sample data in the training set. The convergence judgment unit 124 can judge whether the error value E calculated by the error calculation unit 122 is small enough, for example, lower than a predetermined value. For example, when the error values E obtained by multiple calculations are all lower than the predetermined value, the convergence judgment unit 124 can judge that the error has converged to the predetermined range, and thus the training process of the model can be completed. Otherwise, the convergence judgment unit 124 can judge that the error has not converged to the predetermined range, and more training set data still needs to be used to train the model.
[0057] If the convergence judgment device 120 determines that the model error has not converged to the predetermined range and the model still needs to be trained, the sampling unit 130 and the parameter update device 140 can be used to continue the training process of the network model. Specifically, the sampling unit 130 can perform probability sampling on the magnetic tunnel junction (MTJ) 132 according to the output probability value O j of each node y yj in the network model calculated by the operation unit 110. For example, the magnetic tunnel junction 132 can be continuously sampled 2n times according to each probability value O yj , where n is an integer greater than or equal to 1. Here, sampling the magnetic tunnel junction according to the probability value O yj means that according to the relationship between the switching probability P of the magnetic tunnel junction or the probability P of the value being 1 and the driving voltage V (i.e., the voltage for driving the magnetic tunnel junction to perform the switching operation, also called the switching voltage), the magnetic tunnel junction is sampled at the driving voltage V corresponding to the probability P yj (where P yj = O yj ), that is, measure the output of the magnetic tunnel junction as "0" (for example, low resistance state) or "1" (for example, high resistance state). In some embodiments, the magnetic tunnel junction 132 can be continuously sampled 2 times according to each probability value O yj , that is, the above-mentioned case where the value of n is equal to 1. In other embodiments, according to each probability value Oyj The magnetic tunnel junction 132 is sampled continuously 2n times, where n is an integer greater than 1. For example, when n is 100, for each probability value O yj 100 pairs (i.e., 200) of sampled data are obtained.
[0058] When the sampling unit 130, according to each probability value O yj performs 2 (i.e., n = 1) consecutive sampling operations on the magnetic tunnel junction 132, if O yj the corresponding two sampled values are 1 and 0, the derivative determination unit 142 can determine the derivative P j of the sigmoid activation function of node y yj (1 - P yj ) to be 1; if O yj the corresponding two sampled values are 0 and 0, 0 and 1, or 1 and 1, the derivative determination unit 142 can determine the derivative P j of the sigmoid activation function of node y yj (1 - P yj ) to be 0. In some other embodiments, if the sampling unit 130 performs 2n (where n > 1) consecutive sampling operations on the magnetic tunnel junction 132 according to each probability value O yj the derivative determination unit 142 can calculate / stat the ratio occupied by the cases where the sampled values of every two consecutive samplings among all the sampled values are 1 and 0, and use it as the derivative P j of the sigmoid activation function of node y yj (1 - P yj ) value. For example, if 200 samplings are performed on the magnetic tunnel junction 132 according to the probability value O yj to obtain 100 pairs of sampled values, and each pair of sampled values includes the sampling results of two consecutive samplings, then the ratio of the sampling results being (1, 0) among the 100 pairs of sampled values can be calculated / stat. For example, if there are 23 pairs of sampled values being (1, 0), then the derivative P j of the sigmoid activation function of node y yj (1 - P yj ) can be determined to be 0.23.
[0059] The gradient calculation unit 144 can calculate the gradients of the respective network parameters w yj (1 - P yj ) value determined by the derivative determination unit 142. For i,j the fully connected network shown in Figure 6 , the gradient calculation unit 144 can calculate the gradient Δw i,j of the network parameter w i,j according to the following formula 7.
[0060]
[0061] Then, the parameter update unit 146 can use the gradient value calculated by the gradient calculation unit 144 to update the corresponding network model parameters, thus realizing one round of the training process. For example, for Figure 6 the shown fully connected network, the respective network parameters can be updated according to the following formula 8. The probability calculation accelerator shown in Figure 5 can continuously and iteratively repeat the above process until the error function E converges within a predetermined range, and the convergence determination device 120 determines that the error function E satisfies the predetermined convergence condition, thereby completing the training process.
[0062]
[0063] Figure 7 shows a method for training a network model according to an embodiment of the present invention. Since the training process of the network model of the present invention has been described in detail in the above reference Figures 1 - 6 , the training method will be briefly described here, and its details and examples can be referred to the above description.
[0064] The network model training method of the present invention is applicable to a network model using the sigmoid function as the activation function, and this training method uses the Back Propagation algorithm to train the network model. Referring to Figure 7 , this training method may include:
[0065] Step 710, performing operations on the network model using sample data to obtain the probability value of the node output of the network model;
[0066] Step 720, calculating the error between the output value of the network model corresponding to the sample data and the true value;
[0067] Step 730, determining whether the error has satisfied the convergence condition. If so, the training of the network model is completed and the method can end; if the convergence condition is not satisfied, the subsequent steps can be continued;
[0068] Step 740, sampling the magnetic tunnel junction unit according to the probability value of the node output to obtain a sampling value;
[0069] Step 750, determining the value of the derivative function of the sigmoid function of the node according to the sampling value;
[0070] Step 760, calculating the gradient value of the network model parameters according to the value of the derivative function of the sigmoid function of the node; and
[0071] Step 770: Update the parameters of the network model based on the gradient values of the parameters of the network model.
[0072] In some embodiments, step 740 may include: performing two consecutive samplings on the magnetic tunnel junction unit according to the probability values output by the nodes in the network model to obtain two sampling values. When the two consecutive sampling values are (1, 0), the value of the derivative function of the sigmoid function of the node may be determined to be 1 in step 750; when the two consecutive sampling values are (0, 0), (0, 1), or (1, 1), the value of the derivative function of the sigmoid function of the node is determined to be 0 in step 750.
[0073] In some embodiments, step 740 may include: performing 2n consecutive samplings on the magnetic tunnel junction unit according to the probability values output by the nodes in the network model to obtain n pairs of sampling values, where n is an integer greater than or equal to 1. In step 750, the ratio of the sampling values of (1, 0) in the n pairs of sampling values may be determined, and the ratio may be determined as the value of the derivative function of the sigmoid function of the node.
[0074] In some embodiments, the gradient value of the parameters of the network model is the partial derivative of the error function of the network model with respect to the parameters, and the product of the partial derivative and the learning rate is used to update the parameters.
[0075] It should be understood that the technical solution of the present invention is applicable to any random number generator with sigmoid-type probability tunability, not limited to magnetic tunnel junction true random number generators. As mentioned in the above solution, through two consecutive random sampling operations, and the sampling result of (1, 0) is recognized as the derivative function value of 1; the derivative function results in other sampling result cases (1, 1), (0, 1), and (0, 0) are 0. In fact, as long as the magnetic tunnel junction is in the same initial state, the driving voltage applied to the magnetic tunnel junction is the same, and the two random sampling operations are opposite to each other during these two random sampling operations. It is not necessary that they must be two consecutive random sampling operations. They can be any two independent random sampling operations performed according to the same sampling conditions; however, two consecutive random sampling operations can reduce the sampling time and improve the sampling efficiency.
[0076] Unless the context clearly requires otherwise, throughout the specification and claims, the words "comprise," "comprising," "include," "including," and the like shall be construed in an inclusive sense as opposed to an exclusive or exhaustive meaning. That is to say, they are to be interpreted as "including but not limited to." As commonly used herein, the term "connected" refers to two or more elements that may be directly connected or connected through one or more intermediate elements. As commonly used herein, the word "connected" means two or more elements that may be directly connected or connected through one or more intermediate elements. Additionally, when used in this application, the words "herein," "above," "below," and words of similar import shall refer to the entire application rather than to any particular part of the application. Where the context permits, the word "or" refers to a list of two or more items and encompasses all of the following interpretations of that word: any item in the list, all items in the list, and any combination of items in the list.
[0077] Furthermore, unless specifically stated otherwise or otherwise understood in the context in which it is used, conditional language used herein, such as "can," "could," "might," "may," "for example," "for instance," "such as," and the like, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements, and / or states. Thus, such conditional language is generally not intended to imply that one or more embodiments necessarily require any features, elements, and / or states, or that one or more embodiments must include logic for making a determination, with or without author input or prompting, as to whether such features, elements, and / or states are included in or are to be performed in any particular embodiment.
[0078] Although certain embodiments have been described, these embodiments have been presented by way of example only and are not intended to limit the scope of the disclosure. In fact, the novel facilities, methods, and systems described herein may be embodied in a variety of other forms; furthermore, various omissions, substitutions, and changes in the form of the methods and systems described herein may be made without departing from the spirit of the disclosure. For example, while blocks are presented in a given arrangement, alternative embodiments may perform similar functions with different components and / or circuit topologies, and some blocks may be deleted, moved, added, subdivided, combined, and / or modified. Each of these blocks may be implemented in a variety of different ways. Any suitable combination of the elements and acts of the various embodiments described above may be combined to provide further embodiments. The appended claims and their equivalents are intended to cover such forms or modifications that fall within the scope and spirit of the disclosure.
[0079] The foregoing description has been presented for purposes of illustration and description. In addition, this description is not intended to limit embodiments of the invention to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.
Claims
1. A probability computing accelerator, comprising: A computing unit, used to perform a network model operation using sample data to obtain a probability value of a node output of the network model, wherein the node of the network model uses a sigmoid function as an activation function; A sampling unit, used for sampling the magnetic tunnel junction unit according to the probability value output by the node to obtain a sampling value; as well as The parameter updating device comprises: A derivative function determining unit, used to determine the value of the derivative function of the sigmoid function of the node according to the sampling value obtained by the sampling unit; A gradient calculation unit, used to calculate the gradient value of the parameters of the network model according to the value of the derivative function of the sigmoid function of the node; and The parameter updating unit is used to update the parameters of the network model based on the gradient values of the parameters of the network model.
2. The probability calculation accelerator according to claim 1, further comprising a convergence judgment device, wherein the convergence judgment device comprises: An error calculation unit, used to calculate the error between the output value of the network model corresponding to the sample data and the true value; as well as A convergence judgment unit is used to judge whether the error has satisfied the convergence condition, If the error has satisfied the convergence condition, then the parameter updating device is stopped from updating the parameters of the network model; If the error has not yet satisfied the convergence condition, the parameter updating device continues to be used to update the parameters of the network model.
3. The probability calculation accelerator according to claim 1 or 2, wherein: The sampling unit samples the magnetic tunnel junction unit twice continuously according to the probability value of the node output in the network model to obtain two sampling values. When the two sampling values are (1, 0), the derivative function determining unit determines that the value of the derivative function of the sigmoid function of the node is 1, When the two sampling values are (0, 0), (0, 1) or (1, 1), the derivative function determining unit determines the value of the derivative function of the sigmoid function of the node to be 0.
4. The probability calculation accelerator according to claim 1 or 2, wherein: The sampling unit samples the magnetic tunnel junction unit 2n times continuously according to the probability value of the node output in the network model to obtain n pairs of sampling values, where n is an integer greater than or equal to 1. The derivative function determination unit determines a ratio of the sample values (1, 0) among the n pairs of sample values, and determines the ratio as a value of a derivative function of the sigmoid function of the node.
5. The probability calculation accelerator according to claim 1 or 2, wherein: The gradient calculation unit calculates the gradient value of the parameter of the network model as the partial derivative of the error function of the network model with respect to the parameter, The parameter updating unit updates the parameters of the network model using the product of the gradient value of the parameters of the network model and the learning rate.
6. A training method for a network model, wherein the nodes of the network model use a sigmoid function as an activation function, the training method comprising: Use sample data to perform operations on the network model to obtain probability values of node outputs of the network model; The magnetic tunnel junction unit is sampled according to the probability value output by the node to obtain a sampling value; Determine the value of the derivative function of the sigmoid function of the node according to the sampled value; Calculate the gradient value of the parameters of the network model according to the value of the derivative function of the sigmoid function of the node; and The parameters of the network model are updated based on the gradient values of the parameters of the network model.
7. The training method according to claim 6, further comprising: Calculating the error between the output value of the network model corresponding to the sample data and the true value; as well as Determine whether the error has satisfied the convergence condition, and if the error has satisfied the convergence condition, stop updating the parameters of the network model and complete the training of the network model; If the error has not satisfied the convergence condition, the steps from sampling the magnetic tunnel junction unit to updating the parameters of the network model are continued.
8. The training method according to claim 6 or 7, wherein: Sampling the magnetic tunnel junction unit according to the probability value of the node output includes: According to the probability value of the node output in the network model, the magnetic tunnel junction unit is sampled twice continuously to obtain two sampling values. When the two sample values are (1, 0), the value of the derivative function of the sigmoid function of the node is determined to be 1, When the two sampling values are (0, 0), (0, 1) or (1, 1), the value of the derivative function of the sigmoid function of the node is determined to be 0.
9. The training method according to claim 6 or 7, wherein: Sampling the magnetic tunnel junction unit according to the probability value of the node output includes: The magnetic tunnel junction unit is sampled 2n times continuously according to the probability value of the node output in the network model to obtain n pairs of sampling values, where n is an integer greater than or equal to 1. A ratio of the sample values (1, 0) in the n pairs of sample values is determined, and the ratio is determined as a value of a derivative function of the sigmoid function of the node.
10. The training method according to claim 6 or 7, wherein: The gradient value of the parameters of the network model is the partial derivative of the error function of the network model with respect to the parameters, and the product of the partial derivative and the learning rate is used to update the parameters.
Citation Information
Patent Citations
Spin random number generator with controllable probability
CN115809044A
True random number generator
CN116521132A
True random number generation device, self-calibration method thereof and electronic equipment
CN118550502A
Cited By
Neural network model training method and apparatus
CN122674782A