Methods and apparatuses for training and applying neural networks

By introducing a new initialization method into deep neural networks, training difficulties caused by parameter weight sharing are solved, and the training stability and performance of neural networks are improved.

CN113508401BActive Publication Date: 2025-05-27HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980093322.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-04-29
Publication Date
2025-05-27
Estimated Expiration
2039-04-29

AI Technical Summary

Technical Problem

In the prior art, when training deep neural networks, there are training difficulties due to parameter weight sharing, especially when using Xavier/He initialization, it is difficult to effectively update the highly shared parameters layer.

Method used

A new initialization method is proposed, suitable for compressed neural networks with structured matrices of shared elements. This method ensures a more balanced initialization by adjusting the learning rate of each layer of weight matrix and changing the initialization mode accordingly, and determining the adjustment parameters based on the number of shared elements.

Benefits of technology

Through this method, training difficulties caused by weight sharing can be effectively solved, and training stability and performance of neural networks can be improved, especially in compressed structures such as CirCNN networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113508401B_ABST
    Figure CN113508401B_ABST
Patent Text Reader

Abstract

A neural network training method, comprising: determining parameters of a weight matrix for each layer of the network according to the number of shared elements in a structured implementation having shared elements in each layer; determining the weight matrix for each layer according to the determined parameters; decompressing the determined weight matrix to derive an integral matrix; iteratively performing forward propagation and backward propagation on all layers of the network according to the integral matrix until a preset condition is met, wherein the integral matrix is updated in each iteration; and storing the iterative parameters of the network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to an artificial intelligence (AI) technology, and more particularly, to a method and apparatus for initializing a compressed neural network having a structured matrix with shared elements. Background Art

[0002] In recent years, deep learning methods have achieved remarkable success in a wide range of tasks such as object classification, natural language processing, and speech recognition. In 2016, the Go game software AlphaGo, which is supported by a deep learning algorithm, became the first Go game software to defeat a human champion in a 5-game Go series. How to effectively train a deep neural network has become a hot research field. In the pioneering work of Xavier and Yoshua, they observed that appropriate initialization of learnable parameters plays an important role in training a neural network. Later, He et al. extended the method of Xavier and Yoshua, taking into account the ReLU activation function. Nowadays, these two initialization methods are widely used in deep learning software packages such as TensorFlow, PyTorch, and Keras. Since these two initialization methods only differ by a factor called "gain", which is determined by the activation function, these two methods are regarded as one method and are called Xavier / He initialization.

[0003] The main idea of Xavier / He initialization is to maintain the activation variance and backpropagation gradient variance of different layers. However, when there is weight sharing for some parameters, the activation variance and backpropagation gradient are not good indicators of the variance of learnable parameters. Therefore, the update of highly shared parameters may be much faster than other parameters, which can lead to difficult training. If a layer of a neural network is multiplied by a positive constant and then the positive constant is divided into another layer, it is difficult to train the neural network even with Xavier / He initialization. Summary of the Invention

[0004] An initialization method based on Xavier / He initialization is proposed. For a fully connected layer without weight sharing, the embodiments in the present application are the same as the Xavier / He initialization method. However, these embodiments can handle weight sharing, which is a common phenomenon in the CirCNN implementation of a neural network, and various numerical embodiments are shown to verify the effectiveness of the initialization method in the present application.

[0005] The method is also used to adjust the learning rate of parameters layer by layer, multiplying the weight matrix of each layer by a positive constant scalar while correspondingly changing the initialization of that layer.

[0006] In the first embodiment of the present application, a neural network training method is disclosed, including: determining the parameters of the weight matrix of each layer of the network according to the number of shared elements in the structured implementation with shared elements in each layer; determining the weight matrix of each layer according to the determined parameters; decompressing the determined weight matrix to derive an integral matrix; iteratively performing forward propagation and backward propagation on all layers of the network according to the integral matrix until a preset condition is met, where the integral matrix is updated in each iteration; storing the iteration parameters of the network.

[0007] In a feasible implementation, the parameters include adjustment parameters, where the adjustment parameters are negatively correlated with the number of shared elements.

[0008] It should be noted that the calculation of the adjustment parameter can be 1 / B or 1 / (B×B), etc., and is not limited to this in the present application.

[0009] In a feasible implementation, the number of shared elements in the structured implementation with shared elements is the block size in the cyclic implementation.

[0010] It should be noted that the cyclic implementation is a structured implementation with shared elements, and the number of shared elements or block size is related to the compression ratio of the layer.

[0011] In a feasible implementation, the adjustment parameter is calculated as follows: where m is the number of input channels of a layer of the network, n is the number of output channels of the layer of the network, Cnm is the adjustment parameter, and Bnm is the number of shared elements.

[0012] In a feasible implementation, the parameters further include random numbers uniformly distributed between –Xavier_value and Xavier_value.

[0013] It should be noted that the uniform distribution is a candidate implementation and is not limited to this in the present application. Other distributions can also be used for the solution.

[0014] In a feasible implementation, Xavier_value is determined by the following formula:

[0015]

[0016] where gain is the composition number, v λ(n,m) is a random number.

[0017] In a feasible implementation, the weight matrix is determined by the following formula:

[0018] W nm= C nm ·v λ(n,m)

[0019] where W nm is the weight matrix mentioned above.

[0020] In a feasible implementation, the integration matrix is updated in each iteration, including the random number being updated in each iteration.

[0021] In a feasible implementation, the preset condition includes that the training result of the network converges.

[0022] In a feasible implementation, the training result of the network converges, including that the difference of the weight matrix is less than a preset threshold.

[0023] It should be noted that any criterion that can be used to judge convergence during the neural network training process can be applied in this application.

[0024] In a feasible implementation, the input during the training of the network includes image information data or sound information data.

[0025] In a feasible implementation, the network is used for classifying objects, processing languages, or recognizing voices.

[0026] Therefore, obviously, this application is applicable to this industry, such as the fields of object classification, natural language processing, and speech recognition.

[0027] In the second embodiment of this application, a neural network training device is disclosed, including: a first calculation module for determining the parameters of the weight matrix of each layer of the network according to the number of shared elements in the structured implementation with shared elements in each layer; a second calculation module for determining the weight matrix of each layer according to the determined parameters; a decompression module for decompressing the determined weight matrix to derive an integration matrix; an iteration module for iteratively performing forward propagation and backward propagation on all layers of the network according to the integration matrix until a preset condition is met, where the integration matrix is updated in each iteration; a storage module for storing the iteration parameters of the network.

[0028] In a feasible implementation, the parameters include adjustment parameters, where the adjustment parameters are negatively correlated with the number of shared elements.

[0029] In a feasible implementation, the number of shared elements in the structured implementation with shared elements is the block size in the cyclic implementation.

[0030] In a feasible implementation, the adjustment parameters are calculated as follows: Among them, m is the number of input channels of a layer of the network, n is the number of output channels of the layer of the network, Cnm is the adjustment parameter, and Bnm is the number of shared elements.

[0031] In a feasible implementation manner, the parameter further includes a random number uniformly distributed between -Xavier_value and Xavier_value.

[0032] In a feasible implementation manner, Xavier_value is determined by the following formula:

[0033]

[0034] Among them, gain is a composition number, and v λ(n,m) is a random number.

[0035] In a feasible implementation manner, the weight matrix is determined by the following formula:

[0036] W nm = C nm ·v λ(n,m)

[0037] Among them, W nm is the weight matrix.

[0038] In a feasible implementation manner, the integration matrix is updated in each iteration, including the random number being updated in each iteration.

[0039] In a feasible implementation manner, the preset condition includes that the training result of the network converges.

[0040] In a feasible implementation manner, the training result of the network converges, including that the difference of the weight matrix is less than a preset threshold.

[0041] In a feasible implementation manner, the input of the network during training includes image information data or sound information data.

[0042] In a feasible implementation manner, the network is used for classifying objects, processing languages, or recognizing voices.

[0043] In the third embodiment of the present application, a device for training a neural network is disclosed. The device includes: one or more processors; a non-transitory computer-readable storage medium, coupled to the processor and storing a program for the processor to execute. When the processor executes the program, the device executes the method according to any implementation manner in the first embodiment.

[0044] In a fourth embodiment of the present application, a computer program product is disclosed, including program code for performing the method according to any implementation manner in the first embodiment. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] To more fully understand the present invention and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which like numerals indicate like objects, and in the drawings:

[0046] Figure 1 shows the network losses summarized in Table 1;

[0047] Figure 2 shows the network losses summarized in formula (12a);

[0048] Figure 3 shows an exemplary method for neural network training;

[0049] Figure 4 shows an exemplary compression process of a block circulant matrix;

[0050] Figure 5 shows an exemplary block diagram of a neural network training device;

[0051] Figure 6 shows an exemplary block diagram of a neural network training apparatus. DETAILED DESCRIPTION

[0052] The following discussion Figures 1 to 6 and the various embodiments used to describe the principles of the present invention in this patent document are for illustration only and should not be construed in any way as limiting the scope of the present invention. Those skilled in the art will understand that the principles of the present invention can be implemented in any type of appropriately arranged device or system.

[0053] The following documents are hereby incorporated into the present invention as if fully set forth herein, and the reference numbers of these documents will be used in the following parts of this application.

[0054] [1] Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu and Pavel Kuksa, Natural language processing (almost) from scratch, Journal of Machine Learning Research, 12 (Aug): 2493–2537, 2011.

[0055] [2] Caiwen Ding, Siyu Liao, Yanzhi Wang, Zhe Li, Ning Liu, Youwei Zhuo, Chao Wang, Xuehai Qian, Yu Bai, Geng Yuan, etc., Circnn: Accelerating and Compressing Deep Neural Networks Using Block-Circulant Weight Matrices, Proceedings of the 50th IEEE / ACM International Symposium on Microarchitecture, pp. 395-408, ACM, 2017.

[0056] [3] Xavier Glorot and Yoshua Bengio, Understanding the Difficulty of Training Deep Feedforward Neural Networks, Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, pp. 249-256, 2010.

[0057] [4] Ian Goodfellow, Yoshua Bengio and Aaron Courville, Deep Learning, The MIT Press, 2016. http: / / www.deeplearningbook.org.

[0058] [5] Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun, Delving Deep into Rectifiers: Surpassing Human-Level Performance on Imagenet Classification, International Conference on Computer Vision (ICCV), December 2015.

[0059] [6] Geoffrey Hinton, Li Deng, Dong Yu, George Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Brian Kingsbury, etc., Deep neural networks for acoustic modeling in speech recognition, IEEE Signal Processing Magazine, 29th, 2012.

[0060] [7] Sergey Ioffe and Christian Szegedy, Batch normalization: Accelerating deep network training by reducing internal covariate shift, International Conference on Machine Learning, pp. 448 - 456, 2015.

[0061] [8] Alex Krizhevsky, Ilya Sutskever and Geoffrey E Hinton, Imagenet classification with deep convolutional neural networks, Advances in Neural Information Processing Systems, pp. 1097 - 1105, 2012.

[0062] [9] Dmytro Mishkin and Jiri Matas, All you need is a good init, International Conference on Learning Representations, 2016.

[0063]

[10] Andrew Saxe, James L McClelland and Surya Ganguli, Exact solutions to the nonlinear dynamics of learning in deep linear neural networks, 2013.

[0064]

[11] David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot et al., Mastering the game of go with deep neural networks and tree search nature, 529(7587):484, 2016.

[0065]

[12] Karen Simonyan and Andrew Zisserman, Very deep convolutional networks for large scale image recognition, arXiv preprint arXiv:1409.1556, 2014.

[0066]

[13] Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel Schoenholz and Jeffrey Pennington, Dynamical isometry and a mean field theory of cnns: How to train 10,000-layer vanilla convolutional neural networks, International Conference on Machine Learning, pp. 5389-5398, 2018.

[0067]

[14] Liang Zhao, Siyu Liao, Yanzhi Wang, Zhe Li, Jian Tang, and Bo Yuan, "Theoretical properties for neural networks with weight matrices of low displacement rank," Proceedings of the 34th International Conference on Machine Learning, volume 70, pages 4082-4090, JMLR.org, 2017.

[0068] Some notations and terms used in this application are introduced first.

[0069] I. Single Layer Structure

[0070] For a typical neural network, a layer of the network has the form

[0071] y = act(Wx + b); (1)

[0072] where, is the input vector, is the weight matrix, is the bias vector, and act is a point activation function (e.g., the identity function or the ReLU function ). The entries in W and b are usually parameters to be learned during the training process.

[0073] During the training process, the neural network will go through iterations of forward propagation and backward propagation. Assume the loss function is denoted by L. In forward propagation, y is calculated according to the given x. In backward propagation, according to the given calculate and Assume v is a learnable parameter of this layer. Then, after one backward propagation with a learning rate of r, v will be updated by the following formula

[0074]

[0075] where, will use and calculate

[0076] For any variable z in this layer, is used to represent the value after initialization, while is used to represent the value after the first forward propagation and backward propagation. Here, a definition is introduced. Assume the notations and the single layer structure are as shown in reference [1], and let z be a variable in this layer. Then, when the learning rate is set to 1, the increment of z (denoted as ) is defined as

[0077]

[0078] For the learnable variable v, make For the intermediate variable Wnm corresponding to the learnable variable v, if it is assumed that v only appears in W of this layer, then make

[0079]

[0080] where the last summation is over all terms in W that share the same learnable parameter v. If v also appears in the weights of other layers, the last summation also needs to include the corresponding partial derivatives.

[0081] II. CirCNN Network

[0082] Most state-of-the-art neural networks contain a large number of learnable parameters. For example, the VGG16 network introduced in reference

[12] contains more than 1 million learnable parameters. How to find an efficient method to reduce the number of parameters in deep neural networks has become a hot research area. A promising method is to replace the matrices and convolutions in deep neural networks with structured matrices and structured convolutions. The advantage of this method is that it can preserve the architecture of the network. In reference [2], it is proposed to replace the unstructured matrices in the network with block circulant matrices, and the network is called the CirCNN network. Therefore, a matrix is a circulant matrix if it has the following form.

[0083]

[0084] It should be noted that, different from unstructured matrices, a circulant matrix can be determined if the first row (or the first column) of the matrix is known. For a fully connected layer in the form of y = ReLU(Wx + b), if W consists of block matrices, then W is a block circulant matrix, where each block matrix is a circulant matrix. For a fully connected layer, it is assumed that the unstructured matrix is replaced with a block circulant matrix with a block size of B. Then, for this matrix, the number of parameters is reduced by a factor of B. The number B is called the compression ratio of this layer. It should be noted that for a layer, the number of parameters in the weight matrix W will determine the number of parameters in this layer.

[0085] It should be noted that the unstructured matrix can also be replaced with other structured implementations with shared elements, and the compression ratio B is related to the number of shared elements. The circulant implementation is a structured implementation with shared elements. Other structured implementations with shared elements include Topliize matrices, Hankel matrices, etc. A structured implementation of M×N with shared elements can be represented by fewer than M×N elements.

[0086] Compressing neural networks using block circulant matrices can significantly reduce the number of parameters. For block circulant matrices, fast algorithms such as FFT can be used to speed up matrix-vector multiplication. It has been shown in reference

[14] that if unstructured matrices are replaced with block circulant matrices, the universal approximation property can be retained, which guarantees the expressive power of the CirCNN implementation.

[0087] III. Xavier / He Initialization and Constraints

[0088] Network examples with parameter c indicate that Xavier / He initialization may not be sufficient to achieve fast and stable training.

[0089] For a fully connected layer of the form in Equation (1), the entries in the matrix W and the vector b are learnable parameters. Denoting the final loss function as L, the Xavier / He initialization has the following requirements:

[0090]

[0091] where, it is assumed that all entries in x follow the distribution X and all entries in y follow the distribution Y, all entries in all entries in Unless W is a square matrix, the above conditions cannot be satisfied simultaneously, but the following initialization rules are widely used: the bias vector b should be initialized to the zero vector, and the entries in W should be identically and independently distributed with mean 0 and variance where, gain is a factor determined by the activation function. For the identity function, gain should be 1, while for the ReLU activation, gain should be 2. In practice, the distribution of the entries in W is chosen to be uniform or normal.

[0092] Consider a network without bias of the following form

[0093] y = W 2 (ReLU(W 1 x)), (4)

[0094] where, is the input vector, and W1 and W2 are of sizes and respectively. The entries in the parameter matrices will be learned.

[0095] Decompose the above model into two networks, each with two layers, in the following two ways

[0096]

[0097]

[0098] Among them, c is a fixed positive scalar. For the same W1, W2, and x, these two decompositions will yield the same y. By applying Xavier / He initialization to both networks, it can be ensured that for both networks, the means and variances of the inputs and outputs of each layer are the same, and the means and variances of the partial derivatives of each layer are also the same. Conventional calculations show that the initialization should be as follows

[0099]

[0100] It can be seen that when c = 1, the two decompositions and initializations are the same. To test the influence of the positive scalar c, the above two networks were tested on the MNIST dataset (a collection of handwritten character digits). In this case, M = 784 and N = 10. Set c = 0.01, and use stochastic gradient descent (SGD) (without momentum) as the optimizer. The output y of each network is connected to a soft-max layer, and cross-entropy is used as the loss function in both cases (see reference [4] for details). The learning rates for these two cases were adjusted through trial and error to find the optimal values. It can be observed that the training of the second network (where c = 0.01) is much more difficult than that of the first network. Therefore, a much smaller learning rate needs to be used, which results in a slower convergence rate and a higher final loss value. Apparently, for the case of c = 0.01, there is no observable change in the distribution. This indicates that the parameters in W1 do not learn after iteration. This is because a very small learning rate and a small value of c = 0.01 need to be used. On the other hand, for the case of c = 1, the distribution evolves from a uniform distribution to a bell-shaped curved distribution. This indicates that the learning process of the parameters in W1 is successful. It should be noted that in the case of c = 0.01, the unstable training process cannot be remedied by batch normalization, which is introduced in reference [7] and is widely used to remedy unstable training.

[0101] A new method for initializing the fully connected layer is proposed in this application.

[0102] According to the described problem, in order to obtain stable training and better network performance, the network should be appropriately decomposed. To solve this problem, a new condition is proposed for Xavier / He initialization of layers whose weight matrices do not have sparsity, which can ensure a more balanced initialization.

[0103] Assume that the form of the fully connected layer of a multi-layer network is

[0104]

[0105] where α is a fixed positive scalar, is the input, is the output, is the deviation, is a set of learnable parameters, act is the activation function,

[0106] Input: {0, 1, …, N} × {0, 1, …, M} → {0, 1, …, Λ}

[0107] is a function that constructs the matrix W from the vector v. In this way, it is ensured that each term in W corresponds to an element in v, but allows multiple terms in W to share the same term in v. Let L denote the loss function (in the example described above, L is a combination of soft-max and cross-entropy). Only SGD (without momentum) is considered as the optimizer. The parameters to be trained are v and b.

[0108] Assume the notations in formula (7) and the single-layer structure. Any initialization of formula (7) should satisfy the following conditions

[0109]

[0110] where, denotes the term the distribution followed by.

[0111] It should be noted that for a standard fully connected layer, where all α equal 1, the terms in the matrix W have a one-to-one correspondence with the terms in v, and the last condition in formula (8) will be automatically satisfied. When considering a more general fully connected layer with parameter sharing, the importance of initializing formula (8) will become obvious.

[0112] For the split formula (5b), let α = c, the calculation shows

[0113]

[0114] To satisfy the initialization condition formula (8), c should be equal to 1. For c = 0.01, even if the initialization formulas (6a) and (6b) both satisfy Xavier / He initialization,

[0115] only the initialization formula (6a) satisfies the initialization condition, and it gives a better network decomposition.

[0116] In another embodiment, for the CirCNN network in reference [2], the fully connected layer can be represented as formula (7), where α = 1. Routine calculations show

[0117]

[0118] denotes V nm = {W n′m′: {λ(n′, m′) = λ(n, m)}, and use Bnm to represent the number of elements in Vnm. It should be noted that weight sharing comes from the circulant matrix in W.

[0119] Suppose

[0120]

[0121] is irrelevant, making

[0122]

[0123] The only way to satisfy the initialization condition formula (8) is to make Bnm = 1. However, this corresponds to a circulant matrix of size 1×1, which simply means that CirCNN does not compress the original network at all.

[0124] To achieve the desired compression ratio, it is necessary to introduce an additional positive scalar α and reconstruct the fully connected layer in the form of formula (7) using CirCNN in reference [2]. It should be emphasized that the scalar α will be constant for this layer, and the learnable variables are still v and b. Therefore, by introducing α, the number of learnable parameters does not increase, and the increase in the number of operations for this layer can be ignored.

[0125] Suppose the compression factor B of this layer is required. Then the size of the circulant matrix in W should be B×B. The previous calculations show

[0126]

[0127] where Bnm = B. Then, to achieve the initialization condition formula (8), set

[0128]

[0129] where gain should be determined by the activation function. To see the effect of the positive scalar α on the update of v λ(n,m) , the calculation is as follows. Suppose the learning rate of the SGD optimizer is r, then after one forward propagation and one backward propagation, make

[0130]

[0131] indicating that the effective learning rate of v λ(n,m) now changes from r to αr. Therefore, by formulating each layer as in formula (7), the effective learning rate of the learnable parameters can be changed layer by layer.

[0132] The effectiveness of the proposed initialization method is demonstrated through numerical experiments.

[0133] Example 1. shows the training process of a simple network using the initialization formula (10) on the MNIST dataset. The network structure is summarized in Table 1, where the first fully connected layer is compressed using CirCNN. It should be noted that after compression, the total number of learnable parameters is only about 0.76% of the original model. From Figure 1 the dot plot of the loss function in

[0134]

[0135] Table 1 Network structure of Example 1. For compression ratio B >= 1, a cyclic implementation with block size B. The number of parameters implemented by CirCNN should be divided by B.

[0136] Example 2. A simple network where the weight matrices of many layers share a common weight matrix V. The details of the network and the initialization of the common weight V by the proposed formula (8) are summarized in formulas (12a) and (12b)

[0137]

[0138]

[0139] where, as before, for the proposed initialization, the weight matrix has the form W i = αV, where i = 1, 2,..., 50. Figure 2 summarizes the comparison with the initialization (12b) and Xavier / He initialization.

[0140] Example 3. Tests the proposed initialization method on the CirCNN implementation of the VGG16 network. For the second fully connected layer in the VGG16 network with 4096 input channels and 4096 output channels, a cyclic matrix with a block size of 4096 is used. Tests the Cifar10 dataset for 5 runs with Xavier / He initialization and the proposed initialization method. The original VGG16 network can achieve a top-1 accuracy of 92.24%. In 5 runs, the network using the proposed initialization is consistently better than the network using Xavier / He initialization in terms of top-1 accuracy. The average top-1 accuracy of the proposed method is 92.00%, while the average top-1 accuracy of Xavier / He initialization is 90.98%.

[0141] Figure 3 shows a method for neural network training. The training scheme using the proposed initialization method can be summarized as follows:

[0142] S101: Determine the parameters of the initialization weight matrix of the layer according to the block size in the recurrent implementation of the neural network layer.

[0143] It should be noted that for another structured implementation with shared elements, S101 includes determining the parameters of the weight matrix of each layer of the network according to the number of shared elements in the structured implementation with shared elements in each layer.

[0144] In one implementation, the parameters include:

[0145]

[0146] Among them, for the identity function, gain is equal to 1, and when initializing with ReLU, gain is equal to 2. M or m represents the number of input channels, N or n represents the number of output channels, and Bnm represents the number of parameter sharings of Wnm, which depends on the parameter sharing method. For example, the above-mentioned circulant matrix is a matrix with a parameter sharing form. In the circnn network, where B is the block size in the layer recurrent implementation. Xavier_value is equal to used to generate v λ(n,m) , that is, a random number uniformly distributed between -Xavier_value and Xavier_value, that is, [-Xavier_value, Xavier_value].

[0147] It should be noted that Cnm is negatively correlated with the number of shared weights. The more weight sharing, the smaller the value of Cnm. In this case, the learning rate of the shared weights becomes lower, and the stability of training the neural network is improved.

[0148] In some other embodiments, Cnm can be equal to 1 / Bnm, 1 / (Bnm×Bnm), etc.

[0149] S102: Determine the initialization weight matrix according to the parameters.

[0150] More specifically,

[0151] W nm = C nm ·v λ(n,m)

[0152] S103: Decompress Wnm to derive the integral block circulant matrix Wt.

[0153] Figure 4 An example shows the compression of the block circulant matrix. Decompression is the inverse process of compression. In lossless compression, decompression is an absolute inverse process.

[0154] S104: Forward propagation of training.

[0155] In one implementation, the training data is trained by formula (1), Y = Wt·X + B, where X represents the input data, Y represents the output data, B represents the bias, and Wt is derived from step S103. It should be noted that when the layer is a convolutional layer, X and B are the data converted by the "im2col" operation.

[0156] S105: Backpropagation of the training.

[0157] In one implementation, the output data of step S104 is trained by the following formula:

[0158]

[0159] where L is the loss function of X and Y, and W′ t is the updated value of Wt. Correspondingly, v λ(n,m) is updated.

[0160] S106: Perform iterations of steps S104 and S105 until the training result converges.

[0161] It should be noted that the convergence criterion in this application is not limited to this. For example, the criterion can be that the difference between any parameters (v λ(n,m) 、Wt, loss, etc.) involved in the training between two adjacent iterations is less than a threshold.

[0162] S107: Store the trained parameters of the neural network.

[0163] In one implementation, steps S101 - S103 are implemented on each layer of the neural network, and steps S104 - S107 are implemented based on the entire network.

[0164] It should be noted that the training data can be images, videos, voice signals, shapes, colors, and other information to be recognized. Therefore, the trained neural network can be used in related applications, such as object classification, natural language processing, and speech recognition mentioned above.

[0165] In an embodiment of this application, a neural network training method is disclosed, including: determining the parameters of the weight matrix of each layer of the network according to the number of shared elements in the structured implementation with shared elements in each layer; determining the weight matrix of each layer according to the determined parameters; decompressing the determined weight matrix to derive an integral matrix; iteratively performing forward propagation and backward propagation on all layers of the network according to the integral matrix until a preset condition is met, where the integral matrix is updated in each iteration; storing the iterative parameters of the network.

[0166] In a feasible implementation, the parameter includes an adjustment parameter, where the adjustment parameter is negatively correlated with the number of shared elements.

[0167] In a feasible implementation, the number of shared elements in the structured implementation with shared elements is the block size in the cyclic implementation.

[0168] In a feasible implementation, the adjustment parameter is calculated as follows: where m is the number of input channels of a layer of the network, n is the number of output channels of the layer of the network, Cnm is the adjustment parameter, and Bnm is the number of shared elements.

[0169] In a feasible implementation, the parameter further includes a random number uniformly distributed between -Xavier_value and Xavier_value.

[0170] In a feasible implementation, Xavier_value is determined by the following formula:

[0171]

[0172] where gain is the composition number, v λ(n,m) is a random number.

[0173] In a feasible implementation, the weight matrix is determined by the following formula:

[0174] W nm = C nm ·v λ(n,m)

[0175] where W nm is the weight matrix.

[0176] In a feasible implementation, the integration matrix is updated in each iteration, including the random number being updated in each iteration.

[0177] In a feasible implementation, the preset condition includes that the training result of the network is convergent.

[0178] In a feasible implementation, the training result of the network being convergent includes that the difference of the weight matrix is less than a preset threshold.

[0179] In a feasible implementation, the input of the network during training includes image information data or sound information data.

[0180] In a feasible implementation, the network is used for classifying objects, processing languages, or recognizing voices.

[0181] In one embodiment of the present application, as Figure 5 shown, a neural network training device 400 is disclosed, including: a first calculation module 401, configured to determine parameters of the weight matrix of each layer of the network according to the number of shared elements in the structured implementation with shared elements in each layer; a second calculation module 402, configured to determine the weight matrix of each layer according to the determined parameters; a decompression module 403, configured to decompress the determined weight matrix to derive an integral matrix; an iteration module 404, configured to iteratively perform forward propagation and backward propagation on all layers of the network according to the integral matrix until a preset condition is met, where the integral matrix is updated in each iteration; and a storage module 405, configured to store the iteration parameters of the network.

[0182] In a feasible implementation manner, the parameters include adjustment parameters, where the adjustment parameters are negatively correlated with the number of shared elements.

[0183] In a feasible implementation manner, the number of shared elements in the structured implementation with shared elements is the block size in the cyclic implementation.

[0184] In a feasible implementation manner, the adjustment parameters are calculated as follows: where m is the number of input channels of a layer of the network, n is the number of output channels of the layer of the network, Cnm is the adjustment parameter, and Bnm is the number of shared elements.

[0185] In a feasible implementation manner, the parameters further include random numbers uniformly distributed between -Xavier_value and Xavier_value.

[0186] In a feasible implementation manner, Xavier_value is determined by the following formula:

[0187]

[0188] where gain is the composition number, and v λ(n,m) is a random number.

[0189] In a feasible implementation manner, the weight matrix is determined by the following formula:

[0190] W nm = C nm ·v λ(n,m)

[0191] where W nm is the weight matrix.

[0192] In a feasible implementation, the integration matrix is updated in each iteration, including that the random number is updated in each iteration.

[0193] In a feasible implementation, the preset condition includes that the training result of the network is convergent.

[0194] In a feasible implementation, that the training result of the network is convergent includes that the difference value of the weight matrix is less than a preset threshold.

[0195] In a feasible implementation, the input in the training of the network includes image information data or sound information data.

[0196] In a feasible implementation, the network is used for classifying objects, processing languages or recognizing voices.

[0197] In an embodiment of the present application, a device for training a neural network is disclosed. The device includes: one or more processors; a non-transitory computer-readable storage medium, coupled to the processor and storing a program for the processor to execute. When the processor executes the program, the device executes the method for training a neural network according to the present application.

[0198] In an embodiment of the present application, a computer program product is disclosed, including program code for executing the method for training a neural network according to the present application.

[0199] Figure 6 It is a simplified block diagram of device 500 which can be used as a device for training a neural network provided by the exemplary embodiment.

[0200] The processor 502 in device 500 may be a central processing unit. Alternatively, the processor 502 may be any other type of device or multiple devices that can manipulate or process information, existing or to be developed in the future. Although a single processor such as processor 502 shown in the figure can be used to implement the disclosed implementation, using more than one processor can improve speed and efficiency.

[0201] In one implementation, the memory 504 in the apparatus 500 may be a read only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 that can be accessed by the processor 502 via the bus 512. The memory 504 may also include an operating system 508 and application programs 510, where the application programs 510 include at least one program that allows the processor 502 to execute the methods described herein. For example, the application programs 510 may include Applications 1 to N, and may also include a video decoding application that executes the methods described herein.

[0202] The apparatus 500 may also include one or more output devices, such as a display 518. In one example, the display 518 may be a touch-sensitive display that combines a display with a touch-sensitive element that can be used to sense touch inputs. The display 518 may be coupled to the processor 502 via the bus 512.

[0203] Although the bus 512 of the apparatus 500 is described herein as a single bus, the bus 512 may include multiple buses. In addition, the auxiliary memory 514 may be directly coupled to other components of the apparatus 500 or may be accessible via a network, and may include a single integrated unit (such as a memory card) or multiple units (such as multiple memory cards). Therefore, the apparatus 500 may be implemented in a variety of configurations.

[0204] In some embodiments, some or all of the functions or processes of one or more devices are implemented or supported by a computer program formed of computer-readable program code and embodied in a computer-readable medium. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer, such as a read only memory (ROM), a random access memory (RAM), a hard disk drive, a compact disc (CD), a digital video disc (DVD), or any other type of memory.

[0205] It may be advantageous to set forth definitions of some of the words and phrases used in this patent document. The terms "comprising" and "including" and their derivatives mean including but not limited to. The term "or" is inclusive and means and / or. The phrases "associated with" and "associated therewith" and their derivatives mean including, included within, interconnected with, containing, contained within, connected to or connected with, coupled to or coupled with, communicating with, cooperating with, interlacing, juxtaposed, proximate, bound to or bound with, having, owning, and the like.

[0206] While the invention has been described with respect to certain embodiments and the methods generally associated therewith, changes and permutations of these embodiments and methods will be apparent to those skilled in the art. Accordingly, the foregoing description of the exemplary embodiments does not define or limit the invention. Other changes, substitutions and alterations are also possible without departing from the scope of the invention, as defined by the following claims.

Claims

1. A neural network training method, characterized in that, comprising: determining parameters of the weight matrix of each layer of the network according to the number of shared elements in the structured implementation with shared elements in each layer, the network being a compressed neural network, and the input in the training of the network including image information data or sound information data; determining the weight matrix of each layer according to the determined parameters; decompressing the determined weight matrix to derive an integral matrix; iteratively performing forward propagation and backward propagation on all layers of the network according to the integral matrix until a preset condition is satisfied, wherein the integral matrix is updated in each iteration; storing the iterative parameters of the network.

2. The method according to claim 1, characterized in that, the parameters include adjustment parameters, wherein the adjustment parameters are negatively correlated with the number of shared elements.

3. The method according to claim 1 or 2, characterized in that, the number of shared elements in the structured implementation with shared elements is the block size in the cyclic implementation.

4. The method according to claim 2, characterized in that, The adjustment parameter is calculated as follows: , where m is the number of input channels of a layer of the network, n is the number of output channels of the layer of the network, Cnm is the adjustment parameter, and Bnm is the number of shared elements.

5. The method according to claim 4, characterized in that, the parameters further include random numbers uniformly distributed between -Xavier_value and Xavier_value.

6. The method according to claim 5, characterized in that, Xavier_value is determined by the following formula: * Among them, , , gain is a component number, is a random number.

7. The method according to claim 6, characterized in that, the weight matrix is determined by the following formula: Among them, is the weight matrix.

8. The method according to any one of claims 5 to 7, characterized in that, the integral matrix is updated in each iteration, including the random numbers being updated in each iteration.

9. The method according to any one of claims 5 to 8, characterized in that, the preset condition includes that the training result of the network is convergent.

10. The method according to claim 9, characterized in that, the training result of the network being convergent includes that the difference of the weight matrix is less than a preset threshold.

11. The method according to any one of claims 1 to 10, characterized in that, the network is used for classifying objects, processing languages or recognizing voices.

12. A neural network training device, characterized in that, comprising: a first calculation module for determining parameters of the weight matrix of each layer of the network according to the number of shared elements in the structured implementation with shared elements in each layer, the network being a compressed neural network, and the input in the training of the network including image information data or sound information data; a second calculation module for determining the weight matrix of each layer according to the determined parameters; a decompression module for decompressing the determined weight matrix to derive an integral matrix; an iteration module for iteratively performing forward propagation and backward propagation on all layers of the network according to the integral matrix until a preset condition is satisfied, wherein the integral matrix is updated in each iteration; a storage module for storing the iterative parameters of the network.

13. The device according to claim 12, characterized in that, The parameter includes an adjustment parameter, wherein the adjustment parameter is negatively correlated with the number of the shared elements.

14. The device according to claim 12 or 13, wherein, the number of the shared elements in the structured implementation having the shared elements is the block size in the cyclic implementation.

15. The device according to claim 13, wherein, The adjustment parameter is calculated as follows: , where m is the number of input channels of a layer of the network, n is the number of output channels of the layer of the network, Cnm is the adjustment parameter, and Bnm is the number of shared elements.

16. The device according to claim 15, wherein, the parameter further includes a random number uniformly distributed between –Xavier_value and Xavier_value.

17. The device according to claim 16, wherein, Xavier_value is determined by the following formula: * Among them, , , gain is a component number, is a random number.

18. The device according to claim 17, wherein, the weight matrix is determined by the following formula: Among them, is the weight matrix.

19. The device according to any one of claims 16 to 18, wherein, the integration matrix is updated in each iteration, including the random number being updated in each iteration.

20. The device according to any one of claims 16 to 19, wherein, the preset condition includes that the training result of the network is convergent.

21. The device according to claim 20, wherein, the training result of the network being convergent includes that the difference of the weight matrix is less than a preset threshold.

22. The device according to any one of claims 12 to 21, wherein, the network is used for classifying objects, processing languages or recognizing voices.

23. An apparatus for training a neural network, wherein, the apparatus includes: one or more processors; a non-transitory computer-readable storage medium coupled to the processor and storing a program executed by the processor, wherein when the processor executes the program, the apparatus is caused to execute the method according to any one of claims 1 to 11.

24. A computer program product, wherein, it includes program code for executing the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Deep neural network learning method, processor and deep neural network learning system

    CN104899641A

  • Image processing method and device as well as equipment

    CN107547773A