Neural network training method and device, equipment, storage medium and product
By using matrix expansion and matrix point multiplication operations in neural network training, combined with the update of feedforward network weight errors, the problem of large storage and computing requirements in the existing technology is solved, and more efficient resource utilization is achieved.
Patent Information
- Application Number
- CN202510211701.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
The existing neural network training algorithms have large storage requirements and complex computing requirements, resulting in large resource overhead.
By determining the output signal error of the convolution layer and matrix expansion according to the error, then performing matrix point multiplication with the weights of each channel of the feedback network, the input signal error of the convolution layer is obtained, and the weight error of the feedforward network is determined and the weight is updated.
It improves the accuracy and computing efficiency of data propagation, reduces storage and computing requirements, and reduces resource overhead.
Smart Images

Figure CN120146137A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of machine learning and artificial intelligence, and particularly relates to a neural network training method, apparatus, device, storage medium, and product. Background Art
[0002] Neural network training, as a key force driving the development of the field of artificial intelligence, its core lies in continuously adjusting the network weights through the error backpropagation algorithm to achieve high-precision prediction or classification. However, this training mechanism encounters a significant challenge in biology, that is, the biological infeasibility of weight transmission. To address this challenge, the feedback alignment algorithm emerged, which abandons the practice of reusing the feedforward network weight information in error backpropagation and instead uses random weight information for error backpropagation.
[0003] However, the feedback alignment algorithm also has some key defects. This algorithm significantly increases the storage requirements because it needs to maintain both the feedforward weights and the feedback weights simultaneously. These two sets of weights have the same dimension, resulting in doubling of the memory occupancy. And the feedback alignment algorithm does not bring a significant improvement in computational efficiency. It only calculates the error backpropagation signal using different feedback weights and does not simplify any part of the calculation process, which makes the computational requirements for neural network training still huge and complex. Summary of the Invention
[0004] The main purpose of this application is to provide a neural network training method, apparatus, device, storage medium, and product, aiming to solve the technical problem that the existing neural network training algorithms have large resource overheads due to large storage requirements and complex computational requirements.
[0005] To achieve the above purpose, this application proposes a neural network training method, and the method includes:
[0006] Determine the error of the output signal of the convolutional layer, and perform matrix expansion on the sum of the output signal errors of the convolutional layer according to the error of the output signal of the convolutional layer to obtain the expanded sum of the output signal errors of the convolutional layer;
[0007] Obtain the error of the input signal of the convolutional layer by performing matrix dot product operation on the expanded sum of the output signal errors of the convolutional layer and the weights of each channel of the feedback network;
[0008] Determine the error of the feedforward network weights according to the error of the output signal of the convolutional layer and the error of the input signal of the convolutional layer;
[0009] Update the feedforward network weights of the neural network based on the error of the feedforward network weights to obtain the trained neural network.
[0010] In one embodiment, the step of determining the error of the output signal of the convolutional layer based on the loss function and obtaining the sum of the errors of the output signal of the convolutional layer according to all channel signals of the output signal error of the convolutional layer includes:
[0011] Obtaining the error of the forward propagation result based on the loss function, and obtaining the error of the input signal of the fully connected layer according to the error of the forward propagation result and the feedback weight of the fully connected layer of the feedback network;
[0012] Processing the error of the input signal of the fully connected layer based on the spatial dimension of the output signal of the pooling layer to obtain the error of the output signal of the pooling layer;
[0013] Performing an unpooling operation on the error of the output signal of the pooling layer to obtain the error of the output activation signal of the convolutional layer;
[0014] Determining the error of the output signal of the convolutional layer according to the error of the output activation signal of the convolutional layer and the derivative of the activation function, and obtaining the sum of the errors of the output signal of the convolutional layer according to all channel signals of the error of the output signal of the convolutional layer.
[0015] In one embodiment, the step of obtaining the forward propagation result of the input signal and determining the loss function of the input signal according to the forward propagation result includes:
[0016] Performing a convolution operation on the input signal and the weights of the convolutional layer of the feedforward network to obtain the output signal of the convolutional layer of the feedforward network;
[0017] Obtaining the input signal of the fully connected layer of the feedforward network according to the output signal of the convolutional layer of the feedforward network;
[0018] Performing a vector matrix multiplication operation on the input signal of the fully connected layer of the feedforward network and the weights of the fully connected layer of the feedforward network to obtain the output signal of the fully connected layer of the feedforward network;
[0019] Processing the output signal of the fully connected layer of the feedforward network through an activation function to obtain the forward propagation result of the input signal, and determining the loss function according to the forward propagation result and the target vector of the input signal.
[0020] In one embodiment, the step of obtaining the input signal of the fully connected layer in the feedforward network according to the output signal of the convolutional layer of the feedforward network includes:
[0021] Processing the output signal of the convolutional layer of the feedforward network through an activation function to obtain the output activation signal of the convolutional layer of the feedforward network;
[0022] Performing a pooling operation on the output activation signal of the convolutional layer of the feedforward network to obtain the output signal of the pooling layer;
[0023] Perform a flattening operation on the input signal of the next layer based on the output signal of the pooling layer to obtain the input signal of the fully connected layer in the feedforward network.
[0024] In one embodiment, the step of updating the feedforward network weights of the neural network based on the feedforward network weight error to obtain the trained neural network includes:
[0025] Based on the feedforward network weight error, update the feedforward network weights of the neural network through a weight update algorithm to obtain the initially trained neural network;
[0026] Determine the loss function. If the loss function does not meet the preset target loss value, continue to update the feedforward network weights of the initially trained neural network;
[0027] If the loss function meets the preset target loss value, stop updating the feedforward network weights of the initially trained neural network and obtain the trained neural network.
[0028] In addition, to achieve the above object, the present application also proposes a neural network training device, which includes:
[0029] An output signal error determination module, configured to determine the output signal error of the convolutional layer and perform matrix expansion on the sum of the output signal errors of the convolutional layer according to the output signal error of the convolutional layer to obtain the expanded sum of the output signal errors of the convolutional layer;
[0030] An input signal error determination module, configured to obtain the input signal error of the convolutional layer by performing a matrix dot product operation on the expanded sum of the output signal errors of the convolutional layer and the weights of each channel of the feedback network;
[0031] A weight error determination module, configured to determine the feedforward network weight error according to the output signal error of the convolutional layer and the input signal error of the convolutional layer;
[0032] A network training module, configured to update the feedforward network weights of the neural network based on the feedforward network weight error to obtain the trained neural network.
[0033] In addition, to achieve the above object, the present application also proposes a neural network training device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the neural network training method as described above.
[0034] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the neural network training method described above are implemented.
[0035] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the neural network training method described above are implemented.
[0036] The technical solution proposed by the present application is based on determining the error of the output signal of the convolutional layer, and expanding the sum of the output signal errors of the convolutional layer according to the error of the output signal of the convolutional layer to obtain the expanded sum of the output signal errors of the convolutional layer. By performing a matrix dot product operation on the expanded sum of the output signal errors of the convolutional layer and the weights of each channel of the feedback network, the error of the input signal of the convolutional layer is obtained. According to the error of the output signal of the convolutional layer and the error of the input signal of the convolutional layer, the weight error of the feedforward network is determined, and the weights of the feedforward network of the neural network are updated based on the weight error of the feedforward network to obtain the trained neural network. Through matrix expansion, matrix dot product operation, and weight update according to the weight error of the feedforward network, the accuracy of data propagation and the calculation efficiency are improved, the storage and calculation requirements are reduced, and the resource overhead is reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0038] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the neural network training method of the present application;
[0040] Figure 2 It is a schematic diagram of the summation and matrix expansion of the output signal error of the convolutional layer in the neural network training method of the present application;
[0041] Figure 3 It is a schematic diagram of the first matrix expansion method in the neural network training method of the present application;
[0042] Figure 4 It is a schematic diagram of the second matrix expansion method in the neural network training method of the present application;
[0043] Figure 5 Schematic diagram for obtaining the error of the input signal of the convolutional layer in the neural network training method of the present application;
[0044] Figure 6 Schematic flow chart provided by the second embodiment of the neural network training method of the present application;
[0045] Figure 7 Schematic diagram for digital recognition in the neural network training method of the present application;
[0046] Figure 8 Schematic diagram for the hardware implementation of the convolutional computing memristor in the present application;
[0047] Figure 9 Schematic diagram for the hardware implementation of the matrix-matrix element product operator memristor in the present application;
[0048] Figure 10 Schematic diagram of the module structure of the neural network training device in the embodiment of the present application;
[0049] Figure 11 Schematic diagram of the device structure of the hardware operating environment involved in the neural network training method in the embodiment of the present application.
[0050] The implementation, functional features and advantages of the present application will be further described in conjunction with the embodiments with reference to the accompanying drawings. Detailed implementation manners
[0051] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0052] In order to better understand the technical solutions of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0053] The existing feedback alignment algorithm has some key defects. This algorithm significantly increases the storage requirements because it needs to maintain both the feedforward weights and the feedback weights simultaneously, and these two sets of weights have the same dimension, resulting in doubling of the memory occupation. Moreover, the feedback alignment algorithm does not bring a significant improvement in computational efficiency. It only calculates the error backpropagation signal using different feedback weights and does not simplify the calculation process at all, which makes the computational requirements for neural network training still huge and complex.
[0054] Therefore, in order to overcome the above defects, the present application provides a solution, which improves the accuracy of data propagation and computational efficiency, reduces the storage and computational requirements, and reduces the resource overhead by matrix expansion, matrix dot product operation, and updating weights according to the feedforward network weight error.
[0055] It should be noted that the execution subject of each embodiment of this application can be a computing service system with data processing, network communication, and program running functions, such as an electronic system, a neural network training system, etc. that can implement the above functions. Hereinafter, taking a neural network training system as an example (hereinafter referred to as "system"), the following embodiments will be described.
[0056] Based on this, an embodiment of this application provides a neural network training method. Referring to Figure 1 , Figure 1 is a schematic flowchart of the first embodiment of the neural network training method of this application.
[0057] In this embodiment, the neural network training method includes steps S10 to S40:
[0058] Step S10, determine the convolution layer output signal error, and perform matrix expansion on the convolution layer output signal error sum according to the convolution layer output signal error to obtain the expanded convolution layer output signal error sum.
[0059] It should be noted that in the neural network training of deep learning, the convolution layer, as one of the core components, undertakes the key tasks of feature extraction and pattern recognition. In order to continuously optimize the performance of the neural network, this application accurately calculates the output signal error of the convolution layer and adjusts the network weights accordingly. Among them, a crucial step is to perform matrix expansion on the sum of the convolution layer output signal errors according to the convolution layer output signal error to obtain an error matrix that matches the convolution input signal and the feedback network weights.
[0060] It should be understood that the steps of determining the convolution layer output signal error and performing matrix expansion on the convolution layer output signal error sum according to this error to obtain the expanded convolution layer output signal error sum refer to that in the process of neural network training, first calculate the partial derivative of the loss function with respect to the signal output by the convolution layer. This value reflects the difference direction between the actual output signal and the expected output signal of the convolution layer. The convolution layer output signal error sum is the accumulation of this error on all channels. Subsequently, in order to perform matching operations with the spatial dimension of the convolution layer input signal and the weights of the feedback network, it is necessary to perform matrix expansion on this error sum, that is, according to the spatial dimension of the input signal (for example, the width, height, and color channel number of an image), adjust the size of the error sum matrix through zero-padding expansion or edge expansion, etc., so that its dimension is consistent with the weight matrix dimension of the feedback network.
[0061] Step S20, obtain the convolution layer input signal error by performing matrix dot product operation on the expanded convolution layer output signal error sum and the weights of each channel of the feedback network.
[0062] It can be understood that the process of obtaining the error of the input signal of the convolutional layer by performing a matrix dot product operation on the sum of the output signal errors of the dilated convolutional layer and the weights of each channel of the feedback network is part of the novel feedback alignment algorithm (improved feedback alignment algorithm) proposed in the neural network training method of this application. In this process, first, the sum of the output signal errors of the dilated convolutional layer is subjected to a matrix dot product operation with the weights of each channel of the feedback network (also known as the transposed convolutional layer or deconvolutional layer). The matrix dot product operation refers to the operation of multiplying corresponding elements and then summing them, which requires the dimensions of the two matrices to be the same. Through this operation, the error at the output end of the convolutional layer can be "feedback-aligned" to the input end of the convolutional layer, that is, the error of the input signal of the convolutional layer is calculated. By continuously iterating this process, the weights of the neural network can be gradually optimized to better fit the training data.
[0063] Step S30: Determine the feedforward network weight error according to the output signal error of the convolutional layer and the input signal error of the convolutional layer.
[0064] It should be understood that in this process, the error at the output end of the convolutional layer (i.e., the output signal error of the convolutional layer) and the error that propagates from the output end of the convolutional layer to the input end of the convolutional layer through the feedback network (i.e., the input signal error of the convolutional layer) have been obtained through feedback alignment. Next, using these two error signals, the activation function of the convolutional layer, and the input signal, the gradient or error of the feedforward network (i.e., the weights of the convolutional layer) is calculated through the chain rule. This gradient reflects the contribution degree of the current weight to the output error of the network, that is, the direction and magnitude in which the weight should be adjusted.
[0065] Specifically, the system uses the input signal error of the convolutional layer, the input signal of the convolutional layer (or its derivative) processed by the activation function, and the specific shape and size of the convolutional kernel (i.e., the weight) to calculate the gradient of the weight through a series of mathematical operations (such as convolution, matrix multiplication, element multiplication, etc.). Once the gradient of the weight is obtained, an optimization algorithm (such as gradient descent, stochastic gradient descent, etc.) can be used to update the weights of the feedforward network. This process will be continuously iterated until the performance of the network on the training set reaches the preset stop condition (such as the loss function converges, the number of iterations reaches the upper limit, etc.).
[0066] Step S40: Update the weights of the feedforward network of the neural network based on the feedforward network weight error to obtain the trained neural network.
[0067] It can be understood that, based on the feedforward network weight error (i.e., the gradient of the weights) calculated in the above steps, and the selected optimization algorithm (such as gradient descent, Adam, RMSProp, etc.), the direction and step size of weight update are determined. The optimization algorithm will calculate a weight update amount according to the magnitude and sign of the weight error, as well as possible other factors (such as learning rate, momentum, etc.). Then, this update amount is applied to the current feedforward network weights to obtain the updated weight values. This process usually traverses and updates all the weights in the network to ensure that the entire network can be adjusted and optimized according to the current training data. By continuously iterating this process, that is, continuously calculating the weight error according to the training data, updating the weights, and then calculating the new weight error, the feedforward network weights of the neural network can gradually converge to an optimal or approximately optimal solution. This solution can minimize (or approach the minimum) the loss function of the network on the training set, and at the same time, it is hoped that the network can exhibit good generalization ability on the unseen test set. Finally, when the training process is stopped, a trained neural network is obtained. This network has learned the effective mapping relationship from the input data to the output labels and can be used for subsequent inference or prediction tasks.
[0068] Further, in this embodiment, the above step S40 may include: based on the feedforward network weight error, updating the feedforward network weights of the neural network through a weight update algorithm to obtain an initially trained neural network; determining a loss function, and if the loss function does not meet the preset target loss value, continuing to update the feedforward network weights of the initially trained neural network; if the loss function meets the preset target loss value, stopping updating the feedforward network weights of the initially trained neural network and obtaining a trained neural network.
[0069] It should be understood that the system uses the novel feedback alignment algorithm in the neural network training method of the present application to calculate the feedforward network weight error. According to this error and the selected weight update algorithm, the update amount of the weights is calculated and applied to the current weights, thus obtaining the updated weight values, that is, completing one weight update iteration. After this iteration, an initially trained neural network is obtained. However, at this time, the network may not have reached the best performance because its loss function value (an index measuring the difference between the network output and the target value) may still be relatively high.
[0070] Therefore, it is necessary to continue to evaluate the performance of the neural network after the initial training. Specifically, the system inputs the training data into the network again, calculates the current value of the loss function, and compares it with the preset target loss value. If the current value of the loss function is still higher than the target value, it indicates that there is still room for further optimization of the network. Therefore, it is necessary to continue to perform the weight update iteration, that is, to calculate the weight error again, update the weights, and repeat this process. On the contrary, if the current value of the loss function has met or is lower than the preset target loss value, it indicates that the network has reached the desired performance level. At this time, the weight update iteration can be stopped, and it is considered that a trained neural network has been obtained. This network has learned an effective mapping relationship from the input data to the target output and can be used for subsequent inference or prediction tasks. This process is an iterative optimization process. By continuously adjusting the network weights according to the value of the loss function, the network performance can gradually converge to the optimal or approximately optimal state.
[0071] For the sake of understanding, reference is made to Figures 2 to 5 for illustration, but it is not intended to limit the neural network training method of the present application. Figure 2 FIG. is a schematic diagram of the summation of the output signal errors of the convolutional layer of the neural network training method of the present application and its matrix expansion. The output signal errors of the convolutional layer are distributed on n channels, and the spatial dimension of each channel is s 1 ×s 1 , that is, each channel is a matrix of s 1 ×s 1 . The errors at the same position on different channels are summed as an element of the summation result of the output errors of the convolutional layer . The summation result of the output errors of the convolutional layer is a single-channel matrix with a size of s 1 ×s 1 .
[0072] According to the spatial dimension (s l ) of the input data x 2 ×s 2 ) of the convolutional layer, it is selected whether to perform matrix expansion on the summation result of the output errors of the convolutional layer . There are two ways of matrix expansion. One is to perform matrix expansion by padding zeros around the summation matrix of the output errors of the convolutional layer , and the other is to perform complementary data expansion on the original matrix using part of the data of the summation matrix of the output errors of the convolutional layer so that the expanded matrix is consistent with the spatial dimension of the input data of the convolutional layer. The specific implementation method is as follows:
[0073] Expansion method 1: The expansion operation of the summation matrix of the output errors of the convolutional layer is to pad zeros around the summation matrix of the output errors of the convolutional layer Pad with zeros and sum the error matrices of the convolutional layer output to expand it into a matrix with the same spatial dimension as the input data x of the convolutional layer l Assume that the dimension of the error summation matrix of the convolutional layer output is s 1 ×s 1 , and the target dimension (i.e., the spatial dimension of the input data of the convolutional layer ) is s 2 ×s 2 , and s 2 >s 1 . The expanded matrix from position to position is equal to , and the remaining positions are 0. When is not an integer, round down. The specific zero-padding operation is as shown in Figure 3 ( Figure 3 is a schematic diagram of the first matrix expansion method of the neural network training method in this application).
[0074] Expansion method 2: The expansion operation of the error summation matrix of the convolutional layer output is to expand the matrix by padding numbers on the right and bottom sides of the error summation matrix of the convolutional layer output . The numbers to be filled are the contents of . The expanded matrix can be divided into 4 parts, as shown in Figure 4 ( Figure 4 is a schematic diagram of the second matrix expansion method of the neural network training method in this application). Part 1: from position (1,1) to position (s 1 ,s 1 ) is equal to ; Part 2: from position (1,s 1 +1) to position (s 1 ,s 2 ) is equal to the content in from position (1,2s 1 -s 2 +1) to position (s 1 ,s 1 ; Part 3: from position (s 1 +1,1) to position (s 2 ,s 1 ) is equal to the content in from position (2s 1 -s 2 +1,1) to position (s 1 ,s1 ) The content in the area is equal; Part 4: From position (s 1 +1, s 1 +1) to the area at (s 2 , s 2 ) is equal to the area from the position (2s -s 1 +1, 2s 2 -s 1 +1) to the position (s 2 , s 1 ) in 1 terms of content.
[0075] Perform a matrix-matrix element multiplication (i.e., matrix dot product operation) on the sum of the convolution output layer errors after expansion and each channel of the feedback network weights respectively to obtain the convolution layer input signal error. As Figure 5 shown ( Figure 5 is the schematic diagram for obtaining the convolution layer input signal error of the neural network training method in this application), the feedback weight network is a randomly generated three-dimensional matrix with the same spatial dimension (s 2 × s 2 ) and the number of channels m as the convolution layer input data, that is, m two-dimensional matrices with a spatial dimension of s 2 × s 2 . Each element in the sum matrix of the convolution output layer errors after expansion is multiplied by the corresponding element at a certain weight channel, and the obtained result is the convolution layer input signal error of this channel. Then, calculate the error of the feedforward network weights according to the convolution layer input signal error and the convolution layer output error, update the feedforward network weights according to the error of the feedforward network weights, and perform repeated training according to the number of training rounds until the neural network training ends. When using the neural network for inference, use the trained neural network weight information, perform forward signal propagation calculation on the neural network input data, and complete the inference process.
[0076] In this embodiment, the convolution layer output signal error is determined, the sum of the convolution layer output signal errors is matrix-expanded according to the convolution layer output signal error to obtain the sum of the convolution layer output signal errors after expansion, the sum of the convolution layer output signal errors after expansion is dot-multiplied with the weights of each channel of the feedback network to obtain the convolution layer input signal error, the feedforward network weight error is determined according to the convolution layer output signal error and the convolution layer input signal error, the feedforward network weights of the neural network are updated based on the feedforward network weight error to obtain the trained neural network, and through matrix expansion, matrix dot product operation, and weight update according to the feedforward network weight error, the accuracy of data propagation and calculation efficiency are improved, the storage and calculation requirements are reduced, and the resource overhead is reduced.
[0077] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be elaborated hereinafter. On this basis, please refer to Figure 6 , step S10 may include steps S101 to S103:
[0078] Step S101, obtaining the forward propagation result of the input signal, and determining the loss function of the input signal according to the forward propagation result.
[0079] It can be understood that the input signal is input into the neural network, and forward propagation is performed through the feed-forward path of the network (i.e., applying weights and activation functions layer by layer). This forward propagation process will generate an output signal, which is the prediction or classification result of the neural network for the current input signal.
[0080] Next, it is necessary to evaluate the quality of the output signal above, that is, the degree of difference between it and the expected target output. To quantify this difference, a loss function (also called a cost function or an error function) is introduced. The loss function is a mathematical expression that receives the output signal and the target output of the neural network as inputs and outputs a scalar value (i.e., the loss value), which reflects the degree of difference between the output signal and the target output. Specifically, the choice of the loss function depends on the specific problem and task type to be solved. For example, in a regression problem, the mean squared error (MSE) can be selected as the loss function; in a classification problem, the cross-entropy loss function can be selected. Once the loss function is determined, the loss value can be calculated according to the forward propagation result (i.e., the output signal of the neural network) and the target output. This loss value will be used to guide the training process of the neural network, that is, to update the weights of the network through the novel feedback alignment algorithm in the neural network training method of the present application to minimize the loss value and improve the performance of the network.
[0081] Step S102, determining the convolutional layer output signal error based on the loss function, and obtaining the sum of the convolutional layer output signal errors according to all channel signals of the convolutional layer output signal error.
[0082] It should be understood that based on the already determined loss function, the convolutional layer output signal error is used to evaluate the degree of mismatch between the convolutional layer output signal and the target value under the current weight configuration.
[0083] However, a convolutional layer usually contains multiple channels, and each channel outputs a feature map. Therefore, the convolutional layer output signal error also contains multiple channels, with the error of each channel corresponding to the error of a feature map. To integrate the errors of these channels, it is necessary to sum up the convolutional layer output signal errors of all channels to obtain the total convolutional layer output signal error, and this total error reflects the overall difference degree between the entire convolutional layer output signal and the target value. After obtaining the total convolutional layer output signal error, this error signal is used to guide the training process of the neural network.
[0084] Step S103: Based on the spatial dimension of the input signal, perform zero-padding or edge-padding on the total convolutional layer output signal error to obtain the total convolutional layer output signal error after padding.
[0085] It should be noted that during the training process of a neural network, especially a convolutional neural network, the spatial dimension of the total convolutional layer output signal error needs to match the input dimension of the feedback network to correctly perform feedback alignment and weight update. However, due to convolutional operations that may cause the spatial dimension of the feature map to decrease (for example, due to using a convolutional kernel with a stride greater than 1 or edge effects), the spatial dimension of the total convolutional layer output signal error may be smaller than the spatial dimension of the original input signal.
[0086] To solve this problem, based on the spatial dimension of the input signal, perform zero-padding (i.e., adding zero values around the error matrix) or edge-padding (i.e., adding edge values around the error matrix, and these values can be zero or other values selected according to specific strategies) on the total convolutional layer output signal error. Both of these padding methods can increase the spatial dimension of the error matrix to match the input dimension of the feedback network.
[0087] Among them, zero-padding is a simple and commonly used method. It adds one or more circles of zero values around the error matrix, thereby increasing the width and height of the matrix. This method does not introduce additional information but can increase the size of the matrix to make it suitable for feedback alignment. Edge-padding is a more flexible method. It allows adding non-zero values around the error matrix, and these values can be selected according to specific strategies. For example, values adjacent to the edge of the error matrix can be copied, or calculated according to a certain interpolation method. Edge-padding can retain more edge information and may help improve the effect of feedback alignment.
[0088] Through zero-padding or edge-padding, the total convolutional layer output signal error after padding is obtained. This padded error matrix has a spatial dimension that matches the input of the feedback network. In this way, this padded error matrix can be passed to the feedback network for subsequent feedback alignment and weight update steps.
[0089] In this embodiment, the above step S103 may include: obtaining the forward propagation result error based on the loss function, and obtaining the fully connected layer input signal error according to the forward propagation result error and the feedback weights of the fully connected layer of the feedback network; processing the fully connected layer input signal error based on the spatial dimension of the pooling layer output signal to obtain the pooling layer output signal error; performing an unpooling operation on the pooling layer output signal error to obtain the convolutional layer output activation signal error; determining the convolutional layer output signal error according to the convolutional layer output activation signal error and the derivative of the activation function, and obtaining the sum of the convolutional layer output signal errors according to the signals of all channels of the convolutional layer output signal error.
[0090] It should be understood that the process of calculating the error based on the loss function and aligning it back to each layer of the network to guide the weight update can be as follows: First, calculate the error of the forward propagation result according to the loss function, and use this forward propagation result error and the weights of the next fully connected layer to calculate the input signal error of the fully connected layer of the feedforward network through matrix multiplication (or dot product operation) and appropriate dimension adjustment.
[0091] However, since the network may include pooling layers, these layers change the spatial dimension of the signal. Therefore, before aligning the fully connected layer input signal error back to the pooling layer, it is necessary to process the fully connected layer input signal error according to the spatial dimension of the pooling layer output signal. Next, perform an unpooling operation on the processed pooling layer output signal error. Among them, unpooling is the inverse process of the pooling operation. It "restores" the error signal to the spatial dimension before pooling according to the index information recorded by the pooling layer (in max pooling) or by using other strategies (in average pooling). This process is a key step in the feedback alignment process, which ensures that the error signal can be correctly propagated to the convolutional layer.
[0092] Finally, use the convolutional layer output activation signal error obtained by the unpooling operation and the derivative of the convolutional layer activation function (which depends on the activation function used, such as ReLU, Sigmoid, etc.) to calculate the convolutional layer output signal error through the chain rule, and sum the errors of each channel of the convolutional layer to obtain the sum of the convolutional layer output signal errors.
[0093] In this embodiment, by obtaining the forward propagation result of the input signal, determining the loss function of the input signal according to the forward propagation result, determining the convolutional layer output signal error based on the loss function, obtaining the sum of the convolutional layer output signal errors according to the signals of all channels of the convolutional layer output signal error, and padding zeros or padding the edges to the sum of the convolutional layer output signal errors based on the spatial dimension of the input signal to obtain the sum of the expanded convolutional layer output signal errors, so as to match the input dimension of the convolutional layer, improve the accuracy of error propagation, and ensure the accuracy of error calculation.
[0094] As an implementation manner, in this embodiment, the above step S102 may include: performing a convolution operation on the input signal and the convolution layer weights of the feedforward network to obtain the output signal of the feedforward network convolution layer; obtaining the input signal of the feedforward network fully connected layer according to the output signal of the feedforward network convolution layer; performing a vector matrix multiplication operation on the input signal of the feedforward network fully connected layer and the fully connected layer weights of the feedforward network to obtain the output signal of the feedforward network fully connected layer; processing the output signal of the feedforward network fully connected layer through an activation function to obtain the forward propagation result of the input signal, and determining a loss function according to the forward propagation result and the target vector of the input signal.
[0095] It can be understood that during the forward propagation process of the neural network, especially in a network architecture involving a convolution layer and a fully connected layer, the input signal will pass through these layers in sequence to generate the final output. The input signal will perform a convolution operation with the convolution layer weights of the feedforward network. Among them, the convolution operation is a linear operation. It generates an output signal by sliding a convolution kernel (or called a filter) on the input signal and multiplying each element of the convolution kernel with the corresponding element of the input signal and then summing them up.
[0096] Next, the output signal of the feedforward network convolution layer obtained by the convolution operation will be further processed to generate the input signal of the feedforward network fully connected layer. This process may include pooling operations (such as max pooling or average pooling) to reduce the spatial dimension of the feature map and reduce the computational amount, as well as possible tiling or flattening operations to convert the multi-dimensional feature map into a one-dimensional vector for matrix multiplication with the weights of the fully connected layer. Then, the input signal of the feedforward network fully connected layer will perform a vector matrix multiplication operation with the fully connected layer weights of the feedforward network. Each neuron in the fully connected layer is connected to all neurons in the previous layer, so this operation is a dense connection pattern. The result of the operation is the output signal of the feedforward network fully connected layer, which is a one-dimensional vector, and the number of its elements depends on the number of neurons in the fully connected layer.
[0097] The output signal of the feedforward network fully connected layer will be processed through an activation function. The activation function is non-linear, which introduces non-linear factors and enables the network to learn and represent complex patterns. The signal processed by the activation function constitutes the forward propagation result of the input signal. Finally, according to the forward propagation result and the target vector of the input signal (i.e., the expected output), the loss function can be determined.
[0098] Further, the step of obtaining the input signal of the fully connected layer of the feedforward network according to the output signal of the feedforward network convolutional layer in the above steps includes: processing the output signal of the feedforward network convolutional layer through an activation function to obtain the output activation signal of the feedforward network convolutional layer; performing a pooling operation on the output activation signal of the feedforward network convolutional layer to obtain the output signal of the pooling layer; and flattening the input signal of the next layer based on the output signal of the pooling layer to obtain the input signal of the fully connected layer in the feedforward network.
[0099] It should be noted that in the neural network training method of the present application, the entire network structure has a high degree of flexibility. On the one hand, the network can include more than two convolutional layers, and these convolutional layers can be connected in sequence to extract different features in the input signal. The output signal of each convolutional layer can be used as the input signal of the next convolutional layer, thereby constructing a deep convolutional neural network structure.
[0100] On the other hand, the pooling layer is not necessarily directly connected or adjacently connected after the convolutional layer. Although the pooling layer is usually used to reduce the spatial dimension of the data, reduce the amount of calculation, and prevent overfitting, in the method of the present application, whether to add a pooling layer and the number and position of the pooling layers can be flexibly adjusted according to specific network designs and task requirements. For example, in some cases, in order to retain more detailed information, it can be selected not to add a pooling layer after some convolutional layers; while in other cases, in order to reduce the dimension and complexity of the data, a pooling layer can be added after multiple convolutional layers. Therefore, the neural network training method of the present application is applicable not only to networks with fixed numbers of layers and structures, but also to diverse networks with different numbers of layers and structures, thereby providing a wider range of application scenarios and stronger adaptability.
[0101] Exemplarily, the step of obtaining the input signal of the fully connected layer of the feedforward network based on the output signal of the convolutional layer of the feedforward network can be described as follows: The output signal of the convolutional layer of the feedforward network is processed through an activation function and transformed into an activation signal, that is, the convolutional layer of the feedforward network outputs an activation signal. The activation signal output by the convolutional layer of the feedforward network may include the activation signal output by the first convolutional layer of the feedforward network and the activation signal output by the second convolutional layer of the feedforward network. Then, a pooling operation is performed on it to obtain the output signal of the pooling layer, so as to reduce the computational amount and avoid overfitting. The output signal of the pooling layer will be used as the input signal of the next convolutional layer. However, before entering the fully connected layer, a flattening operation is required. The flattening operation is a process of converting a multi-dimensional feature map into a one-dimensional signal, so that each element can serve as an input neuron of the fully connected layer. The flattened signal retains all the information in the feature map, but converts its form into a one-dimensional vector suitable for processing by the fully connected layer. Finally, the flattened one-dimensional signal is used as the input signal of the fully connected layer. Each neuron in the fully connected layer is connected to all neurons in the previous layer. Through vector matrix multiplication operations and possible activation function processing, the final output signal is generated. This output signal can be the class probability of a classification task, the predicted value of a regression task, etc., depending on the task and architecture of the network.
[0102] If no pooling layer is connected after the convolutional layer, the activation signal output by the convolutional layer of the feedforward network is flattened, and then the input signal of the fully connected layer in the feedforward network is obtained.
[0103] For the sake of easy understanding, the following takes the task of a convolutional neural network for digit recognition of a handwritten dataset as an example, referring to Figure 7 , to describe the specific steps when using the neural network training method of the present application to train and infer a convolutional neural network, but it does not limit the neural network training method of the present application. Figure 7 FIG. is a schematic diagram of digit recognition for the neural network training method of the present application, and the specific steps are as follows:
[0104] Step 1.1: Convolve the input signal x 1 with the feedforward network weight w 1 in convolutional layer 1 to obtain the output signal y 1 = conv(x 1 , w 1 , b 1 );
[0105] Step 1.2: Use the activation function f(·) to operate on the output signal y 1 of convolutional layer 1 to obtain the output activation signal z 1 = f(y 1 );
[0106] Step 1.3: For the output activation signal z of convolutional layer 1 1 perform a pooling operation to obtain the output signal p of pooling layer 2 1 = maxpool(z 1 );
[0107] Step 1.4: Use the output signal p of pooling layer 2 1 as the input signal x of convolutional layer 3 2 , that is, x 2 = p 1 . Convolve the input signal x of convolutional layer 3 2 with the weights w in convolutional layer 3 2 to obtain the output signal y of convolutional layer 3 2 = conv(x 2 , w 2 , b 1 , method);
[0108] There are two methods for the above convolution calculation, namely:
[0109] 1) same method: Pad zeros to the input signal x of convolutional layer 3 2 . The spatial dimension size of the output signal y of convolutional layer 3 2 is the same as that of the input signal x of convolutional layer 3 2 ;
[0110] 2) valid method: Do not pad zeros to the input signal x of convolutional layer 3 2 . The spatial dimension size of the output signal y of convolutional layer 3 2 is not the same as that of the input signal x of convolutional layer 3 2 . In the subsequent step S407, a matrix expansion operation needs to be performed on the sum signal of the loss function L with respect to the output error of convolutional layer 3 to complete the calculation of the error 2 of the input signal x of convolutional layer 3 with respect to the loss function L .
[0111] Step 1.5: Use the activation function f(·) to operate on the output signal y of convolutional layer 3 2 to obtain the output activation signal z of convolutional layer 3 2 = f(y 2 );
[0112] Step 1.6: For the output activation signal z of convolutional layer 3 2 perform a pooling operation to obtain the output signal p of pooling layer 4 2 = maxpool(z 2 );
[0113] Step 1.7: Use the output signal p of pooling layer 4 2 as the input signal x of the next layer 3 , that is, x 3 = p 2 , and perform a flattening operation on x 3 to obtain a one-dimensional signal flatten(x 3 ), which is used as the input signal of fully connected layer 5;
[0114] Step 1.8: Perform a vector-matrix multiplication on the input signal flatten(x 3 ) of fully connected layer 5 and the weight w 3 in fully connected layer 5 to obtain the output signal y 3 = flatten(x 3 )·w 3 ;
[0115] Step 1.9: Use the activation function f(·) to operate on the output signal y 3 of fully connected layer 5 to obtain the output activation signal z 3 = f(y 3 );
[0116] Step 2: The output activation signal z 3 of fully connected layer 5 is the output signal of the feedforward neural network. Use the output signal z 3 of the feedforward neural network and the target vector t to calculate the loss function L;
[0117] Step 3.1: Calculate the error signal 3 of the loss function L with respect to the output activation signal z
[0118] Step 3.2: Through (the error signal of the loss function L with respect to the output activation signal z 3 ) and the weight w b 3 in the feedback network's fully connected layer 6 to perform a vector-matrix multiplication operation to obtain a proxy value of the error of the loss function L with respect to the input signal x 3 of fully connected layer 5
[0119] Step 3.3: According to the dimension of the output signal p 2 of pooling layer 4 (the proxy value of the partial derivative of the loss function L with respect to the input signal x 3 ) to rearrange the elements to obtain a proxy value of the error of the loss function L with respect to the output signal p 2 of pooling layer 4
[0120] Step 3.4: In the unpooling layer operation, for (the proxy value of the error of the loss function L with respect to the output signal p of pooling layer 4 2 ), perform unpooling operation (maxpool) to obtain the proxy value of the partial derivative of the loss function L with respect to the output activation signal z of convolutional layer 3 2 .
[0121] Step 3.5: In the deactivation layer operation, according to the proxy value of the partial derivative of the loss function L with respect to the output activation signal z of convolutional layer 3 2 and the derivative f'(·) of the activation function, calculate the proxy value of the partial derivative of the loss function L with respect to the output signal y of convolutional layer 3 . 2
[0122] Step 3.6: Sum the error signals at the same positions on different channels of (the proxy value of the partial derivative of the loss function L with respect to the output signal y of convolutional layer 3 2 ) as an element of the sum result of the output error of convolutional layer 3, to obtain the sum signal of the output error of the loss function L with respect to convolutional layer 3
[0123] Step 3.7: According to the spatial dimension of the input signal x of convolutional layer 3 2 and the convolution calculation method (same or valid) in step S204, perform an expansion operation on the sum signal of the output error of the loss function L with respect to convolutional layer 3 . When the same method is used for convolution in step S204, has the same spatial dimension as the input signal x of convolutional layer 3 2 . When performing matrix-matrix element multiplication between and ( has the same dimension as x 2 ), no expansion operation is required. When the valid method is used for convolution in step S204, perform an expansion operation on to obtain whose dimension is equal to the spatial dimension of the input signal x of convolutional layer 3 2 2 .
[0124] Step 3.8: Use the matrix after the expansion operation of the sum signal of the output error of the loss function L with respect to convolutional layer 3 and the feedback weight (which is the same as x 2Perform a matrix-matrix element product (denoted using the " " symbol) on matrices of the same dimension to obtain a proxy value for the error of the loss function L with respect to the input signal x to convolutional layer 3 2
[0125] Step 3.9: The error of the loss function L with respect to the input signal x to convolutional layer 3 2 is the same as the error of the loss function L with respect to the output signal p of pooling layer 2 1 That is Perform an unpooling operation (maxunpool) on the error of the loss function L with respect to the output signal p of pooling layer 2 1 to obtain a proxy value for the error of the loss function L with respect to the output activation signal z of convolutional layer 1 1
[0126] Step 3.10: Calculate a proxy value for the error of the loss function L with respect to the output signal y of convolutional layer 1 based on the proxy value of the partial derivative of the loss function L with respect to the output activation signal z of convolutional layer 1 1 and the derivative f'(·) of the activation function 1
[0127] Step 4.1: Perform a convolution operation on the input signal x to convolutional layer 3 2 and the error of the loss function L with respect to the output signal of convolutional layer 3 to obtain a proxy value for the error of the loss function L with respect to the weights w in convolutional layer 3 2
[0128] Step 4.2: Perform a convolution operation on the input signal x 1 and the error of the loss function L with respect to the output signal of convolutional layer 1 to obtain a proxy value for the error of the loss function L with respect to the weights w in convolutional layer 1 1
[0129] Step 5: Update the weights w according to the proxy value for the error of the loss function L with respect to the weights w in convolutional layer 3 2 and the weight update algorithm; update the weights w according to the proxy value for the error of the loss function L with respect to the weights w in convolutional layer 1 and the weight update algorithm; 2 1 1
[0130] Step 6: Repeat the above steps 1 to 5 to gradually reduce the loss function until the training of the convolutional neural network is completed.
[0131] The above is the training process of the neural network. When performing inference, only steps 1.1 to 1.9 need to be carried out, and the classification judgment is made according to the output activation signal z of the fully connected layer 5 of the neural network 3 to complete the inference process.
[0132] It should be noted that applying the neural network training method of the present application can implement a memristor hardware. In combination with Figure 8 and Figure 9 for explanation, this embodiment places restrictions on this. Figure 8 This is a schematic diagram of the hardware implementation of the convolutional computing memristor of the present application, Figure 9 and this is a schematic diagram of the hardware implementation of the matrix-matrix element product operator memristor of the present application.
[0133] The specific steps of the design scheme for implementing convolutional computing using memristor hardware are as follows:
[0134] 1. Use the mapping method to convert the input information into a voltage value signal;
[0135] 2. Use the mapping method to convert the convolutional kernel weight information into a memristor conductance value signal;
[0136] 3. Input the voltage signal into the memristor array, and based on Kirchhoff's law and Ohm's law, perform vector matrix multiplication and addition operations. The calculated current value represents the calculation result;
[0137] 4. Use a transconductance amplifier to convert the output current signal into a voltage signal as the input voltage signal for the next layer of calculation.
[0138] During the calculation process, multiple memristor arrays with the same resistance value can be used to represent the same convolutional weight, thereby realizing parallel calculation of signals; or through the method of time multiplexing, the voltage signal can be input into the same memristor array for sequential calculation. How many identical memristor arrays are specifically used to represent the convolutional kernel for parallel acceleration calculation can be determined by comprehensively considering two factors: hardware overhead and calculation time.
[0139] The specific steps of the design scheme for the hardware implementation of the matrix-matrix element product operator using memristors are as follows:
[0140] 1. Use the mapping method to convert the input value signal into a voltage value signal;
[0141] 2. Use the mapping method to convert the weight information into a memristor conductance value signal;
[0142] 3. Input the voltage value into the memristor array. Based on Ohm's law, calculate the matrix-matrix element product operation, and the calculated current value represents the calculation result.
[0143] 4. Use a transconductance amplifier to convert the output current signal into a voltage signal, which is used as the input voltage signal for the next layer of calculation.
[0144] In this embodiment, the input signal is convolved with the weights of the convolutional layer of the feedforward network to obtain the output signal of the convolutional layer of the feedforward network. The input signal of the fully connected layer of the feedforward network is obtained according to the output signal of the convolutional layer of the feedforward network. The input signal of the fully connected layer of the feedforward network is subjected to a vector-matrix multiplication operation with the weights of the fully connected layer of the feedforward network to obtain the output signal of the fully connected layer of the feedforward network. The output signal of the fully connected layer of the feedforward network is processed through an activation function to obtain the forward propagation result of the input signal, and the loss function is determined according to the forward propagation result and the target vector of the input signal, so that the neural network can learn the deep features of the input signal, enhance the expression ability of the neural network, and ensure the quality and stability of the input signal.
[0145] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the neural network training method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0146] This application also provides a neural network training device. Please refer to Figure 10 , the neural network training device includes:
[0147] The output signal error determination module 10 is used to determine the output signal error of the convolutional layer and perform matrix expansion on the sum of the output signal errors of the convolutional layer according to the output signal error of the convolutional layer to obtain the sum of the expanded output signal errors of the convolutional layer.
[0148] The input signal error determination module 20 is used to obtain the input signal error of the convolutional layer by performing a matrix dot product operation on the sum of the expanded output signal errors of the convolutional layer and the weights of each channel of the feedback network.
[0149] The weight error determination module 30 is used to determine the weight error of the feedforward network according to the output signal error of the convolutional layer and the input signal error of the convolutional layer.
[0150] The network training module 40 is used to update the weights of the feedforward network of the neural network based on the weight error of the feedforward network to obtain the trained neural network.
[0151] The neural network training device provided by this application adopts the neural network training method in the above embodiment, and can solve the technical problem that the existing neural network training algorithm has large storage requirements and complex computing requirements, resulting in large resource overhead. Compared with the prior art, the beneficial effects of the neural network training device provided by this application are the same as those of the neural network training method provided by the above embodiment, and other technical features in the neural network training device are the same as those disclosed in the method of the above embodiment, which will not be elaborated here.
[0152] This application provides a neural network training device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the neural network training method in the first embodiment above.
[0153] Refer to the following Figure 11 , which shows a schematic structural diagram of a neural network training device suitable for implementing the embodiments of this application. The neural network training device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Desctions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 11 The neural network training device shown is only an example, and should not impose any limitation on the functions and usage scope of the embodiments of this application.
[0154] As Figure 11As shown, the neural network training device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in the read-only memory 1002 or a program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the neural network training device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output interface 1006 is also connected to the bus. Generally, the following systems may be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the neural network training device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a neural network training device having various systems, it should be understood that it is not required to implement or have all the shown systems. Instead, more or fewer systems may be implemented or had.
[0155] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0156] The neural network training device provided by the present application adopts the neural network training method in the above embodiments, and can solve the technical problem that the existing neural network training algorithms have large storage requirements and complex computing requirements, resulting in large resource overheads. Compared with the prior art, the beneficial effects of the neural network training device provided by the present application are the same as those of the neural network training method provided by the above embodiments, and other technical features in the neural network training device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0157] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0158] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0159] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the neural network training method in the above embodiments.
[0160] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0161] The above computer-readable storage medium can be included in the neural network training device; it can also exist separately without being assembled into the neural network training device.
[0162] The above computer-readable storage medium carries one or more programs, which, when executed by a neural network training device, cause the neural network training device to: determine the error of the output signal of the convolutional layer, perform matrix expansion on the sum of the output signal errors of the convolutional layer according to the error of the output signal of the convolutional layer, obtain the expanded sum of the output signal errors of the convolutional layer, obtain the error of the input signal of the convolutional layer by performing matrix dot product operation on the expanded sum of the output signal errors of the convolutional layer and the weights of each channel of the feedback network, determine the error of the feedforward network weights according to the error of the output signal of the convolutional layer and the error of the input signal of the convolutional layer, update the feedforward network weights of the neural network based on the error of the feedforward network weights, and obtain the trained neural network.
[0163] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by connecting through the Internet service provider via the Internet).
[0164] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above neural network training method, which can solve the technical problem that the existing neural network training algorithm has large storage requirements and complex computing requirements, resulting in large resource overhead. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the neural network training method provided by the above embodiment, and will not be elaborated here.
[0165] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the neural network training method as described above.
[0166] The computer program product provided by the present application can solve the technical problem that the existing neural network training algorithms have large storage requirements and complex computing requirements, resulting in large resource overheads. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the neural network training method provided in the above embodiments, and will not be elaborated herein.
[0167] The above are only some embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or any direct / indirect application in other related technical fields shall be included in the patent protection scope of the present application.
Claims
1. A neural network training method, characterized in that: The method comprises the following steps: Determine the convolution layer output signal error, and perform matrix expansion on the convolution layer output signal error sum according to the convolution layer output signal error to obtain the expanded convolution layer output signal error sum; The convolution layer input signal error is obtained by performing a matrix dot multiplication operation on the sum of the output signal errors of the expanded convolution layer and the weights of each channel of the feedback network; Determining a feedforward network weight error according to the convolutional layer output signal error and the convolutional layer input signal error; The feedforward network weights of the neural network are updated based on the feedforward network weight errors to obtain a trained neural network.
2. The neural network training method according to claim 1, characterized in that: The step of determining the convolution layer output signal error, and performing matrix expansion on the convolution layer output signal error sum according to the convolution layer output signal error to obtain the expanded convolution layer output signal error sum includes: Obtaining a forward propagation result of an input signal, and determining a loss function of the input signal according to the forward propagation result; Determine the convolution layer output signal error based on the loss function, and obtain the sum of the convolution layer output signal errors according to all channel signals of the convolution layer output signal error; Based on the spatial dimension of the input signal, the sum of the errors of the convolutional layer output signals is zero-padded or edge-expanded to obtain the sum of the errors of the convolutional layer output signals after expansion.
3. The neural network training method according to claim 2, characterized in that: The step of determining the convolution layer output signal error based on the loss function and obtaining the sum of the convolution layer output signal errors according to all channel signals of the convolution layer output signal error comprises: Obtaining a forward propagation result error based on the loss function, and obtaining a fully connected layer input signal error according to the forward propagation result error and a fully connected layer feedback weight of a feedback network; Processing the fully connected layer input signal error based on the spatial dimension of the pooling layer output signal to obtain the pooling layer output signal error; Performing a de-pooling operation on the output signal error of the pooling layer to obtain an output activation signal error of the convolutional layer; According to the convolution layer output activation signal error and the activation function derivative, the convolution layer output signal error is determined, and the sum of the convolution layer output signal errors is obtained according to all channel signals of the convolution layer output signal error.
4. The neural network training method according to claim 2, characterized in that: The step of obtaining a forward propagation result of the input signal and determining a loss function of the input signal according to the forward propagation result comprises: Perform convolution operation on the input signal and the convolution layer weight of the feedforward network to obtain the output signal of the convolution layer of the feedforward network; Obtaining a feedforward network fully connected layer input signal according to the feedforward network convolutional layer output signal; Performing a vector-matrix multiplication operation on the fully connected layer input signal of the feedforward network and the fully connected layer weight of the feedforward network to obtain the fully connected layer output signal of the feedforward network; The output signal of the fully connected layer of the feedforward network is processed by an activation function to obtain a forward propagation result of the input signal, and a loss function is determined based on the forward propagation result and a target vector of the input signal.
5. The neural network training method according to claim 4, characterized in that: The step of obtaining the fully connected layer input signal in the feedforward network according to the feedforward network convolutional layer output signal comprises: Processing the output signal of the feedforward network convolution layer through an activation function to obtain an output activation signal of the feedforward network convolution layer; Performing a pooling operation on the output activation signal of the convolutional layer of the feedforward network to obtain a pooling layer output signal; Based on the output signal of the pooling layer, a flattening operation is performed on the input signal of the next layer to obtain the input signal of the fully connected layer in the feedforward network.
6. The neural network training method according to any one of claims 1 to 5, characterized in that: The step of updating the feedforward network weights of the neural network based on the feedforward network weight error to obtain the trained neural network comprises: Based on the feedforward network weight error, the feedforward network weight of the neural network is updated by a weight update algorithm to obtain an initially trained neural network; Determining a loss function, and if the loss function does not satisfy a preset target loss value, continuing to update the feedforward network weights of the neural network after the initial training; If the loss function satisfies the preset target loss value, the updating of the feedforward network weights of the neural network after initial training is stopped, and the trained neural network is obtained.
7. A neural network training device, characterized in that: The neural network training device comprises: An output signal error determination module is used to determine the convolution layer output signal error, and perform matrix expansion on the convolution layer output signal error sum according to the convolution layer output signal error to obtain the expanded convolution layer output signal error sum; An input signal error determination module, used for obtaining a convolution layer input signal error by performing a matrix dot multiplication operation on the sum of the expanded convolution layer output signal errors and the weights of each channel of the feedback network; A weight error determination module, used to determine a feedforward network weight error according to the convolution layer output signal error and the convolution layer input signal error; The network training module is used to update the feedforward network weights of the neural network based on the feedforward network weight errors to obtain a trained neural network.
8. A neural network training device, characterized in that: The neural network training device comprises: a memory, a processor, and a neural network training program stored in the memory and executable on the processor. When the neural network training program is executed by the processor, the neural network training method according to any one of claims 1 to 6 is implemented.
9. A storage medium, characterized in that: The storage medium stores a neural network training program, and when the neural network training program is executed by the processor, the neural network training method according to any one of claims 1 to 6 is implemented.
10. A computer program product, characterized in that The computer program product comprises a neural network training program, and when the neural network training program is executed by a processor, the neural network training method according to any one of claims 1 to 6 is implemented.