Operation method and training method and device of neural network model
By transforming the input and weight matrices using the Winograd algorithm and replacing dot multiplication with addition, the high computational complexity of neural network models is solved, resulting in faster computation speed and lower power consumption.
Patent Information
- Application Number
- CN202180094093.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-30
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2041-04-30
AI Technical Summary
The high complexity of matrix operations in neural network models leads to longer computation time, affecting processing efficiency and increasing power consumption.
The Winograd algorithm is used to transform the input data and weight matrix, replacing the dot product operation with an addition operation to calculate the L1 distance, thus reducing the computational load of the feature extraction process.
It improves the running speed of neural network models, reduces computational overhead, and decreases power consumption.
Smart Images

Figure CN116888605B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and more specifically, to computational methods, training methods, and apparatus for neural network models. Background Technology
[0002] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making capabilities. Research in the field of AI includes robotics, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, and fundamental AI theories.
[0003] Neural network models typically involve a large number of matrix operations. Taking convolution as an example, convolution involves multiplication, which has high computational complexity and long latency. Matrix operations usually consume a large portion of the computation time of a neural network model, and the latency of matrix operations becomes a major factor limiting computational efficiency, affecting the overall processing efficiency of the neural network model and causing significant power consumption losses.
[0004] Therefore, how to reduce the computational overhead of neural network models has become an urgent problem to be solved. Summary of the Invention
[0005] This application provides a computational method, training method, and apparatus for a neural network model, which can reduce the computational overhead of the neural network model and improve processing efficiency.
[0006] Firstly, a computational method for a neural network model is provided. This method includes performing the following operations in at least one feature extraction layer of the neural network model: transforming the input data matrix of the data to be processed using the input transformation matrix of the Winograd algorithm to obtain a transformed input data matrix, where the data to be processed includes image data, speech data, or text data; extracting features from the transformed input data matrix using the transformed weight matrix to obtain an intermediate matrix, wherein the transformed weight matrix is obtained by transforming the weight matrix of at least one feature extraction layer using the weight transformation matrix of the Winograd algorithm, and each element in the intermediate matrix is determined based on the L1 distance between the corresponding elements in the transformed input data matrix and the transformed weight matrix; and transforming the intermediate matrix using the output transformation matrix of the Winograd algorithm to obtain an output data matrix.
[0007] According to the scheme of this application embodiment, the dot product operation in winograd is replaced with addition operations such as calculating L1 distance, which reduces the amount of computation in the feature extraction process, improves the running speed of the model, and reduces the computational overhead.
[0008] The data to be processed includes image data, voice data, or text data.
[0009] The type of data to be processed is related to the task of the neural network model. For example, if the neural network model is used for image processing tasks, the data to be processed can be images. Specifically, image processing tasks include image classification, image detection, image segmentation, image recognition, or image generation. As another example, if the neural network model is used for text processing tasks, the data to be processed can be text. Specifically, text processing tasks include text recognition or text translation. Similarly, if the neural network model is used for speech processing tasks, the data to be processed can be speech data. Specifically, speech processing tasks include speech recognition. The embodiments of this application do not limit the type of data to be processed.
[0010] The input data matrix of the data to be processed refers to the data matrix input to at least one feature extraction layer.
[0011] The input data matrix can be part or all of the input feature map. The input feature map can be the data itself input into the neural network model, such as the image to be processed. The input feature map can also be the feature map obtained after one or more feature extraction layers in the neural network model.
[0012] Specifically, the input data matrix is transformed by the input transformation matrix to obtain the transformed input data matrix. The weight matrix is the parameter of one or more feature extraction kernels in the at least one feature extraction layer in the neural network model.
[0013] Specifically, the transformed weight matrix can be obtained by performing a winograd transformation on the weight matrix using a weight transformation matrix.
[0014] Specifically, the intermediate matrix is transformed by the output transformation matrix using a winograd transformation to obtain the output data matrix.
[0015] The output data matrix can be part or all of the output feature map.
[0016] In conjunction with the first aspect, in some implementations of the first aspect, the transformed weight matrix is obtained through offline transformation.
[0017] The transformed weight matrix is obtained before the neural network model is executed. For example, the transformed weight matrix may be obtained before the neural network model is deployed.
[0018] According to the scheme of the embodiments of this application, the weight matrix remains unchanged during the inference process of the neural network model. Obtaining the transformed weight matrix through offline calculation can further improve the calculation speed and reduce the calculation overhead.
[0019] In conjunction with the first aspect, in some implementations of the first aspect, each element in the intermediate matrix is the negative of the L1 distance between the corresponding element in the transformed input data matrix and the transformed weight matrix.
[0020] In conjunction with the first aspect, in some implementations of the first aspect, the output data matrix satisfies the following formula:
[0021] Y = A T [-|[GgG T ]-[B T dB]|]A;
[0022] Where Y represents the output data matrix, A represents the output transformation matrix, and A T Let G denote the transpose of A, and let G denote the weight transformation matrix. T Let G be the transpose of the matrix, g be the weight matrix before the transformation, and B be the input transformation matrix. T Let d denote the transpose of B, and let d denote the input data matrix before the transformation.
[0023] In conjunction with the first aspect, in some implementations of the first aspect, the value of an element in the output transformation matrix is any one of 0, -1, or 1.
[0024] According to the scheme of the embodiment of this application, the elements in the output transformation matrix take the values of 0, 1 or -1, which can reduce multiplication operations, further reduce the number of calculations, and help reduce the computational load of the model.
[0025] In conjunction with the first aspect, in some implementations of the first aspect, the output transformation matrix is:
[0026]
[0027] Where c0, c1, and c2 are any of the terms 0, -1, and 1, respectively.
[0028] In conjunction with the first aspect, in some implementations of the first aspect, at least one row of the output transformation matrix contains elements that are the negatives of the elements at the corresponding positions in the first matrix. The elements of the other rows of the output transformation matrix are the same as the elements at the corresponding positions in the other rows of the first matrix. The first matrix is:
[0029]
[0030] Where c0, c1, and c2 are any of the terms 0, -1, and 1, respectively.
[0031] In other words, any row of A' can be inverted without affecting the final result of the winograd transformation.
[0032] For example, c0 is 0, c1 is -1, and c2 is 1. Or, c0 is -1, c1 is 1, and c2 is 0.
[0033] In conjunction with the first aspect, in some implementations of the first aspect, the weight transformation matrix is:
[0034]
[0035] According to the scheme of the embodiments of this application, the above-mentioned output transformation matrix and weight transformation matrix satisfy the general solution form of the Winograd algorithm, which can ensure that the convolution calculation result obtained by the Winograd algorithm is the same as the result of conventional convolution calculation. In this way, when using the existing input transformation matrix, the above-mentioned output transformation matrix and weight transformation matrix can still be applied to convolution calculation.
[0036] In conjunction with the first aspect, in some implementations of the first aspect, the number of positive numbers in each column of the output transformation matrix is the same, and the number of negative numbers in each column is the same.
[0037] According to the scheme of this application embodiment, the number of positive numbers in each column of the output transformation matrix is the same, and the number of negative numbers in each column is the same. For example, the number of +1s in each column of the output transformation matrix is the same, and the number of -1s in each column is the same. This can balance the magnitude of each position in the output data matrix, that is, reduce the imbalance of eigenvalue accumulation, which is beneficial to model training. In addition, it is beneficial to perform subsequent processing on the output data matrix, such as batch normalization.
[0038] Secondly, a method for training a neural network model is provided. This method includes: transforming the input data matrix of the training data using the input transformation matrix of the Winograd algorithm to obtain a transformed input data matrix; extracting features from the transformed input data matrix using the transformed weight matrix to obtain an intermediate matrix, wherein the transformed weight matrix is obtained by transforming the weight matrix of at least one feature extraction layer using the weight transformation matrix of the Winograd algorithm, and each element in the intermediate matrix is determined based on the Lp distance between corresponding elements in the transformed input data matrix and the transformed weight matrix; transforming the intermediate matrix using the output transformation matrix of the Winograd algorithm to obtain an output data matrix; determining the value of the loss function based on the output data matrix; and training the neural network model based on the value of the loss function. In the m-th iteration of training the neural network model, p is 2, and in the n-th iteration, p is 1, where m and n are positive integers, and m is less than n.
[0039] According to the scheme of this application embodiment, in the early stage of training, L2 distance is used to assist training. L2 distance is more favorable to the Winograd algorithm, which can improve the convergence speed of the training process and thus improve the training effect of the model. In the later stage of training, training is based on L1 distance to further improve the training effect of the model using L1 distance. Using L1 distance in the trained model is more hardware-friendly.
[0040] The type of training data depends on the task of the neural network model. For example, if the neural network model is used for image processing tasks, the training data can be images. Specifically, image processing tasks include image classification, image detection, image segmentation, or image generation. Similarly, if the neural network model is used for text processing tasks, the training data can be text. Specifically, text processing tasks include text recognition or text translation. Furthermore, if the neural network model is used for speech processing tasks, the training data can be speech data. Specifically, speech processing tasks include speech recognition. This application does not limit the type of training data.
[0041] Specifically, the output data matrix is part or all of the output feature map. The processing result of the training data can be determined based on the output feature map, and the value of the loss function can be calculated based on the processing result of the training data.
[0042] The results of training data processing depend on the type of training data and the task of the neural network model.
[0043] For example, the training data is image data, and image processing may include image super-resolution processing, image denoising processing, image recognition processing, etc. Correspondingly, the image processing results may include image super-resolution, image denoising, or image classification, etc. This application embodiment does not limit this.
[0044] For example, the training data is speech data, and speech processing may include speech recognition, etc. Accordingly, the speech processing results include speech recognition results, etc. This application embodiment does not limit this.
[0045] In conjunction with the second aspect, in some implementations of the second aspect, during the training of the neural network model, the initial value of p is 2, and the value of p decreases as the number of iterations increases.
[0046] In other words, p is reduced from 2 to 1 during training.
[0047] In conjunction with the second aspect, in some implementations of the second aspect, the neural network model is trained based on the value of the loss function, including: adjusting the weights in the first weight matrix according to the partial derivative of the loss function with respect to the weights in the first weight matrix, wherein the first weight matrix includes the weight matrix before transformation or the weight matrix after transformation.
[0048] During training, the gradient of the weight matrix before transformation can be calculated, and the values of the weight matrix before transformation can be adjusted accordingly. Alternatively, the gradient of the weight matrix after transformation can be calculated, and the values of the weight matrix after transformation can be adjusted accordingly.
[0049] In conjunction with the second aspect, in some implementations of the second aspect, the partial derivatives of the loss function with respect to the weights in the first weight matrix satisfy the following formula:
[0050]
[0051] Where p is the norm of the calculation, p∈[1,2], w represents the weight, x represents the data in the input data matrix before or after the transformation, L represents the loss function, i represents the number of layers in the neural network model, and sign() represents the sign function.
[0052] When p = 2, the above equation represents the partial derivative of the loss function obtained based on the L2 distance with respect to the first weight matrix, meaning that backpropagation is performed based on the L2 distance. When p = 1, the above equation represents the partial derivative of the loss function obtained based on the L1 distance with respect to the first weight matrix, meaning that backpropagation is performed based on the L2 distance.
[0053] Thirdly, a computing device for a neural network model is provided, the device comprising a module or unit for executing the methods described in the first aspect and any implementation thereof.
[0054] Fourthly, a training apparatus for a neural network model is provided, the apparatus comprising a module or unit for performing the methods described in the second aspect and any implementation thereof.
[0055] It should be understood that the extensions, limitations, interpretations and descriptions of the relevant content in the first aspect above also apply to the same content in the second, third and fourth aspects.
[0056] Fifthly, a computing device for a neural network model is provided, the device comprising: a memory for storing a program; and a processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to execute the method of the first aspect and any implementation thereof.
[0057] The processor mentioned in the fifth aspect above can be a central processing unit (CPU) or a combination of a CPU and a neural network processing processor. This neural network processing processor can include graphics processing units (GPUs), neural network processing units (NPUs), and tensor processing units (TPUs), etc. The TPU is a dedicated integrated circuit developed by Google for a fully customized artificial intelligence accelerator for machine learning.
[0058] In a sixth aspect, a training apparatus for a neural network model is provided, the apparatus comprising: a memory for storing a program; and a processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to execute the method of the second aspect and any implementation thereof.
[0059] The processor mentioned in the sixth aspect above can be a central processing unit (CPU) or a combination of a CPU and a neural network processing processor. The neural network processing processor can include a graphics processing unit (GPU), a neural network processor (NNP), and a tensor processor, among others. The TPU is a Google-designed application-specific integrated circuit (ASIC) for a fully customized AI accelerator designed for machine learning.
[0060] A seventh aspect provides a computer-readable storage medium storing program code for execution by a device, the program code including methods for performing any implementation of the first or second aspect.
[0061] Eighthly, a computer program product containing instructions is provided, which, when run on a computer, causes the computer to perform the method in any one of the implementations of the first or second aspect described above.
[0062] Ninth aspect, a chip is provided, the chip including a processor and a data interface, the processor reading instructions stored in a memory through the data interface and executing the method in any one of the implementations of the first or second aspect described above.
[0063] Optionally, as one implementation, the chip may further include a memory storing instructions, and the processor is used to execute the instructions stored in the memory. When the instructions are executed, the processor is used to execute the method in either the first aspect or the second aspect.
[0064] The aforementioned chip can be a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). Attached Figure Description
[0065] Figure 1 This is a schematic diagram of an artificial intelligence main framework provided in an embodiment of this application;
[0066] Figure 2 A schematic diagram of a system architecture provided in an embodiment of this application;
[0067] Figure 3 This is a schematic diagram of the structure of a convolutional neural network provided in an embodiment of this application;
[0068] Figure 4 A schematic diagram of the hardware structure of a chip provided in an embodiment of this application;
[0069] Figure 5A schematic diagram of a system architecture provided for an embodiment of this application;
[0070] Figure 6 A schematic block diagram of a computing device for a neural network model provided in an embodiment of this application;
[0071] Figure 7 A schematic flowchart illustrating a computation method for a neural network model provided in an embodiment of this application;
[0072] Figure 8 A schematic flowchart illustrating a training method for a neural network model provided in an embodiment of this application;
[0073] Figure 9 This is a schematic block diagram of a training device for a neural network model provided in an embodiment of this application;
[0074] Figure 10 This is a schematic block diagram of a computing device for a neural network model provided in an embodiment of this application;
[0075] Figure 11 This is a schematic block diagram of a training device for another neural network model provided in an embodiment of this application;
[0076] Figure 12 This is a schematic block diagram of a computing device for another neural network model provided in an embodiment of this application. Detailed Implementation
[0077] The technical solutions in this application will now be described with reference to the accompanying drawings.
[0078] Figure 1 A schematic diagram of an artificial intelligence framework is shown, which describes the overall workflow of an artificial intelligence system and is applicable to general artificial intelligence domain needs.
[0079] The above-mentioned artificial intelligence framework will be elaborated in detail from two dimensions: the "intelligent information chain" (horizontal axis) and the "information technology (IT) value chain" (vertical axis).
[0080] The "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it could be a general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. In this process, data undergoes a condensation process of "data—information—knowledge—wisdom."
[0081] The "IT value chain" reflects the value that artificial intelligence brings to the information technology industry, from the underlying infrastructure of human intelligence, information (provided and processed by technology) to the industrial ecosystem of systems.
[0082] (1) Infrastructure:
[0083] Infrastructure provides computing power to support artificial intelligence systems, enables them to communicate with the outside world, and provides support through basic platforms.
[0084] Infrastructure can communicate with the outside world through sensors, and its computing power can be provided by smart chips.
[0085] The intelligent chips here can be hardware acceleration chips such as central processing units (CPUs), neural network processing units (NPUs), graphics processing units (GPUs), application-specific integrated circuits (ASICs), and field-programmable gate arrays (FPGAs).
[0086] The basic platform of the infrastructure can include distributed computing frameworks and related platform guarantees and support, such as cloud storage and computing, and interconnected networks.
[0087] For example, for infrastructure, data can be acquired through sensors and external communication, and then this data can be provided to intelligent chips in the distributed computing system provided by the basic platform for computation.
[0088] (2) Data:
[0089] The data at the next layer of infrastructure is used to represent data sources in the field of artificial intelligence. This data includes graphics, images, voice, text, and IoT data from traditional devices, including business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.
[0090] (3) Data processing:
[0091] The aforementioned data processing typically includes data training, machine learning, deep learning, search, reasoning, and decision-making.
[0092] Among them, machine learning and deep learning can perform intelligent information modeling, extraction, preprocessing, and training on data, including symbolization and formalization.
[0093] Reasoning refers to the process in which, in a computer or intelligent system, the machine thinks and solves problems by simulating human intelligent reasoning, based on reasoning control strategies and using formalized information. Typical functions include search and matching.
[0094] Decision-making refers to the process of making decisions based on intelligent information after reasoning, and it typically provides functions such as classification, sorting, and prediction.
[0095] (4) General abilities:
[0096] After the data processing mentioned above, the results of the data processing can be used to form some general capabilities, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0097] (5) Smart products and industry applications:
[0098] Intelligent products and industry applications refer to products and applications of artificial intelligence systems in various fields. They are the encapsulation of overall artificial intelligence solutions, productizing intelligent information decision-making and realizing practical applications. Their application areas mainly include: intelligent manufacturing, intelligent transportation, smart home, intelligent healthcare, intelligent security, autonomous driving, and intelligent terminals.
[0099] The embodiments of this application can be applied to many fields of artificial intelligence, such as intelligent manufacturing, intelligent transportation, smart home, intelligent healthcare, intelligent security, and autonomous driving.
[0100] Specifically, the embodiments of this application can be applied to fields that require the use of (deep) neural networks, such as autonomous driving, image classification, image retrieval, image semantic segmentation, image quality enhancement, image super-resolution, and natural language processing.
[0101] The following is a brief introduction to two application scenarios: image classification and surveillance.
[0102] Image Categories:
[0103] When users store a large number of pictures on their terminal devices (such as mobile phones) or cloud storage, recognizing the images in the album can make it easier for users or the system to classify and manage the album, thus improving the user experience.
[0104] The computational method using the neural network model in this application reduces hardware overhead and is more user-friendly for terminal devices. Furthermore, it improves the speed of image classification using the neural network, facilitating real-time tagging of images into different categories for easier user viewing and searching. Additionally, these image classification tags can be provided to the photo album management system for categorized management, saving users' management time, improving album management efficiency, and enhancing the user experience.
[0105] monitor:
[0106] Monitoring scenarios include: smart cities, field monitoring, indoor monitoring, outdoor monitoring, and vehicle monitoring. In smart city scenarios, multiple attribute recognition is required, such as pedestrian and cyclist attribute recognition. Deep neural networks, with their powerful capabilities, play a crucial role in this multi-attribute recognition.
[0107] By employing the computational method of the neural network model in the embodiments of this application, the processing efficiency of the neural network model can be improved, which is beneficial for real-time processing of the input road images, faster identification of different attribute information in the road images, and reduction of power consumption.
[0108] Since the embodiments of this application involve a large number of neural network applications, for ease of understanding, the relevant terms and concepts of neural networks that may be involved in the embodiments of this application will be introduced below.
[0109] (1) Neural Network
[0110] Neural networks can be composed of neural units, which can refer to units represented by x. s The arithmetic unit that takes an intercept of 1 as input can output the following:
[0111]
[0112] Where s = 1, 2, ..., n, n is a natural number greater than 1, W s For x s The weights are denoted by b, where b is the bias of the neural unit.
[0113] f represents the activation function of the neural network, which introduces nonlinear characteristics to transform the input signal into the output signal. The output signal of this activation function can be used as the input to the next layer. For example, the activation function can be ReLU, tanh, or sigmoid.
[0114] A neural network is a network formed by connecting multiple individual neural units, meaning that the output of one neural unit can be the input of another. The input of each neural unit can be connected to the local receptive field of the previous layer to extract features from the local receptive field, which can be a region composed of several neural units.
[0115] (2) Deep Neural Networks
[0116] A deep neural network (DNN), also known as a multilayer neural network, can be understood as a neural network with multiple hidden layers. Based on the position of the layers, the internal neural network of a DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. The layers are fully connected, meaning that any neuron in the i-th layer is connected to any neuron in the (i+1)-th layer.
[0117] Although DNNs seem complex, the operation of each layer is actually not complicated. Simply put, it involves the following linear relationship expression: in, It is the input vector. It is the output vector. α is the offset vector, W is the weight matrix (also called coefficients), and α() is the activation function. Each layer is simply an adjustment of the input vector. The output vector is obtained through this simple operation. Due to the large number of layers in a DNN, the coefficients W and the offset vector... The number of these parameters is also relatively large. The definitions of these parameters in DNNs are as follows: Taking the coefficient W as an example: Assuming a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as... The superscript 3 represents the layer number where coefficient W is located, while the subscript corresponds to the third layer index 2 of the output and the second layer index 4 of the input.
[0118] In summary, the coefficient from the k-th neuron in layer L-1 to the j-th neuron in layer L is defined as...
[0119] It's important to note that the input layer has no parameters. In deep neural networks, more hidden layers allow the network to better represent complex real-world situations. Theoretically, the more parameters a model has, the higher its complexity and "capacity," meaning it can perform more complex learning tasks. Training a deep neural network is essentially the process of learning the weight matrix, with the ultimate goal of obtaining the weight matrix of all layers in the trained deep neural network (a weight matrix formed by vectors from many layers).
[0120] (3) Convolutional Neural Network
[0121] A convolutional neural network (CNN) is a deep neural network with a convolutional structure. A CNN contains a feature extractor consisting of convolutional layers and subsampling layers, which can be viewed as a filter. A convolutional layer is a layer of neurons in a CNN that performs convolutional processing on the input signal. In a convolutional layer of a CNN, a neuron may only be connected to some of its neighboring neurons. A convolutional layer typically contains several feature planes, each composed of rectangularly arranged neural units. Neural units on the same feature plane share weights, which are called the convolutional kernel. Shared weights can be understood as the way image information is extracted regardless of location. The convolutional kernel can be formalized as a matrix of random size, and during the training process of the CNN, the kernel can learn appropriate weights. Furthermore, the direct benefit of shared weights is reducing the connections between layers in the CNN, while also reducing the risk of overfitting.
[0122] (4) Loss Function
[0123] In training deep neural networks, to ensure the output closely approximates the desired predicted value, we compare the network's prediction with the target value and update the weight vector of each layer based on the difference. (Of course, there's usually a pre-configuration process before the first update, where parameters are pre-configured for each layer.) For example, if the prediction is too high, the weight vector is adjusted to predict a lower value. This adjustment continues until the deep neural network can predict the target value or a value very close to it. Therefore, we need to predefine "how to compare the difference between the predicted and target values," which is the loss function or objective function. These are important equations used to measure the difference between the predicted and target values. Taking the loss function as an example, a higher output value (loss) indicates a greater difference, and training the deep neural network becomes a process of minimizing this loss. Generally, a smaller loss indicates higher training quality, while a larger loss indicates lower training quality. Similarly, smaller loss fluctuations result in more stable training, while larger loss fluctuations lead to less stable training.
[0124] (5) Backpropagation algorithm
[0125] Neural networks can employ backpropagation (BP) to correct the parameters of the neural network model during training, thereby reducing the reconstruction error loss. Specifically, forward propagation of the input signal to the output generates error loss; this error loss information is then propagated back to update the parameters of the neural network model, leading to convergence of the error loss. The backpropagation algorithm is an error-loss-driven backpropagation process aimed at obtaining the optimal parameters of the neural network model, such as the weight matrix.
[0126] For example, the loss value generated during each training iteration of a neural network model is passed from the back to the front layer of the model. At each layer, the update amount of the layer's parameters is calculated (partial derivative operation), and this update amount is related to the gradient.
[0127] (6) Adder neural network (AdderNet)
[0128] The structure of AdderNet is similar to that of CNN. Convolutional layers in CNNs can be used for feature extraction or filtering of input data. Similarly, adder layers in AdderNet can also be used for feature extraction or filtering of input data. The parameters of the convolutional kernels in CNNs can be understood as the parameters of the adders in AdderNet. Both the convolutional kernels in CNNs and the adders in AdderNet can be understood as filters.
[0129] In a CNN, each convolutional layer extracts feature information from the input data through convolution operations. AdderNet uses L1 distance to calculate the output features. Specifically, the adder layer in AdderNet extracts feature information from the input data through addition (or subtraction) operations and absolute value operations.
[0130] Because the computational complexity of addition is much lower than that of multiplication, AdderNet's power consumption is significantly lower than that of comparable CNNs. For example, AdderNet is obtained by replacing convolution operations in a CNN with addition (or subtraction) operations and absolute value operations. This significantly reduces the power consumption of CNNs while maintaining performance.
[0131] The output feature map of the adder layer in AdderNet satisfies the following formula:
[0132]
[0133] Where Y(m,n,t) represents the element in the m-th row, n-th column, and t-th page of the output feature map, X(m+i,n+j,k) represents the element in the (m+i)-th row, n+j-th column, and k-th page of the input feature map, F(i,j,k,t) represents the element in the i-th row, j-th class, and k-th page of the filter, t represents the number of channels of the filter, d represents the number of rows of the filter, and c in This indicates the number of channels in the input feature map.
[0134] AdderNet uses L1 distance to extract features during the forward computation process, that is, it uses addition operations to replace multiplication operations, which reduces computational complexity, power consumption and hardware area.
[0135] (7) Winograd algorithm
[0136] The Winograd algorithm is a commonly used method for fast convolution calculation. It can significantly reduce the computational cost of CNNs and improve computational efficiency without affecting the calculation results.
[0137] The winograd algorithm satisfies the following formula:
[0138]
[0139] Where Y represents the output data matrix; g represents the convolution kernel, i.e. the weight matrix; and d represents the title of the input feature map, i.e. the input data matrix. This represents element-wise multiplication, i.e., the dot product of matrices. A, G, and B all represent transformation matrices; specifically, A represents the output transformation matrix, G represents the weight transformation matrix, and B represents the input transformation matrix. During the inference process of the neural network, g remains unchanged; therefore, the transformed weight matrix... It can be pre-computed, that is, the transformed weight matrix is pre-computed before the forward computation begins. This can further reduce the computational power consumption of the neural network model.
[0140] The winograd algorithm, in the form F(m, r), is used to quickly compute convolution operations with a kernel size of r and an output feature map size of m. Size can also be understood as dimension.
[0141] The transformation matrix can be determined based on the dimension of the convolution kernel and the stride. Alternatively, B, G, and A are fixed combinations of kernel size and stride for a given size and can be derived using the Winograd algorithm.
[0142] In practical applications, the Winograd algorithm is commonly used in the form of F(2×2, 3×3).
[0143] The output data matrix Y is a 2×2 matrix, and the weight matrix g is a 3×3 matrix.
[0144] For a 3×3 weight matrix and a combination with a stride of 1, the transformation matrix satisfies the following formula:
[0145]
[0146] The following example illustrates the computation process of the winograd algorithm for F(2×2,3×3).
[0147] 1) Transform the 4×4 input data matrix d using the 4×4 input transformation matrix B to obtain the transformed 4×4 input data matrix.
[0148] For example, the input data matrix d satisfies the following formula.
[0149]
[0150] The transformed input data matrix satisfies the following formula:
[0151]
[0152] 2) Transform the 3×3 weight matrix g using a 4×3 weight transformation matrix G to obtain the transformed 4×4 weight matrix.
[0153] For example, the weight matrix g satisfies the following formula.
[0154]
[0155] The transformed weight matrix satisfies the following formula:
[0156]
[0157] 3) Perform element-wise multiplication between the transformed input data matrix and the transformed weight matrix, that is, multiply the corresponding elements in the two matrices to obtain a 4×4 intermediate matrix.
[0158] 4) Transform the intermediate matrix using a 4×2 output transformation matrix A to obtain a 2×2 output data matrix. The output matrix is the result of convolution between the input data matrix d and the weight matrix g.
[0159] If convolution is used to obtain four results in the output data matrix, each result requires 9 (3*3) multiplication operations, and the four results require 36 multiplication operations. If the winograd algorithm is used, excluding the transformation overhead, only 16 (4*4) multiplication operations are needed in step 3), resulting in a speedup of 2.25 (36 / 16) times. The elements in the transformation matrix are all values of 0, ±1, and ±1 / 2, which can be accomplished through lightweight hardware operations such as changing the sign bit and shifting operations. In other words, the transformation overhead is usually small.
[0160] like Figure 2 As shown, this application embodiment provides a system architecture 100. In Figure 2 In this embodiment, the data acquisition device 160 is used to acquire training data. For example, in the training method of the neural network model according to this application embodiment, if the training data is image data, the training data may include training images and the corresponding processing results of the training images. For example, the classification result corresponding to the training image may be a result pre-annotated manually.
[0161] After collecting the training data, the data acquisition device 160 stores the training data in the database 130, and the training device 120 trains the target model / rule 101 based on the training data maintained in the database 130.
[0162] The following describes how the training device 120 obtains the target model / rule 101 based on the training data. The training device 120 processes the input raw data and compares the output value with the target value until the difference between the output value of the training device 120 and the target value is less than a certain threshold, thereby completing the training of the target model / rule 101.
[0163] The target model / rule 101 in this embodiment can specifically be a neural network model, such as a convolutional neural network or a residual network. It should be noted that in practical applications, the training data maintained in the database 130 may not all come from the data acquisition device 160; it may also be received from other devices. Furthermore, it should be noted that the training device 120 may not necessarily train the target model / rule 101 entirely based on the training data maintained in the database 130; it may also obtain training data from the cloud or other sources for model training. The above description should not be construed as limiting the embodiments of this application.
[0164] The target model / rule 101 trained using training device 120 can be applied to different systems or devices, such as... Figure 2The execution device 110 shown can be a terminal, such as a mobile phone terminal, tablet computer, laptop computer, augmented reality (AR) / virtual reality (VR) device, vehicle terminal, etc., or it can be a server or cloud service. Figure 2 In this embodiment, the execution device 110 is configured with an input / output (I / O) interface 112 for data interaction with external devices. Users can input data to the I / O interface 112 through the client device 140. The input data may include data to be processed input by the client device.
[0165] During the preprocessing of input data by the execution device 110, or during the calculation module 111 of the execution device 110 performing calculations and other related processes, the execution device 110 can call data, code, etc. in the data storage system 150 for corresponding processing, or store the data, instructions, etc. obtained from the corresponding processing into the data storage system 150.
[0166] Finally, I / O interface 112 returns the processing result, such as the data processing result obtained above, to client device 140, thereby providing it to the user.
[0167] It is worth noting that the training device 120 can generate corresponding target models / rules 101 based on different training data for different objectives or tasks. The corresponding target models / rules 101 can be used to achieve the above objectives or complete the above tasks, thereby providing the user with the required results.
[0168] exist Figure 2 In the scenario shown, the user can manually provide input data, which can be done through the interface provided by I / O interface 112. Alternatively, the client device 140 can automatically send input data to I / O interface 112. If user authorization is required for the client device 140 to automatically send input data, the user can set the corresponding permissions in the client device 140. The user can view the output results of the execution device 110 on the client device 140, which can be presented in various forms such as display, sound, or animation. The client device 140 can also act as a data acquisition terminal, collecting the input data and output results of the input I / O interface 112 as new sample data and storing them in the database 130. Alternatively, data can be collected directly from the I / O interface 112 without going through the client device 140, using the input data and output results of the input I / O interface 112 as new sample data and storing them in the database 130.
[0169] It is worth noting that, Figure 2 This is merely a schematic diagram of a system architecture provided in an embodiment of this application. The positional relationships between the devices, components, modules, etc., shown in the diagram do not constitute any limitation. For example, in Figure 2 In this context, the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 may also be placed within the execution device 110.
[0170] like Figure 2 As shown, the target model / rule 101 is obtained by training the training device 120. The target model / rule 101 can be the neural network in this application embodiment. Specifically, the neural network in this application embodiment can be a CNN or a residual network, etc.
[0171] CNN is a very common type of neural network. Below, we will combine... Figure 3 This section focuses on a detailed explanation of the structure of CNNs. As mentioned in the basic concept introduction above, a Convolutional Neural Network (CNN) is a deep neural network with a convolutional structure. It is a deep learning architecture, which refers to learning at multiple levels of abstraction through machine learning algorithms. As a deep learning architecture, CNN is a feed-forward artificial neural network, in which each neuron can respond to the input image.
[0172] like Figure 3 As shown, the convolutional neural network (CNN) 200 may include an input layer 210, a convolutional / pooling layer 220 (where the pooling layer is optional), and a fully connected layer 230.
[0173] Convolutional / pooling layers 220:
[0174] Convolutional layers:
[0175] like Figure 3 The convolutional / pooling layer 220 shown may include layers as in Examples 221-226. For instance, in one implementation, layer 221 is a convolutional layer, layer 222 is a pooling layer, layer 223 is a convolutional layer, layer 224 is a pooling layer, layer 225 is a convolutional layer, and layer 226 is a pooling layer; in another implementation, layers 221 and 222 are convolutional layers, layer 223 is a pooling layer, layers 224 and 225 are convolutional layers, and layer 226 is a pooling layer. That is, the output of the convolutional layer can be used as the input to a subsequent pooling layer, or as the input to another convolutional layer to continue the convolution operation.
[0176] The following section will use convolutional layer 221 as an example to introduce the internal working principle of a convolutional layer.
[0177] Convolutional layer 221 can include multiple convolution operators, also known as kernels. In image processing, a convolution operator acts as a filter to extract specific information from the input image matrix. Essentially, a convolution operator can be a weight matrix, which is usually predefined. During the convolution operation, the weight matrix typically processes the input image pixel by pixel (or two pixels by two pixels, depending on the stride) along the horizontal direction, thus extracting specific features from the image. The size of the weight matrix should be related to the image size. It's important to note that the depth dimension of the weight matrix is the same as the depth dimension of the input image; during convolution, the weight matrix extends to the entire depth of the input image. Therefore, convolution with a single weight matrix produces a single-depth convolutional output. However, in most cases, a single weight matrix is not used; instead, multiple weight matrices of the same size (rows × columns) are applied—multiple identical matrices. The outputs of each weight matrix are stacked to form the depth dimension of the convolutional image; this dimension can be understood as being determined by the "multiple" mentioned above. Different weight matrices can be used to extract different features from an image. For example, one weight matrix can be used to extract image edge information, another weight matrix can be used to extract specific colors of the image, and yet another weight matrix can be used to blur unwanted noise in the image. These multiple weight matrices have the same size (rows × columns), and the feature maps extracted by these multiple weight matrices of the same size also have the same size. The extracted feature maps of the same size are then merged to form the output of the convolution operation.
[0178] The weight values in these weight matrices need to be obtained through extensive training in practical applications. The weight matrices formed by the weight values obtained through training can be used to extract information from the input image, thereby enabling the convolutional neural network 200 to make correct predictions.
[0179] When a convolutional neural network 200 has multiple convolutional layers, the initial convolutional layers (e.g., 221) tend to extract more general features, which can also be called low-level features. As the depth of the convolutional neural network 200 increases, the features extracted by later convolutional layers (e.g., 226) become more and more complex, such as high-level semantic features. Features with higher semantic levels are more suitable for the problem to be solved.
[0180] Pooling layer:
[0181] Because it is often necessary to reduce the number of training parameters, pooling layers are often introduced periodically after convolutional layers, such as... Figure 3Layers 221-226 in example 220 can be a convolutional layer followed by a pooling layer, or multiple convolutional layers followed by one or more pooling layers. In image processing, the sole purpose of pooling layers is to reduce the spatial size of the image. Pooling layers can include average pooling and / or max pooling operators to sample the input image to obtain a smaller image size. The average pooling operator calculates the average value of pixel values within a specific range as the result of average pooling. The max pooling operator takes the pixel with the largest value within a specific range as the result of max pooling. Furthermore, just as the size of the weight matrix in a convolutional layer should be related to the image size, the operators in a pooling layer should also be related to the image size. The size of the output image after processing by the pooling layer can be smaller than the size of the input image of the pooling layer. Each pixel in the output image of the pooling layer represents the average or maximum value of the corresponding sub-region of the input image of the pooling layer.
[0182] Fully connected layer 230:
[0183] After processing by the convolutional / pooling layers 220, the convolutional neural network 200 is still insufficient to output the required information. As mentioned earlier, the convolutional / pooling layers 220 only extract features and reduce the parameters introduced by the input image. However, to generate the final output information (the required class information or other relevant information), the convolutional neural network 200 needs to utilize fully connected layers 230 to generate one or a set of outputs representing the required number of classes. Therefore, the fully connected layers 230 can include multiple hidden layers (such as...). Figure 3 As shown in 231, 232 to 23n), the parameters contained in this multi-layer hidden layer can be pre-trained based on relevant training data for a specific task type. For example, the task type may include image recognition, image classification, image super-resolution reconstruction, etc.
[0184] After the multiple hidden layers in the fully connected layer 230, the final layer of the entire convolutional neural network 200 is the output layer 240. This output layer 240 has a loss function similar to the classification cross-entropy, specifically used to calculate the prediction error. Once the entire convolutional neural network 200 has propagated forward (e.g., ... Figure 3 Propagation from 210 to 240 degrees is considered forward propagation, while backward propagation (e.g.) is completed. Figure 3 The propagation from 240 to 210 (backpropagation) will begin to update the weight values and biases of the layers mentioned above, in order to reduce the loss of the convolutional neural network 200 and the error between the output of the convolutional neural network 200 through the output layer and the ideal result.
[0185] It should be noted that, as Figure 3The convolutional neural network 200 shown is merely an example of a convolutional neural network. In specific applications, convolutional neural networks can also exist in the form of other network models, for example, including only... Figure 3 As shown in the network structure, for example, the convolutional neural network used in the embodiments of this application may only include an input layer 210, a convolutional / pooling layer 220, and an output layer 240.
[0186] Figure 4 The present application provides a hardware structure for a chip, which includes a neural network processor 50. This chip can be configured as follows: Figure 2 The execution device 110 shown is used to perform the calculations of the calculation module 111. This chip can also be located in, for example... Figure 2 The training device 120 shown is used to complete the training work of the training device 120 and output the target model / rule 101. The method in this embodiment can be used as follows: Figure 4 This is achieved in the chip shown.
[0187] The neural network processor 50 can be any processor suitable for large-scale XOR operations, such as a neural network processing unit (NPU), tensor processing unit (TPU), or graphics processing unit (GPU). Taking an NPU as an example: the neural network processor NPU 50 is mounted as a coprocessor on the main central processing unit (CPU) (host CPU), and tasks are assigned by the host CPU. The core of the NPU is the arithmetic circuit 503, and the controller 504 controls the arithmetic circuit 503 to retrieve data from memory (weight memory or input memory) and perform operations. The TPU is a Google-customized application-specific integrated circuit for machine learning AI accelerators.
[0188] In some implementations, the arithmetic circuit 503 internally includes multiple process engines (PEs). In some implementations, the arithmetic circuit 503 is a two-dimensional pulsating array. The arithmetic circuit 503 can also be a one-dimensional pulsating array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 503 is a general-purpose matrix processor.
[0189] For example, suppose we have an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from the weight memory 502 and caches it in each PE of the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from the input memory 501 and performs matrix operations with matrix B. The partial result or the final result of the obtained matrix is stored in the accumulator 508.
[0190] The vector computation unit 507 can further process the output of the arithmetic circuit, such as vector multiplication, vector addition, exponentiation, logarithmic operations, size comparisons, etc. For example, the vector computation unit 507 can be used for network computation in non-convolutional / non-FC layers of neural networks, such as pooling, batch normalization (BN), local response normalization, etc.
[0191] In some implementations, the vector computation unit 507 can store the processed output vector into a unified buffer 506. For example, the vector computation unit 507 can apply a nonlinear function to the output of the arithmetic circuit 503, such as a vector of accumulated values, to generate activation values. In some implementations, the vector computation unit 507 generates normalized values, merged values, or both. In some implementations, the processed output vector can be used as activation input to the arithmetic circuit 503, for example, for use in subsequent layers of a neural network.
[0192] The unified memory 506 is used to store input data and output data.
[0193] The weight data is directly transferred from the external memory to the input memory 501 and / or the unified memory 506 through the direct memory access controller 505 (DMAC), the weight data in the external memory is stored in the weight memory 502, and the data in the unified memory 506 is stored in the external memory.
[0194] The bus interface unit (BIU) 510 is used to enable interaction between the main CPU, DMAC and instruction fetch memory 509 via a bus.
[0195] The instruction fetch buffer 509, which is connected to the controller 504, is used to store the instructions used by the controller 504.
[0196] The controller 504 is used to call the instructions cached in the instruction memory 509 to control the operation of the computing accelerator.
[0197] Generally, the unified memory 506, input memory 501, weight memory 502, and instruction fetch memory 509 are all on-chip memories, while the external memory is memory outside the NPU. This external memory can be double data rate synchronous dynamic random access memory (DDR SDRAM), high bandwidth memory (HBM), or other readable and writable memory.
[0198] The above-mentioned Figure 2 The execution device 110 or Figure 4 The chip in the chip is capable of executing each step of the computation method of the neural network model in the embodiments of this application. The above-described... Figure 2 Training equipment 120 or Figure 4 The chip in the application is capable of executing the various steps of the training method for the neural network model of the present application embodiment.
[0199] like Figure 5 As shown, this application embodiment provides a system architecture 300. The system architecture includes a local device 301, a local device 302, an execution device 310, and a data storage system 350, wherein the local devices 301 and 302 are connected to the execution device 310 through a communication network.
[0200] The execution device 310 can be implemented by one or more servers. Optionally, the execution device 310 can be used in conjunction with other computing devices, such as data storage devices, routers, load balancers, etc. The execution device 310 can be deployed on a single physical site or distributed across multiple physical sites. The execution device 310 can use data in the data storage system 350 or call program code in the data storage system 350 to implement the computation method or training method of the neural network model of the embodiments of this application.
[0201] Specifically, in one implementation, the execution device 110 can perform the following process:
[0202] In at least one feature extraction layer of the neural network model, the following operations are performed: The input data matrix of the data to be processed is transformed using the Winograd algorithm to obtain a transformed input data matrix; Features are extracted from the transformed input data matrix using the transformed weight matrix to obtain an intermediate matrix, wherein the transformed weight matrix is obtained by transforming the weight matrix of at least one feature extraction layer using the Winograd algorithm, and each element in the intermediate matrix is determined based on the L1 distance between the corresponding elements in the transformed input data matrix and the transformed weight matrix; The intermediate matrix is then transformed using the Winograd algorithm to obtain an output data matrix.
[0203] Users can interact with execution device 310 by operating their respective user devices (e.g., local device 301 and local device 302). Each local device can represent any computing device, such as a personal computer, computer workstation, smartphone, tablet, smart camera, smart car or other type of cellular phone, media consumption device, wearable device, set-top box, game console, etc.
[0204] Each user's local device can interact with the execution device 310 through a communication network of any communication mechanism / standard. The communication network can be a wide area network, a local area network, a point-to-point connection, or any combination thereof.
[0205] In one implementation, local devices 301 and 302 obtain relevant parameters of the neural network from execution device 310, deploy the neural network on local devices 301 and 302, and use the neural network to perform image classification, image processing, speech processing, or text processing, etc.
[0206] In another implementation, a neural network can be directly deployed on the execution device 310. The execution device 310 obtains the data to be processed from local devices 301 and 302 and processes the data using a neural network model.
[0207] The aforementioned execution device 310 can also be a cloud device, in which case the execution device 310 can be deployed in the cloud; or, the aforementioned execution device 310 can also be a terminal device, in which case the execution device 310 can be deployed on the user terminal side. This application embodiment does not limit this.
[0208] Neural network models typically involve a large number of matrix operations, such as convolution. The latency of matrix operations becomes a major factor limiting computational efficiency, affecting the overall processing efficiency of the neural network model and making it difficult to deploy in scenarios with high real-time requirements. Furthermore, when the computational load of a neural network model is large, it places high demands on the hardware's computing power, making it difficult to deploy on hardware devices with lower computing power, such as mobile phones and other terminal devices.
[0209] Therefore, how to reduce the computational overhead of neural network models has become an urgent problem to be solved.
[0210] This application provides a method for operating a neural network model, which replaces the element-wise multiplication operation in the Winograd algorithm with an addition operation, further reducing the computational load of the neural network model, improving the operation speed of the neural network model, and reducing computational overhead.
[0211] To better illustrate the methods of the embodiments of this application, the computing apparatus of the neural network model of the embodiments of this application will be described below with reference to the accompanying drawings.
[0212] Figure 6 This application illustrates a computing device for a neural network model provided in an embodiment, such as... Figure 6 As shown, the device 600 includes an input data preprocessing module 610, a weight preprocessing module 620, and an acceleration module 630.
[0213] The input data preprocessing module 610 uses the Winograd algorithm to transform the input data matrix, obtaining the transformed input data matrix. The input data matrix can be part or all of the input feature map. For example, if the dimension of the input feature map is 8*8, the input data matrix can be 4*4 or 8*8.
[0214] The weight preprocessing module 620 performs weight transformation on the weight matrix using the winograd algorithm to obtain the transformed weight matrix.
[0215] It should be noted that the weight preprocessing module 620 is an optional module.
[0216] For example, the transformed weight matrix may be pre-stored in device 600, or the transformed weight matrix may be sent to device 600 by other devices.
[0217] In other words, the transformed weight matrix can be calculated either before the neural network model's computation (i.e., offline computation) or during the neural network model's computation (i.e., online computation). During the neural network model's inference process, the weight matrix remains unchanged. Obtaining the transformed weight matrix through offline computation can further improve computational speed and reduce computational overhead.
[0218] The acceleration module 630 is used to extract features from the transformed input data matrix using the transformed weight matrix to obtain an intermediate matrix; and to perform a Winograd transformation on the intermediate matrix to obtain an output data matrix. Each element in the intermediate matrix is determined based on the L1 distance between the corresponding element in the transformed input data matrix and the transformed weight matrix. The output data matrix can be part or all of the output feature map.
[0219] Specifically, the acceleration module can obtain an intermediate matrix by performing subtraction, absolute value calculation, and inversion operations on the transformed input data matrix and the transformed weight matrix.
[0220] For details of the process, please refer to Method 700 below, which will not be repeated here.
[0221] For example, the acceleration module 630 may include a matrix operation module and a post-processing module. The matrix operation module is used to extract features from the transformed input data matrix using the transformed weight matrix to obtain an intermediate matrix. The post-processing module is used to perform output data transformation on the intermediate matrix using the Winograd algorithm to obtain an output data matrix.
[0222] The method described in this application embodiment can be executed by a device capable of performing the winograd algorithm. Specifically, Figure 6 The operations in each module can be performed by software, hardware, or a combination of both. This application does not limit this approach.
[0223] For example, Figure 6 The operations of each module in the process can be performed by Figure 4 The chip executes within it.
[0224] Specifically, Figure 4 The external memory in the neural network processor 50 can be used to store input data matrices and computation results, etc. The neural network processor 50 can retrieve the input data matrix from the external memory.
[0225] The operations in the input data preprocessing module 610 can be performed by the vector calculation unit 507. Alternatively, the vector calculation unit 507 may include the input data preprocessing module 610.
[0226] Furthermore, the external memory can be used to store the transformed weight matrix obtained through offline computation. In this case, the neural network processor 50 can retrieve the input data matrix and the transformed weight matrix from the external memory.
[0227] If the transformed weight matrix obtained through offline calculation is not stored in the external memory, the operations in the weight preprocessing module 620 can be performed by the vector calculation module 507. Alternatively, the vector calculation module 507 may include the weight preprocessing module 620.
[0228] The matrix operation module in acceleration module 630 can perform operations by operation circuit 503. Alternatively, operation circuit 503 may include the matrix operation module.
[0229] The operations performed by the post-processing module in acceleration module 630 can be executed by vector calculation module 507. Alternatively, vector calculation module 507 may include the post-processing module.
[0230] Alternatively, the acceleration module 630 can also be in Figure 4 The dedicated module set up on the chip shown can further improve the processing speed of the acceleration module 630 compared to the operation performed by the arithmetic circuit and vector calculation module.
[0231] The following is combined with Figure 7 The computational methods of the neural network model in the embodiments of this application are described in detail.
[0232] Figure 7 The present application illustrates a computation method 700 for a neural network model provided in an embodiment of the present application. Figure 7 The method shown can be executed by an execution device for a neural network model. This device can be a cloud service device or a terminal device, such as a computer, server, or other device with sufficient computing power to execute neural network model operations. It can also be a system composed of cloud service devices and terminal devices. For example, method 700 can be executed by... Figure 2 Execution device 110 in Figure 4 The neural network processor 50 or Figure 5 The method can be executed by the execution device 310 or a local device. Alternatively, method 700 can also be executed by a device that provides AutoML services. For example, the device that provides AutoML services can be a cloud service device.
[0233] For example, method 700 can be specifically derived from, such as Figure 2 The execution device 110 shown executes the process, and the data to be processed in method 700 can be as follows: Figure 2 The input data provided by the client device 140 shown.
[0234] Method 700 includes steps S710 to S730. Steps S710 to S730 are described in detail below.
[0235] Perform the following steps in at least one feature extraction layer of the neural network model:
[0236] S710 uses the winograd algorithm to transform the input data matrix of the data to be processed, and obtains the transformed input data matrix.
[0237] The data to be processed includes image data, voice data, or text data.
[0238] The type of data to be processed is related to the task of the neural network model. For example, if the neural network model is used for image processing tasks, the data to be processed can be images. Specifically, image processing tasks include image classification, image detection, image segmentation, image recognition, or image generation. As another example, if the neural network model is used for text processing tasks, the data to be processed can be text. Specifically, text processing tasks include text recognition or text translation. Similarly, if the neural network model is used for speech processing tasks, the data to be processed can be speech data. Specifically, speech processing tasks include speech recognition. The embodiments of this application do not limit the type of data to be processed.
[0239] For example, the data to be processed is an image. The image to be processed can be an image captured by a camera of a terminal device (or a computer, server, or other device or equipment), or the image to be processed can be an image obtained from within the terminal device (or a computer, server, or other device or equipment) (e.g., an image stored in the photo album of the terminal device, or an image obtained by the terminal device from the cloud). This application embodiment does not limit this.
[0240] The neural network model in this application embodiment can be an existing neural network model, such as a residual network. Alternatively, the neural network model can also be a self-constructed neural network model with other structures. This application embodiment does not limit this.
[0241] The input data matrix of the data to be processed refers to the data matrix input to at least one feature extraction layer.
[0242] For example, the feature extraction layer includes an adder layer, which uses addition operations as filters for feature extraction. For a detailed description, please refer to the feature extraction layer in AdderNet mentioned earlier; it will not be repeated here.
[0243] Alternatively, the feature extraction layer may include a convolutional layer, i.e., using convolutional kernels as filters for feature extraction. For a detailed description, please refer to the section on convolutional neural networks above; it will not be repeated here. In this case, the convolutional layer can be replaced with an adder layer, and then method 700 can be applied to the feature extraction layer.
[0244] Specifically, the input data matrix is transformed by the input transformation matrix using a winograd transformation to obtain the transformed input data matrix.
[0245] Winograd transformation includes input data transformation, weight transformation, and output data transformation performed by the Winograd algorithm. In other words, all three types of transformation can be understood as Winograd transformation.
[0246] For example, the transformed input data matrix Satisfy the following formula:
[0247]
[0248] Where B represents the input transformation matrix, B T Let d denote the transpose of B, and let d denote the input data matrix before the transformation.
[0249] For example, the input transformation matrix can be an existing winograd input transformation matrix.
[0250] For example, if the input data matrix d before transformation is a 4×4 matrix, and the weight matrix g before transformation is a 3×3 matrix, the input transformation matrix can be:
[0251]
[0252] The input data matrix can be part or all of the input feature map. For example, if the input feature map has a dimension of 8*8, the input data matrix can be a 4*4 matrix of the input feature map, or the input data matrix can be the input feature map itself.
[0253] It should be noted that the input feature map can be the data itself input into the neural network model, such as the image to be processed. The input feature map can also be a feature map obtained after one or more feature extraction layers in the neural network model, for example, a feature map obtained after one or more feature extractions on the image to be processed. The feature extraction method can employ existing schemes, such as convolutional layer processing. That is, the input feature map in this embodiment can be a feature map obtained after one or more convolutional processes on the image to be processed. Alternatively, the feature extraction method can also employ method 700 in this embodiment. In other words, the input feature map in this embodiment can be a feature map obtained after processing the image to be processed using method 700 in this embodiment. This embodiment does not limit the method of obtaining the input feature map.
[0254] For example, step S710 can be performed by Figure 6 The input data preprocessing module 610 in the middle is executed.
[0255] S720: Feature extraction is performed on the transformed input data matrix using the transformed weight matrix to obtain an intermediate matrix. The transformed weight matrix is obtained by applying a weight transformation to the weight matrix of at least one feature extraction layer using the Winograd algorithm. Each element in the intermediate matrix is determined based on the L1 distance between the corresponding element in the transformed input data matrix and the transformed weight matrix.
[0256] L1 distance can also be called L1 regular distance, L1 norm distance, Manhattan distance, or taxi distance.
[0257] The transformed weight matrix can be obtained by performing a Winograd transformation on the weight matrix using a weight transformation matrix.
[0258] Specifically, the weight matrix represents the parameters of one or more feature extraction kernels in the at least one feature extraction layer of the neural network model. A feature extraction layer may include one or more feature extraction kernels, i.e., filters. Feature extraction kernels are used to extract features from the data input to the neural network model.
[0259] For example, the transformed weight matrix Satisfy the following formula:
[0260]
[0261] Where G represents the weight transformation matrix, G T Let G be the transpose of G, and g be the weight matrix before the transformation.
[0262] For example, the weight transformation matrix can be an existing winograd weight transformation matrix.
[0263] For example, if the input data matrix d before transformation is a 4×4 matrix and the weight matrix g before transformation is a 3×3 matrix, the weight transformation matrix can be:
[0264]
[0265] The transformed weight matrix can be obtained offline or online.
[0266] The transformed weight matrix is obtained offline, meaning it is obtained before the neural network model's computation. For example, the transformed weight matrix can be obtained before the neural network model is deployed. During the inference process of the neural network model, the weight matrix remains unchanged. Obtaining the transformed weight matrix offline can further improve computational speed and reduce computational overhead.
[0267] The transformed weight matrix is obtained through online transformation, meaning that the transformed weight matrix is obtained during the execution of the neural network model.
[0268] Specifically, each element in the intermediate matrix is the negative of the L1 distance between the corresponding element in the transformed input data matrix and the transformed weight matrix.
[0269] For example, the intermediate matrix X satisfies the following formula:
[0270] X = -|[GgG T ]-[B T dB]|;
[0271] In this embodiment of the application, the minus sign between the two terms in the above formula indicates element-wise subtraction. This indicates that the absolute value is calculated for each element in the matrix.
[0272] Step S720 can be achieved by performing subtraction, absolute value calculation, and inversion operations on the transformed input data matrix and the transformed weight matrix.
[0273] Specifically, step S720 includes steps S721 to S723.
[0274] S721, perform element-wise subtraction between the transformed input data matrix and the transformed weight matrix to obtain the difference matrix.
[0275] S722, calculate the absolute value of the difference matrix to obtain the absolute value matrix.
[0276] S723, calculate the opposite of the absolute value matrix to obtain the intermediate matrix.
[0277] It should be noted that addition and subtraction are essentially the same, and the result of subtraction can also be obtained through addition. For the sake of simplicity, in this embodiment, addition and subtraction are collectively referred to as addition.
[0278] For example, step S720 can be performed by Figure 6 The acceleration module 630 in the middle is executed.
[0279] It should be noted that a neural network model may include multiple feature extraction layers, and each feature extraction layer may include one or more feature extraction kernels, i.e., weight matrices. Accordingly, feature extraction processing may be performed once or multiple times in each feature extraction layer, and one or more of these feature extractions may employ method 700 to perform the corresponding computation process. In other words, performing method 700 on a feature extraction layer of a neural network model may include performing method 700 on a feature extraction kernel within that feature extraction layer.
[0280] The S730 uses the winograd algorithm to transform the intermediate matrix into an output data matrix.
[0281] Specifically, the intermediate matrix is transformed by the output transformation matrix using a winograd transformation to obtain the output data matrix.
[0282] For example, the output data matrix Y satisfies the following formula:
[0283] Y = A T XA;
[0284] Specifically, the output data matrix Y satisfies the following formula:
[0285] Y = A T [-|[GgG T ]-[B T dB]|]A;
[0286] Where A represents the output transformation matrix, A T Let A be the transpose of A.
[0287] For example, the output transformation matrix can be an existing winograd input transformation matrix.
[0288] For example, if the input data matrix d before transformation is a 4×4 matrix, and the weight matrix g before transformation is a 3×3 matrix, the output transformation matrix can be:
[0289]
[0290] The output data matrix can be part or all of the output feature map.
[0291] For example, step S730 can be performed by Figure 6 The acceleration module 630 in the middle is executed.
[0292] According to the scheme of this application embodiment, the dot product operation in winograd is replaced with addition operations such as calculating L1 distance, which reduces the amount of computation in the feature extraction process, improves the running speed of the model, and reduces the computational overhead.
[0293] The solution in this application embodiment can be understood as a scheme that merges the Winograd algorithm with AdderNet. Alternatively, it can be understood as a scheme that optimizes the adder layer of AdderNet using the Winograd algorithm. Specifically, from the perspective of AdderNet, this solution reduces the number of addition operations in the adder layer of AdderNet, thereby reducing the computational cost of the model; from the perspective of the Winograd algorithm, this solution replaces multiplication operations with addition operations, thereby reducing the computational cost of the model.
[0294] Optionally, the elements in the output transformation matrix can take any of the following values: 0, -1, 1.
[0295] In this way, the elements in the output transformation matrix can take the values 0, 1, or -1, which can reduce multiplication operations, further reduce the number of calculations, and help reduce the computational load of the model.
[0296] Optionally, when the dimension of the weight matrix before transformation is 3×3 and the dimension of the output data matrix Y is 2×2, the output transformation matrix can be:
[0297]
[0298] Optionally, at least one row of the output transformation matrix has elements that are the negatives of the elements at the corresponding positions in the first matrix. The elements of the other rows of the output transformation matrix are the same as the elements at the corresponding positions in the other rows of the first matrix. The first matrix can be:
[0299]
[0300] Where c0, c1, and c2 are any of the terms 0, -1, and 1, respectively.
[0301] In other words, any row of A' can be inverted without affecting the final result of the winograd transformation.
[0302] Where c0, c1, and c2 are any of the terms 0, -1, and 1, respectively.
[0303] For example, c0 is 0, c1 is -1, and c2 is 1. Or, c0 is -1, c1 is 1, and c2 is 0.
[0304] Optionally, when the dimension of the weight matrix before transformation is 3×3 and the dimension of the output data matrix Y is 2×2, the weight transformation matrix can be:
[0305]
[0306] The output transformation matrix and weight transformation matrix described above satisfy the general solution form of the Winograd algorithm, ensuring that the convolution calculation result obtained by the Winograd algorithm is the same as the result of the conventional convolution calculation. Therefore, even with the existing input transformation matrix, the above output transformation matrix and weight transformation matrix can still be applied to convolution calculations.
[0307] It should be noted that the transformation matrix in the Winograd algorithm can take various forms, as long as it satisfies the general form of Winograd, the convolution result obtained by the Winograd algorithm is guaranteed to be the same as the result of a regular convolution calculation. The output transformation matrix and weight transformation matrix mentioned above are only examples when using the input transformation matrix of the existing Winograd algorithm. When the input data matrix is adjusted, the output transformation matrix and weight transformation matrix can also take other forms.
[0308] However, imbalances in the elements of the output data matrix can lead to a slower decrease in the loss function during training, resulting in lower accuracy of the trained model. Furthermore, this imbalance can also affect subsequent processing of the output data matrix, such as batch normalization, thus impacting model performance.
[0309] The following example illustrates this, using a 4×4 intermediate matrix X and a 2×2 output data matrix Y.
[0310] The intermediate matrix X can be represented in the following form:
[0311]
[0312] Where x0, x1…x 16 This represents an element in the intermediate matrix X.
[0313] The output data matrix can be represented in the following form:
[0314]
[0315] Where y0, y1, y2 and y3 represent elements in the output data matrix Y.
[0316] From Y=A T XA yields the output transformation matrix from the existing Winograd algorithm, i.e.:
[0317]
[0318] The elements in Y satisfy the following formula:
[0319] y0=x0+x1+x2+x4+x5+x6+x8+x9+x 10 ;
[0320] y1 = x1 - x2 - x3 + x5 - x6 - x7 + x9 - x 10 -x 11 ;
[0321] y2=x4+x5+x6-x8-x9-x 10 -x 12 -x 13 -x 14 ;
[0322] y3 = x5 - x6 - x7 - x9 + x 10 +x 11 -x 13 +x 14 +x 15 ;
[0323] As can be seen from the above formula, the number of addition operations in the equations corresponding to each element in Y is different. Correspondingly, the number of subtraction operations in the equations corresponding to each element in Y is different. In other words, the number of positive and negative signs of the elements in the equations corresponding to each element in Y is different. For example, the equation corresponding to y0 includes 9 addition operations, while the equation corresponding to y1 includes 3 addition operations; the number of addition operations in the equations corresponding to y0 and y1 is different. The magnitudes of the elements in X are usually the same, and each element in X is always non-positive. Because the number of addition and subtraction operations in the equations corresponding to each element in Y is different, the magnitudes of the elements in the output data matrix Y are inconsistent, meaning there is an imbalance in the elements at different positions in the output data matrix, which affects and consequently impacts the performance of the model.
[0324] The number of addition and subtraction operations in the equations corresponding to each element in Y is determined by the sign of the elements in the output transformation matrix. In other words, the sign of the elements in the output transformation matrix affects the distribution of eigenvalues in the output data matrix.
[0325] Therefore, this application also provides an output transformation matrix that can improve the computation speed of the neural network model while avoiding performance degradation.
[0326] Optionally, the number of positive numbers in each column of the output transformation matrix is the same, and the number of negative numbers in each column is the same.
[0327] This ensures the balance of elements at each position in the output data matrix.
[0328] As mentioned earlier, the value of an element in the output transformation matrix can be any of 0, 1, or -1. In this case, the number of +1s in each column of the output transformation matrix is the same, and the number of -1s in each column is the same.
[0329] For example, c0 is 0, c1 is -1, and c2 is 1. In this case, the output transformation matrix can include any of the following:
[0330]
[0331] Here, A0, A1, A2, and A3 represent four types of output transformation matrices. It should be understood that these four output transformation matrices are only examples. Other forms of output transformation matrices can be obtained by interchangeing the rows of the four output transformation matrices, with the values of c0, c1, and c2 being 0, 1, and -1 respectively. These are not listed here.
[0332] When using the output transformation matrix A0, from Y = A T From XA, we can obtain that the elements in Y satisfy the following formula:
[0333] y0=x0-x1-x2-x4+x5+x6-x8+x9+x 10 ;
[0334] y1 = -x1 + x2 - x3 + x5 - x6 + x7 + x9 - x 10 +x 11 ;
[0335] y2=x4+x5+x6+x8-x9-x 10 -x 12 +x 13 +x 14 ;
[0336] y3=x5-x6+x7-x9+x 10 -x 11 +x 13 -x 14 +x 15 ;
[0337] As can be seen from the above formula, the number of addition operations in the equations corresponding to each element of Y is the same. Correspondingly, the number of subtraction operations in the equations corresponding to each element of Y is the same. In other words, the number of positive and negative signs in the elements of the intermediate matrix in the equations corresponding to each element of Y is the same. For example, the equation corresponding to y0 contains 5 positive signs, and the equation corresponding to y1 also contains 5 positive signs. The number of positive signs in the equations corresponding to y0 and y1 is the same, ensuring the balance of elements in each position of the output data matrix. It should be understood that A0 is used as an example here; A1, A2, and A3 can all achieve the same effect.
[0338] For example, c0 is 0, c1 is -1, and c2 is 1. In this case, the weight transformation matrix can include any of the following:
[0339]
[0340] G0, G1, G2, and G3 represent four types of weight transformation matrices. It should be understood that these four weight transformation matrices are only examples. Other forms of weight transformation matrices can be obtained by interchangeing the rows of these four weight transformation matrices, with the values of c0, c1, and c2 being 0, 1, and -1 respectively. These are not listed here.
[0341] The weight transformation matrices G0, G1, G2, and G3 have a one-to-one correspondence with the output transformation matrices A0, A1, A2, and A3. These corresponding weight transformation matrices and output transformation matrices are used together. For example, G0 and A0 are used together. When using these four types of weight transformation matrices and output transformation matrices, the input transformation matrix can be the existing input transformation matrix from winograd.
[0342] With the input data matrix using the existing Winograd, the output transformation matrix and weight transformation matrix can take the four forms mentioned above. With adjustments to the input data matrix, the output transformation matrix and weight transformation matrix can also take other forms. In other words, as long as the transformation matrix satisfies the general solution form of Winograd and ensures the balance of elements at each position in the output data matrix, it is acceptable.
[0343] According to the scheme of this application embodiment, the number of positive numbers in each column of the output transformation matrix is the same, and the number of negative numbers in each column is the same. For example, the number of +1s in each column of the output transformation matrix is the same, and the number of -1s in each column is the same. This can balance the magnitude of each position in the output data matrix, that is, reduce the imbalance of eigenvalue accumulation, which is beneficial to model training. In addition, it is beneficial to perform subsequent processing on the output data matrix, such as batch normalization.
[0344] Furthermore, the output transformation matrix and the weight transformation matrix satisfy the general solution form of Winograd, which ensures that the output result is consistent with the actual result of the convolution operation, thus making it suitable for convolution operations in the model.
[0345] As mentioned earlier, using the Winograd algorithm to accelerate convolution operations does not affect the computation results. However, Method 700 uses an absolute value operation, which renders the distributive law of multiplication inapplicable. This means that if the Winograd algorithm is used to accelerate the adder layer, the computation result will differ somewhat from the original result. Combining the Winograd algorithm with the adder layer in the above manner will lead to a performance degradation of the neural network model.
[0346] This application also provides a method for training a neural network model, which can improve the performance of the neural network model.
[0347] The following is combined with Figure 8 The training method of the neural network model in the embodiments of this application is described in detail.
[0348] Figure 8 The present application illustrates a method 800 for training a neural network model according to an embodiment of the present application. Figure 8 The method shown can be executed by a training device for a neural network model. This device can be a cloud service device or a terminal device, such as a computer, server, or other device with sufficient computing power to perform neural network model training. It can also be a system composed of a cloud service device and a terminal device. For example, method 800 can be executed by... Figure 2 Training equipment 120 Figure 4 The neural network processor 50 or Figure 5 The execution device 310 in the middle is used for execution.
[0349] Figure 8 The training method 800 uses computational methods in the forward propagation process of executing a neural network model. Figure 7The method is the same as in Method 700. Simply replace "data to be processed" with "training data" in Method 700. The specific implementation of the forward propagation process in Method 800 can be referred to the aforementioned Method 700. To avoid unnecessary repetition, repeated descriptions will be omitted appropriately when introducing Method 800 below.
[0350] Method 800 includes steps S810 to S850. Steps S810 to S850 are described in detail below.
[0351] Perform the following operations on at least one feature extraction layer in a neural network model:
[0352] S810 uses the winograd algorithm to transform the input data matrix of the training data to obtain the transformed input data matrix.
[0353] The type of training data depends on the task of the neural network model. For example, if the neural network model is used for image processing tasks, the training data can be images. Specifically, image processing tasks include image classification, image detection, image segmentation, or image generation. Similarly, if the neural network model is used for text processing tasks, the training data can be text. Specifically, text processing tasks include text recognition or text translation. Furthermore, if the neural network model is used for speech processing tasks, the training data can be speech data. Specifically, speech processing tasks include speech recognition. This application does not limit the type of training data.
[0354] For example, the training data may be pre-stored. For instance, the training data could be... Figure 2 The training data maintained in the database 130 shown.
[0355] The input data matrix refers to the data matrix input to at least one feature extraction layer.
[0356] S820: Feature extraction is performed on the transformed input data matrix using the transformed weight matrix to obtain an intermediate matrix. The transformed weight matrix is obtained by applying a weight transformation to the weight matrix of at least one feature extraction layer using the Winograd algorithm. Each element in the intermediate matrix is determined based on the Lp distance between the corresponding element in the transformed input data matrix and the transformed weight matrix.
[0357] Lp distance, also known as Minkowski distance, is a set of distance definitions. p is a parameter.
[0358] When p is 1, each element in the intermediate matrix is determined based on the L1 distance between the corresponding element in the transformed input data matrix and the transformed weight matrix.
[0359] For example, the intermediate matrix X satisfies the following formula:
[0360] X = -|[GgG T ]-[B T dB]|;
[0361] When p is 2, each element in the intermediate matrix is determined based on the L2 distance between the corresponding element in the transformed input data matrix and the transformed weight matrix.
[0362] For example, the intermediate matrix X satisfies the following formula:
[0363] X = -([GgG) T ]-[B T dB]) 2 ;
[0364] L2 distance can also be called Euclidean distance.
[0365] The S830 uses the winograd algorithm to transform the intermediate matrix into an output data matrix.
[0366] S840 determines the value of the loss function based on the output data matrix.
[0367] Specifically, the output data matrix is part or all of the output feature map. The processing result of the training data can be determined based on the output feature map, and the value of the loss function can be calculated based on the processing result of the training data.
[0368] The results of training data processing depend on the type of training data and the task of the neural network model.
[0369] For example, the training data is image data, and image processing may include image super-resolution processing, image denoising processing, image recognition processing, etc. Correspondingly, the image processing results may include image super-resolution, image denoising, or image classification, etc. This application embodiment does not limit this.
[0370] For example, the training data is speech data, and speech processing may include speech recognition, etc. Accordingly, the speech processing results include speech recognition results, etc. This application embodiment does not limit this.
[0371] The output feature map can be further processed, such as by activation functions, to obtain the processing results of the training data.
[0372] S850 trains the neural network model based on the value of the loss function.
[0373] During the m-th iteration of training the neural network model, p is 2; during the n-th iteration of training the neural network model, p is 1, where m and n are positive integers, and m is less than n.
[0374] In the m-th iteration, the intermediate matrix is calculated using L2 distance during the forward computation, and the partial derivatives of the loss function with respect to the weights in the first weight matrix are calculated based on L2 distance during the backpropagation. Specifically, forward computation is performed based on the L2 distance between corresponding elements in the transformed input data matrix and the transformed weight matrix, and backpropagation is performed based on the L2 distance between corresponding elements in the transformed input data matrix and the transformed weight matrix. Alternatively, it can be understood as performing forward and backward computations based on L2 distance.
[0375] During the nth iteration, the intermediate matrix is calculated using L1 distance during the forward computation, and the partial derivatives of the loss function with respect to the weights in the first weight matrix are calculated based on the L1 distance during the backpropagation. Specifically, forward computation is performed based on the L1 distance between corresponding elements in the transformed input data matrix and the transformed weight matrix, and backpropagation is performed based on the L1 distance between corresponding elements in the transformed input data matrix and the transformed weight matrix. Alternatively, it can be understood as performing both forward and backpropagation based on the L1 distance.
[0376] The first weight matrix includes either the weight matrix before transformation or the weight matrix after transformation. That is, the first weight matrix can be either the weight matrix g before transformation or the weight matrix g after transformation. During training, the gradient of the weight matrix before transformation can be calculated, and the values of the weight matrix before transformation can be adjusted accordingly. Alternatively, the gradient of the weight matrix after transformation can be calculated, and the values of the weight matrix after transformation can be adjusted accordingly.
[0377] In other words, in the early stages of neural network model training, forward computation and backpropagation are performed based on L2 distance, while in the later stages of neural network model training, forward computation and backpropagation are performed using L1 distance, that is, L2 distance is used to approximate L1 distance.
[0378] Training solely based on L1 distance can hinder Winograd algorithm optimization and potentially prevent network training from converging. This embodiment utilizes L2 distance to assist training in the early stages, as L2 distance is more favorable to the Winograd algorithm, improving convergence speed and ultimately enhancing model performance. In the later stages, training is based on L1 distance to further improve the training effect of the model using L1 distance. Using L1 distance in the trained model is also more hardware-friendly.
[0379] Optionally, the partial derivatives of the loss function with respect to the weights in the first weight matrix satisfy the following formula:
[0380]
[0381] Where p is the calculated norm, p∈[1,2], w represents the weights in the first weight matrix, which can be either the weights in the original weight matrix g or the weights in the transformed weight matrix, x represents the data in the original input data matrix or the data in the transformed input data matrix, where the data in the original input data matrix is the eigenvalue of the input feature map, and L represents the loss function.
[0382] i represents the number of layers in the neural network model, and i is an integer. This represents the partial derivative of the loss function with respect to the eigenvalues of the i-th layer. This represents the partial derivative of the loss function with respect to the eigenvalues of the (i+1)th layer. represents the partial derivative of the loss function with respect to the weights in the first weight matrix of the i-th layer, and sign() represents the sign function.
[0383] When p = 2, the above equation represents the partial derivative of the loss function obtained based on the L2 distance with respect to the first weight matrix, meaning that backpropagation is performed based on the L2 distance. When p = 1, the above equation represents the partial derivative of the loss function obtained based on the L1 distance with respect to the first weight matrix, meaning that backpropagation is performed based on the L2 distance.
[0384] Optionally, the value of p is determined based on the number of iterations in the training process.
[0385] Furthermore, the initial value of p is 2, and the value of p decreases as the number of iterations increases.
[0386] In other words, p is reduced from 2 to 1 during training.
[0387] The value of p can decrease once in each iteration, or it can decrease once every few iterations.
[0388] For example, during training, the value of p decreases by 'a' in each iteration. The value of 'a' can be set as needed. 'a' can be fixed, meaning the amount of decrease in p remains constant in each iteration. For example, 'a' can be 0.05, meaning p decreases by 0.05 in each iteration. Alternatively, 'a' can be variable. For example, the amount of decrease in p gradually increases with the number of iterations. For instance, in the first iteration, p is 2; in the second iteration, 'a' is 0.01, so p is 1.99; and in the third iteration, 'a' is 0.02, so p is 1.98.
[0389] Alternatively, during training, the value of p decreases by a every k iterations.
[0390] Alternatively, train using L2 distance until convergence, reduce the value of p, and retrain.
[0391] It should be noted that the above is only an example, and other methods can be used to reduce p from 2 to 1. This application does not limit this.
[0392] The trained neural network model can be used to perform a target task. For example, the target task can be an image processing task, such as object detection, image segmentation, instance segmentation, image denoising, image super-resolution, etc. Alternatively, the target task can be a speech processing task, such as speech recognition, etc. Or, the target task can be a text processing task, such as text recognition or text translation, etc.
[0393] Table 1 shows a comparison of the experimental results of the proposed solution and existing solutions on CIFAR-10 classification data.
[0394] Table 1
[0395] method accuracy AdderNet 91.84 The computational method used in this application (employing existing transformation matrices) 86.13 The computational method used in this application (employing an adjusted transformation matrix) 88.60 The model obtained using the training method of this application 91.47
[0396] As shown in Table 1, while the computational method of this application improves computational efficiency when using existing transformation matrices, it leads to a decrease in model accuracy. Compared to using existing transformation matrices, using the adjusted transformation matrix improves model accuracy. Furthermore, compared to using existing transformation matrices, training the model using the adjusted transformation matrix further improves model accuracy, approaching the accuracy of AdderNet.
[0397] Table 2 shows a comparison of the experimental results of the proposed solution and existing solutions on low-level vision tasks.
[0398] Table 2
[0399] method PSNR Convolutional Neural Networks 57.31 AdderNet 57.22 The computational method used in this application (employing an adjusted transformation matrix) 57.27
[0400] As shown in Table 2, the computation method of this application embodiment can achieve higher performance metrics than AdderNet. Furthermore, the computation method of this application embodiment can achieve visual effects close to those of AdderNet.
[0401] The following is combined with Figures 9 to 12 The apparatus of the embodiments of this application will be described below. It should be understood that the apparatus described below is capable of performing the methods of the foregoing embodiments of this application. To avoid unnecessary repetition, repeated descriptions will be appropriately omitted when describing the apparatus of the embodiments of this application below.
[0402] Figure 9This is a schematic block diagram of a training apparatus for a neural network model according to an embodiment of this application. Figure 9 The training device 3000 for the neural network model shown includes an acquisition unit 3010 and a processing unit 3020.
[0403] The acquisition unit 3010 and the processing unit 3020 can be used to execute the training method of the neural network model of the present application embodiment, specifically, they can be used to execute method 800.
[0404] The acquisition unit 3010 is used to acquire training data.
[0405] The processing unit 3020 performs the following operations on at least one feature extraction layer of the neural network model:
[0406] The input data matrix of the training data is transformed using the input transformation matrix of the Winograd algorithm to obtain the transformed input data matrix. The transformed weight matrix is then used to extract features from the transformed input data matrix to obtain an intermediate matrix. The transformed weight matrix is obtained by transforming the weight matrices of at least one feature extraction layer using the weight transformation matrix of the Winograd algorithm. Each element in the intermediate matrix is determined based on the Lp distance between corresponding elements in the transformed input data matrix and the transformed weight matrix. The intermediate matrix is then transformed using the output transformation matrix of the Winograd algorithm to obtain the output data matrix. The value of the loss function is determined based on the output data matrix. The neural network model is then trained based on the value of the loss function. In the m-th iteration of training the neural network model, p is 2; in the n-th iteration, p is 1; m and n are positive integers, with m less than n.
[0407] Alternatively, as an example, during the training of the neural network model, the initial value of p is 2, and the value of p decreases as the number of iterations increases.
[0408] Optionally, as an embodiment, training the neural network model based on the value of the loss function includes: adjusting the weights in the first weight matrix according to the partial derivative of the loss function with respect to the weights in the first weight matrix, wherein the first weight matrix includes the weight matrix before transformation or the weight matrix after transformation.
[0409] Optionally, as an example, the partial derivatives of the loss function with respect to the weights in the first weight matrix satisfy the following formula:
[0410]
[0411] Where p is the norm of the calculation, p∈[1,2], w represents the weight, x represents the data in the input data matrix before or after the transformation, L represents the loss function, i represents the number of layers in the neural network model, and sign() represents the sign function.
[0412] Figure 10 This is a schematic block diagram of the computing device 4000 for the neural network model provided in this application embodiment. Figure 10 The apparatus 4000 shown includes an acquisition unit 4010 and a processing unit 4020.
[0413] The acquisition unit 4010 and the processing unit 4020 can be used to execute the operation method of the neural network model of the present application embodiment, for example, they can be used to execute method 700.
[0414] The acquisition unit 4010 is used to acquire data to be processed, including image data, voice data, or text data.
[0415] The processing unit 4020 performs the following operations on at least one feature extraction layer of the neural network model: transforms the input data matrix of the data to be processed using the input transformation matrix of the Winograd algorithm to obtain a transformed input data matrix; extracts features from the transformed input data matrix using the transformed weight matrix to obtain an intermediate matrix, wherein the transformed weight matrix is obtained by transforming the weight matrix of at least one feature extraction layer using the weight transformation matrix of the Winograd algorithm, and each element in the intermediate matrix is determined based on the L1 distance between the corresponding elements in the transformed input data matrix and the transformed weight matrix; and transforms the intermediate matrix using the output transformation matrix of the Winograd algorithm to obtain an output data matrix.
[0416] Alternatively, as an example, the output data matrix satisfies the following formula:
[0417] Y = A T [-|[GgG T ]-[B T dB]|]A;
[0418] Where Y represents the output data matrix, A represents the output transformation matrix, and A T Let G denote the transpose of A, and let G denote the weight transformation matrix. T Let G be the transpose of the matrix, g be the weight matrix before the transformation, and B be the input transformation matrix. T Let d denote the transpose of B, and let d denote the input data matrix before the transformation.
[0419] Alternatively, as an example, the values of the elements in the output transformation matrix are any of 0, -1, or 1.
[0420] Optionally, as an example, the output transformation matrix is:
[0421]
[0422] Where c0, c1, and c2 are any of the terms 0, -1, and 1, respectively.
[0423] Optionally, as an embodiment, the elements of at least one row in the output transformation matrix are the negatives of the elements at the corresponding positions in the first matrix, and the elements of the other rows in the output transformation matrix are the same as the elements at the corresponding positions in the other rows of the first matrix. The first matrix is:
[0424]
[0425] Where A' represents the first matrix, and c0, c1, and c2 are any one of 0, -1, and 1, respectively.
[0426] Optionally, as an example, the weight transformation matrix is:
[0427]
[0428] Where c0, c1, and c2 are any of the terms 0, -1, and 1, respectively.
[0429] Alternatively, as an example, the number of positive numbers in each column of the output transformation matrix is the same, and the number of negative numbers in each column is the same.
[0430] It should be noted that the training device 3000 and device 4000 mentioned above are embodied in the form of functional units. The term "unit" here can be implemented in software and / or hardware, and there is no specific limitation on this.
[0431] For example, a "unit" can be a software program, a hardware circuit, or a combination of both that implements the above functions. The hardware circuit may include an application-specific integrated circuit (ASIC), electronic circuitry, a processor (e.g., a shared processor, a proprietary processor, or a group processor) and memory for executing one or more software or firmware programs, integrated logic circuitry, and / or other suitable components that support the described functions.
[0432] Therefore, the units of the various examples described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0433] Figure 11 This is a schematic diagram of the hardware structure of the training device for the neural network model provided in this application embodiment. Figure 11 The training device 5000 for the neural network model shown (specifically, the device 5000 can be a computer device) includes a memory 5001, a processor 5002, a communication interface 5003, and a bus 5004. The memory 5001, processor 5002, and communication interface 5003 are interconnected via the bus 5004.
[0434] The memory 5001 can be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 5001 can store programs. When the program stored in the memory 5001 is executed by the processor 5002, the processor 5002 executes the various steps of the neural network model training method of this application embodiment. Specifically, the processor 5002 can execute the steps described above... Figure 8 Method 800 is shown.
[0435] The processor 5002 may be a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), graphics processing unit (GPU), or one or more integrated circuits, used to execute related programs to implement the neural network model training method of the method embodiment of this application.
[0436] The processor 5002 can also be an integrated circuit chip with signal processing capabilities; for example, it could be... Figure 4 The chip shown. In the implementation process, each step of the training method for the neural network model of this application can be completed by the integrated logic circuit of the hardware in the processor 5002 or by software instructions.
[0437] The processor 5002 described above can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory 5001, and the processor 5002 reads the information in memory 5001 and combines it with its hardware to complete the task. Figure 9 The training apparatus shown includes units that are required to perform functions, or to perform the methods described in this application. Figure 8 The training method for the neural network model shown.
[0438] The communication interface 5003 uses a transceiver device, such as, but not limited to, a transceiver, to enable communication between the device 5000 and other devices or communication networks. For example, training data can be acquired through the communication interface 5003.
[0439] Bus 5004 may include a pathway for transmitting information between various components of device 5000 (e.g., memory 5001, processor 5002, communication interface 5003).
[0440] Figure 12 This is a schematic diagram of the hardware structure of the computing device for the neural network model in an embodiment of this application. Figure 12 The data processing device 6000 shown includes a memory 6001, a processor 6002, a communication interface 6003, and a bus 6004. The memory 6001, processor 6002, and communication interface 6003 are interconnected via the bus 6004.
[0441] The memory 6001 can be a ROM, a static storage device, or RAM. The memory 6001 can store programs. When the program stored in the memory 6001 is executed by the processor 6002, the processor 6002 and the communication interface 6003 are used to execute the various steps of the operation method of the neural network model in this embodiment. Specifically, the processor 6002 can execute the steps described above... Figure 7 Steps S710 to S730 in the method shown.
[0442] The processor 6002 may be a general-purpose CPU, microprocessor, ASIC, GPU, or one or more integrated circuits, used to execute relevant programs to implement the functions required by the units in the computing device of the neural network model of the present application embodiment, or to execute the computing method of the neural network model of the method embodiment of the present application.
[0443] The processor 6002 can also be an integrated circuit chip with signal processing capabilities; for example, it could be... Figure 4 The chip shown. In the implementation process, each step of the operation method of the neural network model in this application embodiment can be completed by the integrated logic circuit of the hardware in the processor 6002 or by software instructions.
[0444] The processor 6002 described above can also be a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory 6001. The processor 6002 reads the information in memory 6001 and, in conjunction with its hardware, completes the functions required by the units included in the neural network model computing device of the embodiments of this application, or executes the computing method of the neural network model of the method embodiments of this application.
[0445] The communication interface 6003 uses transceiver devices, such as, but not limited to, transceivers, to enable communication between the device 6000 and other devices or communication networks. For example, data to be processed can be obtained through the communication interface 6003.
[0446] Bus 6004 may include a pathway for transmitting information between various components of device 6000 (e.g., memory 6001, processor 6002, communication interface 6003).
[0447] It should be noted that although only the memory, processor, and communication interface are shown in the above-described devices 5000 and 6000, those skilled in the art should understand that in specific implementations, devices 5000 and 6000 may also include other devices necessary for normal operation. Furthermore, depending on specific needs, those skilled in the art should understand that devices 5000 and 6000 may also include hardware devices for implementing other additional functions. In addition, those skilled in the art should understand that devices 5000 and 6000 may only include the devices necessary for implementing the embodiments of this application, and may not necessarily include... Figure 11 and Figure 12 All the devices shown.
[0448] It should be understood that the processor in the embodiments of this application can be a central processing unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0449] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0450] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0451] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0452] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0453] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0454] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0455] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0456] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0457] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0458] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0459] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0460] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A computational method for a neural network model, characterized in that, Perform the following operations in at least one feature extraction layer of the neural network model: The input data matrix of the data to be processed is transformed by the input transformation matrix of the Winograd algorithm to obtain the transformed input data matrix. The data to be processed includes image data, voice data, or text data. The transformed input data matrix is used to extract features to obtain an intermediate matrix. The transformed weight matrix is obtained by transforming the weight matrix of the at least one feature extraction layer using the weight transformation matrix of the winograd algorithm. Each element in the intermediate matrix is determined based on the L1 distance between the corresponding element in the transformed input data matrix and the transformed weight matrix. The intermediate matrix is transformed using the output transformation matrix of the winograd algorithm to obtain the output data matrix.
2. The calculation method according to claim 1, characterized in that, The output data matrix satisfies the following formula: Y=A T [-|[GgG T ]-[B T dB]|]A; Where Y represents the output data matrix, A represents the output transformation matrix, and A T Let G denote the transpose of A, and let G denote the weight transformation matrix. T Let g denote the transpose of G, g denote the weight matrix before the transformation, and B denote the input transformation matrix. T Let d represent the transpose of B, and let d represent the input data matrix before the transformation.
3. The calculation method according to claim 1 or 2, characterized in that, The value of each element in the output transformation matrix is any one of 0, -1, or 1.
4. The calculation method according to claim 3, characterized in that, The output transformation matrix is: Where c0, c1, and c2 are any of the terms 0, -1, and 1, respectively.
5. The calculation method according to claim 3, characterized in that, At least one row of the output transformation matrix contains elements that are the negatives of the elements in the first matrix corresponding to the corresponding rows. The elements in the other rows of the output transformation matrix are the same as the elements in the first matrix corresponding to the corresponding rows. The first matrix is: Where A' represents the first matrix, and c0, c1 and c2 are any one of 0, -1 and 1 respectively.
6. The calculation method according to claim 1 or 2, characterized in that, The weight transformation matrix is: Where c0, c1, and c2 are any of the terms 0, -1, and 1, respectively.
7. The calculation method according to claim 1 or 2, characterized in that, The number of positive numbers in each column of the output transformation matrix is the same, and the number of negative numbers in each column is the same.
8. A method for training a neural network model, characterized in that, include: The input data matrix of the training data is transformed by the input transformation matrix of the winograd algorithm to obtain the transformed input data matrix; The transformed input data matrix is used to extract features to obtain an intermediate matrix. The transformed weight matrix is obtained by transforming the weight matrix of the at least one feature extraction layer using the weight transformation matrix of the winograd algorithm. Each element in the intermediate matrix is determined based on the Lp distance between the corresponding element in the transformed input data matrix and the transformed weight matrix. The intermediate matrix is transformed using the output transformation matrix of the winograd algorithm to obtain the output data matrix. The value of the loss function is determined based on the output data matrix; The neural network model is trained based on the value of the loss function; During the m-th iteration of training the neural network model, p is 2, and during the n-th iteration of training the neural network model, p is 1. m and n are positive integers, and m is less than n.
9. The training method according to claim 8, characterized in that, During the training of the neural network model, the initial value of p is 2, and the value of p decreases as the number of iterations increases.
10. The training method according to claim 8 or 9, characterized in that, Training the neural network model based on the value of the loss function includes: The weights in the first weight matrix are adjusted according to the partial derivative of the loss function with respect to the weights in the first weight matrix. The first weight matrix includes either the weight matrix before the transformation or the weight matrix after the transformation.
11. The training method according to claim 10, characterized in that, The partial derivatives of the loss function with respect to the weights in the first weight matrix satisfy the following formula: Where p∈[1,2], w represents the weight, x represents the data in the input data matrix before transformation or the data in the input data matrix after transformation, L represents the loss function, i represents the number of layers in the neural network model, and sign() represents the sign function.
12. A computing device for a neural network model, characterized in that, include: An acquisition unit is used to acquire data to be processed, which includes image data, voice data, or text data. A processing unit is configured to perform the following operations on at least one feature extraction layer of a neural network model: The input data matrix of the Winograd algorithm is used to transform the input data matrix of the data to be processed to obtain the transformed input data matrix. The transformed input data matrix is used to extract features to obtain an intermediate matrix. The transformed weight matrix is obtained by transforming the weight matrix of the at least one feature extraction layer using the weight transformation matrix of the winograd algorithm. Each element in the intermediate matrix is determined based on the L1 distance between the corresponding element in the transformed input data matrix and the transformed weight matrix. The intermediate matrix is transformed using the output transformation matrix of the winograd algorithm to obtain the output data matrix.
13. The computing device according to claim 12, characterized in that, The output data matrix satisfies the following formula: Y=A T [-|[GgG T ]-[B T dB]|]A; Where Y represents the output data matrix, A represents the output transformation matrix, and A T Let G denote the transpose of A, and let G denote the weight transformation matrix. T Let g denote the transpose of G, g denote the weight matrix before the transformation, and B denote the input transformation matrix. T Let d represent the transpose of B, and let d represent the input data matrix before the transformation.
14. The computing device according to claim 12 or 13, characterized in that, The value of each element in the output transformation matrix is any one of 0, -1, or 1.
15. The computing device according to claim 14, characterized in that, The output transformation matrix is: Where c0, c1, and c2 are any of the terms 0, -1, and 1, respectively.
16. The computing device according to claim 15, characterized in that, At least one row of the output transformation matrix contains elements that are the negatives of the elements in the first matrix corresponding to the corresponding rows. The elements in the other rows of the output transformation matrix are the same as the elements in the first matrix corresponding to the corresponding rows. The first matrix is: Where A' represents the first matrix, and c0, c1 and c2 are any one of 0, -1 and 1 respectively.
17. The computing device according to claim 12 or 13, characterized in that, The weight transformation matrix is: Where c0, c1, and c2 are any of the terms 0, -1, and 1, respectively.
18. The computing device according to claim 12 or 13, characterized in that, The number of positive numbers in each column of the output transformation matrix is the same, and the number of negative numbers in each column is the same.
19. A training device for a neural network model, characterized in that, include: The acquisition unit is used to acquire training data; Processing unit, used for: The input data matrix of the training data is transformed by the input transformation matrix of the winograd algorithm to obtain the transformed input data matrix; The transformed input data matrix is used to extract features to obtain an intermediate matrix. The transformed weight matrix is obtained by transforming the weight matrix of the at least one feature extraction layer using the weight transformation matrix of the winograd algorithm. Each element in the intermediate matrix is determined based on the Lp distance between the corresponding element in the transformed input data matrix and the transformed weight matrix. The intermediate matrix is transformed using the output transformation matrix of the winograd algorithm to obtain the output data matrix. The value of the loss function is determined based on the output data matrix; The neural network model is trained based on the value of the loss function; During the m-th iteration of training the neural network model, p is 2, and during the n-th iteration of training the neural network model, p is 1. m and n are positive integers, and m is less than n.
20. The training device according to claim 19, characterized in that, During the training of the neural network model, the initial value of p is 2, and the value of p decreases as the number of iterations increases.
21. The training device according to claim 19 or 20, characterized in that, Training the neural network model based on the value of the loss function includes: adjusting the weights in the first weight matrix according to the partial derivative of the loss function with respect to the weights in the first weight matrix, wherein the first weight matrix includes the weight matrix before transformation or the weight matrix after transformation.
22. The training device according to claim 21, characterized in that, The partial derivatives of the loss function with respect to the weights in the first weight matrix satisfy the following formula: Where p∈[1,2], w represents the weight, x represents the data in the input data matrix before transformation or the data in the input data matrix after transformation, L represents the loss function, i represents the number of layers in the neural network model, and sign() represents the sign function.
23. A computing device for a neural network model, characterized in that, It includes a processor and a memory, the memory being used to store program instructions, and the processor being used to invoke the program instructions to perform the method as described in any one of claims 1 to 7.
24. A training device for a neural network model, characterized in that, It includes a processor and a memory, the memory being used to store program instructions, and the processor being used to invoke the program instructions to perform the method as described in any one of claims 8 to 11.
25. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program code that can be executed by a device, the program code including methods for performing any one of claims 1 to 7 or claims 8 to 11.
26. A computer program product containing instructions, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as claimed in any one of claims 1 to 7 or 8 to 11.
27. A chip, characterized in that, The chip includes a processor and a data interface, wherein the processor reads instructions stored in a memory through the data interface to execute the method as claimed in any one of claims 1 to 7 or claims 8 to 11.
Citation Information
Patent Citations
Convolutional neural network data processing method and device based on winograd convolution operation
CN110097172A
Neural network processor using dyadic weight matrix and operation method thereof
US20200167637A1