Method, Medium, and Electronic Device for Operating a Neural Network Model
By splicing the weight matrix of the same operation form in the neural network model, the problem of high computing speed and power consumption in the prior art is solved, and more efficient computing speed and lower power consumption are achieved.
Patent Information
- Application Number
- CN202210330475.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-03-30
AI Technical Summary
The existing neural network models have bottlenecks in terms of computing speed and power consumption, especially in edge devices with limited computing resources, which are difficult to effectively improve recognition and reduce computing time.
By splicing the weight matrix in the operation terms with the same operation form and the same input data, using the spliced weight matrix for matrix operation, the input data can be obtained from the storage unit at one time, and the calculation results of each operation term can be obtained.
The speed of the electronic device's computing neural network model is improved, the number of times the computing unit reads data from the storage unit is reduced, power consumption is reduced, and the processing capability of edge devices is improved.
Smart Images

Figure CN114676832B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of machine learning, and particularly to a method for running a neural network model, a medium, and an electronic device. Background Art
[0002] With the rapid development of Artificial Intelligence (AI) technology, neural networks (e.g., deep neural networks, recurrent neural networks) have achieved very good results in the fields of computer vision, speech, natural language, reinforcement learning, etc. in recent years. With the development of neural network algorithms, the complexity of the algorithms is getting higher and higher. In order to improve the recognition rate, the scale of the model is gradually increasing. Correspondingly, the power consumption and the consumption of computing resources of the device deploying the neural network model are also increasing. Especially for some edge devices with limited computing resources, it is particularly important to improve the computing speed of the neural network model, save computing time, and reduce power consumption. Summary of the Invention
[0003] The purpose of this application is to provide a method for running a neural network model, a medium, and an electronic device. Through the method for running a neural network model of this application, by concatenating the operation terms with the same operation form and the same input data (such as "x t ×W xz +h t-1 ×W hz " in formula (1) below and "x t ×W xr +h t-1 ×W hr " with the same operation form and the input data both being x t and h t-1 ), the weight matrices (W xz and W xr as well as W hz and W hr ) are concatenated, and then the operation results of each operation term can be obtained by using the operation unit corresponding to the operation term of this form based on the concatenated weight matrix through one-time obtaining of the input data from the storage unit. Thus, the computing speed of the electronic device for running the neural network model is improved.
[0004] The first aspect of the present application provides a method for running a neural network model, which is applied to an electronic device. The neural network model includes a first operation and a second operation. The method includes: obtaining a data matrix to be operated on for the first operation or the second operation, where the number of operation factors in the first operation and the second operation is the same, and for each operation factor in the first operation, there is a corresponding operation factor in the second operation with the same data matrix to be operated on but different operation coefficient matrices. Performing a third operation on the data matrix to be operated on to generate a result matrix of the third operation, where the third operation is an operation method obtained by combining the operation coefficient matrices of the corresponding operation factors in the first operation and the second operation. Splitting the result matrix of the third operation to obtain the result matrix of the first operation and the result matrix of the second operation respectively.
[0005] For example, taking the neural network model as a gated recurrent unit model, in the GRU network at the t-th time step of the gated recurrent unit model, the first operation can be x in formula (1) below t ×W xz +h t-1 ×W hz , and the second operation can be x in formula (2) below t ×W xz +h t-1 ×W hz . For example, the arithmetic unit of the processor reads the input data x at the t-th time step t , weight coefficient matrix W xz , weight coefficient matrix W hz , weight coefficient matrix W xr , weight coefficient matrix W hr , and the output data h at the (t - 1)-th time step from the storage unit of the processor once t-1 , splices the aforementioned weight coefficient matrix W xz and weight coefficient matrix W xr to obtain a spliced weight matrix W xzr , splices the aforementioned weight coefficient matrix W hz and weight coefficient matrix W hr to obtain a spliced weight matrix W hzr . The arithmetic unit performs a matrix operation logic of X×H1 + Y×H2, takes the spliced weight matrix W xzr as H1, the spliced weight matrix W hzr as H2, the input data x at the t-th time step t as X, and the output data h at the (t - 1)-th time step t-1 as Y, and performs the corresponding matrix operation to obtain x in formula (1) in one operation t ×W xz +h t-1×W hz The operation result and x t ×W xr +h t-1 ×W hr The operation result, and then according to the weight coefficient matrix W xz and the weight coefficient matrix W xr of the dimension value, x t ×W xz +h t-1 ×W hz The operation result and x t ×W xr +h t-1 ×W hr The operation result is split to separately obtain the operation results of x t ×W xz +h t-1 ×W hz The operation result, x t ×W xr +h t-1 ×W hr The operation result. By obtaining the input data from the storage unit once, the operation results of each operation item can be obtained. Thus, the operation speed of the neural network model by the electronic device is improved.
[0006] In a possible implementation of the first aspect above, the operation coefficient matrices of the first operation and the second operation include two data dimensions of height and width. Among them, the height and width of the operation coefficient matrices of the corresponding operation factors in the first operation and the second operation are correspondingly equal. The third operation is an operation method obtained by merging the operation coefficient matrices of the corresponding operation factors in the first operation and the second operation along any data dimension direction.
[0007] In a possible implementation of the first aspect above, the third operation is an operation method obtained by merging the operation coefficient matrices of the corresponding operation factors in the first operation and the second operation along the width direction.
[0008] In a possible implementation of the first aspect above, splitting the result matrix of the third operation to separately obtain the result matrix of the first operation and the result matrix of the second operation includes: The result matrix of the third operation includes two data dimensions of height and width. Splitting the result matrix of the third operation along any data dimension direction separately obtains the result matrix of the first operation and the result matrix of the second operation.
[0009] In a possible implementation of the first aspect above, splitting the result matrix of the third operation along the width direction separately obtains the result matrix of the first operation and the result matrix of the second operation.
[0010] In a possible implementation of the first aspect described above, the third operation, which is an operation mode obtained after combining the operation coefficient matrices of the corresponding operation factors in the first operation and the second operation, includes: the data matrix to be operated on includes a first input data matrix and a second input data matrix, and the operation factors of the first operation and the second operation include the matrix products of the first input data matrices of the first operation and the second operation with the corresponding operation coefficient matrices, and the matrix products of the second input data matrices of the first operation and the second operation with the corresponding operation coefficient matrices.
[0011] The third operation is an operation mode obtained after combining the operation coefficient matrices corresponding to the first input data matrices of the first operation and the second operation, and combining the operation coefficient matrices corresponding to the second input data matrices of the first operation and the second operation.
[0012] In a possible implementation of the first aspect described above, the neural network model is a recurrent neural network model, and the first operation and the second operation are operations of the fully connected layer of the recurrent neural network model.
[0013] In a possible implementation of the first aspect described above, the neural network model includes at least one of the following: gated recurrent unit model, long short-term memory model.
[0014] The second aspect of the present application provides an electronic device, which includes: an operation unit and a storage unit. Among them, the operation unit runs a first matrix operation circuit. The operation unit obtains the data matrix to be operated on in the first operation or the second operation from the storage unit. Among them, the number of operation factors in the first operation and the second operation is the same, and for each operation factor in the first operation, there is a corresponding operation factor in the second operation with the same data matrix to be operated on but different operation coefficient matrices. The operation unit performs a third operation on the data matrix to be operated on by running the first matrix operation circuit to generate a result matrix of the third operation, where the third operation is an operation mode obtained after combining the operation coefficient matrices of the corresponding operation factors in the first operation and the second operation, and the first matrix operation circuit is used to generate a first operation result matrix and a second operation result matrix. The operation unit splits the result matrix of the third operation to obtain the result matrix of the first operation and the result matrix of the second operation respectively.
[0015] For example, when the matrix operation of X×H1 + Y×H2 of the operation unit corresponds to an arithmetic logic unit (i.e., the first matrix operation circuit) that can process a data volume greater than the data volume for generating the output result of the update gate at the t-th time step and the output result of the reset gate at the t-th time step at one time, the operation unit can read the input data x at the t-th time step from the storage unit last time t , the first weight matrix W xz , the second weight matrix W hz, the third weight matrix W xr , the fourth weight matrix W hr and the output data h at the (t-1)th time step t-1 , then, for the first weight matrix W xz and the third weight matrix W xr are concatenated along the width direction to generate the first concatenated weight matrix. For the second weight matrix W xz and the fourth weight matrix W hr are concatenated along the width direction to generate the second concatenated weight matrix. The operation unit runs the arithmetic logic unit corresponding to the matrix operation of X×H1+Y×H2, that is, takes the first concatenated weight matrix as H1 and the second concatenated weight matrix as H2 and inputs them into the above arithmetic logic unit, so that the operation unit can determine the output result z t ' of the update gate and the output result r t ' of the reset gate at the tth time step by reading data from the storage unit once. In this way, the number of times the operation unit reads data from the storage unit through the processing unit is reduced. When the GRU model processes a piece of speech to be recognized, the operation unit only needs to read data from the storage unit n times (reducing the data reading amount by half) to generate the text corresponding to the speech to be recognized, thereby improving the speed of the processor for speech processing, further improving the speech recognition speed of the electronic device, shortening the speech recognition time, and enhancing the user experience.
[0016] The third aspect of the present application provides an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device, and a plurality of processors for running the instructions in the memory to execute the neural network model running method of the first aspect.
[0017] The fourth aspect of the present application provides a computer-readable storage medium, including: instructions are stored on the readable medium of the electronic device, and when the instructions are executed on the electronic device, the electronic device executes the neural network model running method of the first aspect.
[0018] The fifth aspect of the present application provides a computer program product, including instructions for implementing the neural network model running method of the first aspect. Description of the Drawings
[0019] Figure 1A According to some embodiments of the present application, a schematic diagram of the application of a GRU model is shown;
[0020] Figure 1B According to some embodiments of the present application, a schematic diagram of the structure of a GRU model is shown;
[0021] Figure 1CAccording to some embodiments of the present application, there is shown a Figure 1B Schematic diagram of the GRU network at the t-th time step in
[0022] Figure 2 According to some embodiments of the present application, there is shown a schematic diagram of the structure of an electronic device;
[0023] Figure 3 According to some embodiments of the present application, there is shown a schematic diagram of generating the output result of the update gate at the t-th time step;
[0024] Figure 4 According to some embodiments of the present application, there is shown a schematic diagram of generating the output result of the reset gate at the t-th time step;
[0025] Figure 5 According to some embodiments of the present application, there is shown a schematic diagram of concatenating the first weight matrix at the t-th time step and the third weight matrix at the t-th time step;
[0026] Figure 6 According to some embodiments of the present application, there is shown a schematic diagram of concatenating the second weight matrix at the t-th time step and the fourth weight matrix at the t-th time step;
[0027] Figure 7 According to some embodiments of the present application, there is shown a schematic diagram of the first output data at the t-th time step;
[0028] Figure 8 According to some embodiments of the present application, there is shown a schematic diagram of splitting the first output data at the t-th time step into the output result of the update gate and the output result of the reset gate;
[0029] Figure 9 According to some embodiments of the present application, there is shown a schematic diagram of the flow of a GRU model running method;
[0030] Figure 10 According to some embodiments of the present application, there is shown a schematic diagram of the flow of another GRU model running method. Detailed implementation manners
[0031] The illustrative embodiments of the present application include, but are not limited to, a method, apparatus, electronic device, medium, and computer program product for running a neural network model. The embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0032] Since this application involves the content of the Gated Recurrent Unit (GRU) model, to more clearly illustrate the solutions of the embodiments of this application, the GRU model involved in the embodiments of this application will be described in detail below.
[0033] (1) Gated Recurrent Unit (GRU) model
[0034] It is a type of Recurrent Neural Network (RNN) model. Similar to the Long-Short Term Memory (LSTM) model, it is also proposed to solve problems such as long-term memory and gradients in backpropagation.
[0035] Figure 1A Fig. shows an application schematic diagram of a GRU model 10. As Figure 1A shown, the input data of the GRU model 10 is the speech to be recognized, and the output data of the GRU model 10 is the text corresponding to the speech to be recognized.
[0036] In some other embodiments, the GRU model 10 can also be used for text translation between different languages. For example, the input data of the GRU model 10 can be Chinese text, and the output data of the GRU model 10 can be the English text corresponding to the Chinese text. The GRU model 10 can also be used for image classification. For example, the input data of the GRU model 10 can be multiple frames of images, and the output data of the GRU model 10 can be the image type corresponding to each frame of the image. It can be understood that the GRU model 10 is mainly used to process and predict sequence data, and the content recognized by the GRU model 10 in this application is not specifically limited according to the actual application.
[0037] Figure 1B Fig. shows a structural schematic diagram of a GRU model 10. As Figure 1B shown, the GRU model 10 includes n GRU networks, namely GRU1, GRU2, …, GRUt-1, GRUt ……, GRUn. Among them, GRU1 represents the GRU network at the 1st time step, GRU2 represents the GRU network at the 2nd time step, GRUt-1 represents the GRU network at the (t-1)th time step, GRUt represents the GRU network at the tth time step, and GRUn represents the GRU network at the nth time step.
[0038] For example, as Figure 1B shown, {x1, x2, …, x t-1 、x t ……、x n}{ is the speech data to be recognized. Among them, x1 is the speech input data of the GRU network (GRU1) at the first time step, x2 is the speech input data of the GRU network (GRU2) at the second time step, x t is the input data of the GRU network (GRUt) at the t-th time step, ……, x n is the speech input data of the GRU network (GRUn) at the n-th time step. The output data {h1, h2, …, h t-1 、h t ……、h n} of the GRU model 10 can be the text corresponding to the speech to be recognized. Among them, h1 is the output data of the GRU network (GRU1) at the first time step, h2 is the output data of the GRU network (GRU2) at the second time step, h t is the output data of the GRU network (GRUt) at the t-th time step, ……, h n is the output data of the GRU network (GRUn) at the n-th time step.
[0039] It can be understood that the input data or output data of the GRU network at each of the foregoing time steps can be a matrix, a tensor, a vector, etc. For the convenience of description, in the following scheme introduction, the input data or output data matrix of the GRU network is taken as an example to illustrate the data processing of the input or output of each layer of the neural network.
[0040] It is not difficult to see that since the GRU model 10 is mainly used to process and predict sequence data, therefore, the GRU network at the current time step needs to combine the output data of the previous time step, process the input data of the current time step, and generate the output data of the current time step.
[0041] Figure 1C shows a Figure 1B structural schematic diagram of the GRU network at the t-th time step in. According to the different gate structures, the GRU network at the t-th time step can be divided into the following four stages:
[0042] 1. Update stage
[0043] The update stage is used to control the degree to which the state information output by the GRU network at the previous time step is brought into the state of the GRU network at the current time step. The larger the gate value of the update gate, the higher the degree to which the state information of the previous time step is brought into the current state. Specifically, in the GRU network at the t-th time step, the update stage can be based on the output data h t-1 of the GRU network at the (t - 1)-th time step and the input data x t of the GRU network at the t-th time step, calculate the update gate control z t , and through the update gate control z tTo describe the degree of being brought into the current state h t of.
[0044] Exemplarily, the update gate z t can be calculated by the following formula (1):
[0045] z t = σ(x t × W xz + h t-1 × W hz ) (1)
[0046] where σ represents the sigmoid function, that is, the calculation result of x t × W xz + h t-1 × W hz is converted to between (0, 1) through the sigmoid function. W xz represents the weight coefficient matrix of the input data x t of the GRU network at the t-th time step, and W hz represents the weight coefficient matrix of the output data h t-1 of the GRU network at the (t - 1)-th time step. × represents matrix multiplication, and + represents matrix concatenation.
[0047] Exemplarily, the result of matrix multiplication of the input data x t at the t-th time step and the weight matrix W xz is concatenated with the result of matrix multiplication of the output data h t-1 at the (t - 1)-th time step and the weight matrix W xz , that is, x t × W xz + h t-1 × W hz . x t × W xz + h t-1 × W hz can be understood as Figure 1C the vector concatenation operation t11 in. t The update gate z Figure 1C can be understood as
[0048] 2. Reset stage
[0049] The reset phase is used to control how much of the state information output by the GRU network in the previous time step is written into the state of the GRU network in the current time step. The smaller the reset gate, the less state information output by the GRU network in the previous time step is written into the state of the GRU network in the current time step. Specifically, in the GRU network at the t-th time step, the reset phase can be based on the output data h of the GRU network at the (t - 1)-th time step t-1 and the input data x of the GRU network at the t-th time step t to calculate the reset gate r t . The reset gate r t is used to describe the degree to which the state information output by the GRU network in the previous time step is written into the state of the GRU network in the current time step.
[0050] Exemplarily, the reset gate r t can be calculated by the following formula (2):
[0051] r t = σ(x t × W xr + h t-1 × W hr ) (2)
[0052] where σ represents the sigmoid function, that is, the calculation result of x t × W xr + h t-1 × W hr is converted to the range (0, 1) through the sigmoid function. W xr represents the weight coefficient matrix of the input data x of the GRU network at the t-th time step t , and W hr represents the weight coefficient matrix of the output data h of the GRU network at the (t - 1)-th time step t-1 . × represents matrix multiplication, and + represents matrix concatenation.
[0053] It is not difficult to see that the result of multiplying the input data x at the t-th time step by the weight coefficient matrix W t is concatenated with the result of multiplying the output data h at the (t - 1)-th time step by the weight coefficient matrix W xr , that is, x t-1 × W hr + h t × W xr . x t-1 × W hr + h t × W xr + h t-1 × W hr can be understood as Figure 1CThe vector concatenation operation in t13. Reset the gate r t It can be understood as Figure 1C The output of layer t14 in σ.
[0054] 3. Update memory stage
[0055] The update memory phase can update the memory of the input data of the GRU network at the current time step, that is, update the important ones. Specifically, in the GRU network at the tth time step, the update phase can be based on the output data h of the GRU network at the t-1th time step. t-1 And the GRU network input data x at the tth time step t , calculate the reset gate r t , and then use the reset gate r t Realize the intermediate output result c of the tth time step t Calculation.
[0056] For example, the intermediate output result c at the tth time step is t It can be calculated by the following formula (3):
[0057] c t =tanh(x t ×W x +(h t-1 ·r t )W h ) (3)
[0058] Among them, tanh represents the tanh function, that is, for x t ×W x +(h t-1 ·r t )W h The calculated result of W is converted to (-1,1) through the tanh function. x Represents the input data x of the GRU network at the tth time step t The weight coefficient matrix, W h represents the output h of the GRU network at the t-1th time step t-1 and the gate value r of the reset gate t The weight coefficient matrix of the result of bitwise multiplication. · indicates that the data with the same row and column numbers in two matrices are multiplied, × indicates matrix multiplication, and + indicates matrix concatenation.
[0059] It is not difficult to see that the output h of the GRU network at the t-1th time step is t-1 and the gate value r of the reset gate t Bitwise multiplication, i.e. h t-1 ·r t ,h t-1 ·r tIt can be understood as Figure 1C t15 of Figure 1C . Using the input data x at the t-th time step t and the weight coefficient matrix W xr The result of matrix multiplication is combined with the output data h at the (t - 1)-th time step t-1 and the gate value r of the reset gate t The calculation result of bitwise multiplication is combined with the weight coefficient matrix W hr by matrix multiplication, that is, x t ×W x +(h t-1 ·r t )W h . x t ×W x +(h t-1 ·r t )W h It can be understood as Figure 1C The vector concatenation operation t16 in Figure 1C . The intermediate output result c at the t-th time step t It can be understood as Figure 1C The output of the tanh layer t17 in Figure 1C .
[0060] 4. Output stage
[0061] In the output stage, the output and state of the GRU network at the t-th time step can be determined and output. Specifically, in the output stage, based on the output data h of the GRU network at the (t - 1)-th time step t-1 , the gate control z of the update gate t and the intermediate output result c at the t-th time step t , the output data h at the t-th time step can be calculated t .
[0062] Exemplarily, the output data h at the t-th time step t can be calculated by the following formula (4):
[0063] h t =(1 - z t )·c t +z t ·h t-1 (4)
[0064] where z t represents the gate control z of the update gate t , h t-1 represents the output h of the GRU network at the (t - 1)-th time step t-1 , c t represents the intermediate output result c at the t-th time step t . · represents element-wise multiplication of data with the same row and column numbers in two matrices, and + represents vector summation.
[0065] It is not difficult to see that using 1 minus the gate value z of the update gate t is 1 - z t , and 1 - z t can be understood as Figure 1C the operation of t18 in t Using the intermediate output result c at the t-th time step t and performing a bitwise vector multiplication operation on the operation result of 1 - z t , that is, (1 - z t ) · c t , (1 - z t can be understood as Figure 1C the operation of t19 in t-1 Using the output h of the GRU network at the (t - 1)-th time step t and performing a bitwise vector multiplication operation on the gate control z of the update gate t , that is, z t-1 · h t , z t-1 can be understood as Figure 1C the operation of t20 in t-1 Using the output h of the GRU network at the (t - 1)-th time step t and performing a bitwise vector multiplication operation on the gate control z of the update gate t and then performing a vector summation operation on the operation result of the bitwise vector multiplication operation of the output result c at the t-th time step t and 1 - z t , that is, (1 - z t ) · c t + z t-1 · h t , (1 - z t ) · c t + z t-1 can be understood as Figure 1C the operation of t21 in
[0066] It can be understood that in the processor of the electronic device, there are operation units for each operation item in the above formulas (1) to (4). By reading the input data of each operation item from the storage unit of the processor, the operation result of the operation item can be obtained using the operation unit. Among them, the operation item can be at least part of the above formulas, such as x t × W xz + h t-1 × W hz in formula (1), x t × W xz + h t-1 × W hz, sigmoid function σ(), etc.
[0067] When the arithmetic unit of the processor of the electronic device runs the GRU network defined by the foregoing formulas (1) to (4), it needs to first read the input data of the arithmetic terms in each formula from the storage unit of the processor into the corresponding arithmetic logic circuit of the arithmetic unit, and then obtain the arithmetic results of each arithmetic term according to the arithmetic logic of each arithmetic term. That is to say, every time the arithmetic unit of the processor runs an arithmetic term, it needs to read the input data from the storage unit of the processor once, which increases the number of times the arithmetic unit of the processor reads data from the storage unit and reduces the running speed of the neural network model.
[0068] For example, when the processing unit of the processor runs the matrix arithmetic logic of X×H1 + Y×H2 to calculate the x of formula (1) t ×W xz +h t-1 ×W hz When calculating the operation result, it is necessary to first read the input data x t , weight coefficient matrix W xz , weight coefficient matrix W hz , and the output data h of the (t - 1)-th time step t-1 . After the processing unit reads this data, it runs the matrix arithmetic logic of X×H1 + Y×H2 to perform arithmetic on x t ×W xz +h t-1 ×W hz to obtain the operation result of x t ×W xz +h t-1 ×W hz . Then, the processing unit of the processor runs the matrix arithmetic logic of X×H1 + Y×H2 again to calculate the x of formula (2) t ×W xr +h t-1 ×W hr When calculating the operation result, and it is necessary to read the input data x t , weight coefficient matrix W xr , weight coefficient matrix W hr , and the output data h of the (t - 1)-th time step t-1 again. After the processing unit reads this data, it runs the matrix arithmetic logic of X×H1 + Y×H2 to perform arithmetic on x t ×W xr +h t-1 ×W hr to obtain the operation result of x t ×W xr +h t-1 ×Whr The operation result. It is not difficult to see that when the operation unit calculates x t ×W xz +h t-1 ×W hz and the operation result of x t ×W xr +h t-1 ×W hr , it is necessary to read data from the storage unit twice to obtain the operation result of x t ×W xz +h t-1 ×W hz and the operation result of x t ×W xr +h t-1 ×W hr .
[0069] To reduce the number of times the operation unit reads data from the storage unit, the embodiment of the present application provides a method for running a neural network model. By combining operation terms with the same operation form and the same input data (such as "x t ×W xz +h t-1 ×W hz " in formula (1) and "x t ×W xr +h t-1 ×W hr The operation forms are the same, and the input data are both x t and h t-1 ), the weight matrices (W xz and W xr and W hz and W hr ) are concatenated, and then the operation unit corresponding to the operation term of this form is used to obtain the operation results of each operation term by obtaining the input data from the storage unit once based on the concatenated weight matrix. Thereby, the speed of the electronic device for operating the neural network model is improved.
[0070] For example, the operation unit of the processor reads the input data x at the t-th time step, the weight coefficient matrix W t , the weight coefficient matrix W xz , the weight coefficient matrix W hz , the weight coefficient matrix W xr , the weight coefficient matrix W hr and the output data h at the (t - 1)-th time step from the storage unit of the processor once t-1 , and concatenate the aforementioned weight coefficient matrix W xz and the weight coefficient matrix W xr to obtain the concatenated weight matrix W xzr, splice the aforementioned weight coefficient matrix \(W\) hz and the weight coefficient matrix \(W\) hr to obtain a spliced weight matrix \(W\) hzr . The operation unit performs matrix operations of \(X\times H1 + Y\times H2\). Using the spliced weight matrix \(W\) xzr as \(H1\), the spliced weight matrix \(W\) hzr as \(H2\), the input data \(x\) at the \(t\)-th time step t as \(X\), and the output data \(h\) at the \((t - 1)\)-th time step t-1 as \(Y\), perform corresponding matrix operations, and the operation result of the formula (1) \(x\) t \(\times W\) xz \(+ h\) t-1 \(\times W\) hz can be obtained in one operation. And the operation result of \(x\) t \(\times W\) xr \(+ h\) t-1 \(\times W\) hr . Then, according to the dimension values of the weight coefficient matrix \(W\) xz and the weight coefficient matrix \(W\) xr , split the operation results of \(x\) t \(\times W\) xz \(+ h\) t-1 \(\times W\) hz and \(x\) t \(\times W\) xr \(+ h\) t-1 \(\times W\) hr to respectively obtain the operation results of \(x\) t \(\times W\) xz \(+ h\) t-1 \(\times W\) hz , and \(x\) t \(\times W\) xr \(+ h\) t-1 \(\times W\) hr .
[0071] To facilitate understanding of the technical solution of the embodiments of the present application, the electronic device 20 for executing the calculation process of the threshold recurrent unit model 10 will be introduced below.
[0072] Figure 2 According to some embodiments of the present application, a schematic structural diagram of an electronic device 20 is shown. As Figure 2 shown, the electronic device 20 includes a processor 21, a system memory 22, a non-volatile memory 23, an input / output device 24, a communication interface 25, and a system control logic 26 for coupling the processor 21, the system memory 22, the non-volatile memory 23, the input / output device 24, and the communication interface 25. Among them:
[0073] The processor 201 may include one or more processing units, for example, may include processing modules or processing circuits such as a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Digital Signal Processor (DSP), a Micro-programmed Control Unit (MCU), a Field Programmable Gate Array (FPGA), an Artificial Intelligence Processing Unit (AIPU), and a Neural-network Processing Unit (NPU). Among them, different processing units may be independent devices or integrated in one or more processors. In some embodiments, the processor 201 may execute the calculation process of the neural network model 10.
[0074] In some embodiments, the processor 21 may include a control unit 210, an arithmetic unit 211, and a storage unit 212. The control unit 210 is used to schedule the processor 21. In some embodiments, the control unit 210 further includes a Direct Memory Access Controller (DMAC) 2101 for transferring the data in the storage unit 212 to other units, such as transferring it to the system memory 22.
[0075] The arithmetic unit 211 is used to perform specific arithmetic and / or logical operations. In some embodiments, the arithmetic unit 211 may include an arithmetic logic unit, which refers to a combinational logic circuit that can implement multiple sets of arithmetic operations and logical operations and is used to perform arithmetic and logical operations. For example, the arithmetic unit 211 includes arithmetic logic units corresponding to the operators of formulas (1) to (4).
[0076] In some embodiments, the arithmetic unit 211 internally includes multiple processing units (Process Engine, PE). In some implementations, the arithmetic unit 211 is a two-dimensional systolic array. The arithmetic unit 211 may also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic unit 211 is a general matrix processor.
[0077] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The operation unit 211 fetches the corresponding data of the input matrix A and the weight matrix B from the storage unit 212 and caches them on each PE in the operation circuit. The operation circuit performs matrix operations on the data in matrix A and the data in matrix B to obtain partial results or final results of the matrix. It is not difficult to understand that the operation unit 211 can process matrix operations of size 100×100 at a time to obtain the final result of the matrix operation, or can also process matrix operations of size 10×10 at a time to obtain the final result of the matrix operation.
[0078] In some other embodiments, the operation unit 211 may further include multiple application specific integrated circuits (ASICs) suitable for running neural network models, such as convolution calculation units, vector calculation units, etc. In some embodiments, the operation unit 211 can be used to read the input data, the first weight matrix, the second weight matrix, the third weight matrix, the fourth weight matrix at the t-th time step, and the output data at the (t - 1)-th time step from the storage unit 212, and generate the gate value of the reset gate and the gate value of the update gate at the t-th time step through arithmetic operations. The storage unit 212 is used to temporarily store the input and / or output data of the operation unit 211. For example, the storage unit 212 can be used to store the input data, the first weight matrix, the second weight matrix, the third weight matrix, the fourth weight matrix at the t-th time step, and the output data at the (t - 1)-th time step.
[0079] It can be understood that in some other embodiments, the DMAC 2101 may not be integrated in the processor 21, but an independent module coupled to the system control logic 26, which is not limited in the embodiments of the present application.
[0080] The system memory 22 may include memories such as random access memory (RAM), double data rate synchronous dynamic random access memory (DDR SDRAM), etc., and is used to temporarily store data or instructions of the electronic device 20.
[0081] The non-volatile memory 23 may include one or more tangible, non-transitory computer-readable media for permanently storing data and / or instructions. The non-volatile memory 23 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as a hard disk drive (HDD), a compact disc (CD), a digital versatile disc (DVD), a solid-state drive (SSD), etc. In some embodiments, the non-volatile memory 23 may also be a removable storage medium, such as a Secure Digital (SD) memory card, etc. In some embodiments, the non-volatile memory 23 is used to permanently store data or instructions of the electronic device 20, such as instructions for storing the neural network model 10.
[0082] The input / output (I / O) device 24 may include an input device, such as a keyboard, a mouse, a touch screen, etc., for converting a user's operation into an analog or digital signal and transmitting it to the processor 21; and an output device, such as a speaker, a printer, a display, etc., for presenting the information in the electronic device 20 to the user in the form of sound, text, images, etc.
[0083] The communication interface 25 provides a software / hardware interface for the electronic device 20 to communicate with other electronic devices, enabling the electronic device 20 to exchange data with other electronic devices 20. For example, the electronic device 20 can obtain data for running the neural network model from other electronic devices through the communication interface 25, and can also transmit the operation results of the neural network model to other electronic devices through the communication interface 25.
[0084] The system control logic 26 may include any suitable interface controller to provide any suitable interface to other modules of the electronic device 20, enabling the various modules of the electronic device 20 to communicate with each other.
[0085] In some embodiments, at least one of the processors 21 may be logically packaged with one or more controllers for the system control logic 26 to form a System in Package (SiP). In other embodiments, at least one of the processors 21 may also be integrated with the logic of one or more controllers for the system control logic 26 on the same chip to form a System-on-Chip (SoC).
[0086] It can be understood that Figure 2The hardware structure of the illustrated electronic device 20 is only an example. In some other embodiments, the electronic device 20 may also include more or fewer modules, and some modules may be combined or split. The embodiments of the present application do not make any limitations.
[0087] It can be understood that the electronic device 20 can be any electronic device capable of running the GRU model 10, including but not limited to laptop computers, desktop computers, tablet computers, mobile phones, servers, wearable devices, head-mounted displays, mobile email devices, portable game consoles, portable music players, reader devices, and televisions in which one or more processors are embedded or coupled. The embodiments of the present application do not make any limitations.
[0088] For ease of understanding, the following describes the process of the electronic device running the GRU model 10 in conjunction with the structure of the electronic device 20.
[0089] It can be understood that the process of the electronic device 20 running each GRU network in the GRU model 10 is similar. Taking the t-th GRU network as an example, the following introduces the process of the electronic device 20 running the GRU model 10.
[0090] The embodiments of the present application provide a method for running a GRU model. The arithmetic unit 211 processes the input data at the t-th time step, the first weight matrix (i.e., the weight coefficient matrix W in formula (1)), the second weight matrix (i.e., the weight coefficient matrix W in formula (1)), and the output data at the (t - 1)-th time step read from the storage unit 212 for the first time. By running the non-linear operator and linear operator of the update gate, the input data at the t-th time step, the first weight matrix, the second weight matrix, and the output data at the (t - 1)-th time step are processed. The arithmetic unit 211 determines the output result of the update gate (i.e., the calculation result of the operator x in formula (1) × W + h × W) by running the non-linear operator. Then, according to the input data xt at the t-th time step, the third weight matrix (i.e., the weight coefficient matrix W in formula (2)), the fourth weight matrix (i.e., the weight coefficient matrix W in formula (2)), and the output data at the (t - 1)-th time step read from the storage unit 212 for the second time, the arithmetic unit 211 generates the output result of the reset gate (i.e., the calculation result of the operator x in formula (2) × W + h × W) by running the linear operator of the reset gate. xz hz t ×W xz +h t-1 ×W hz xr hr t ×W xr +h t-1 ×W hr
[0091] The following takes the input data x at the t-th time step...t is a matrix of size 1×2, and the output data h at the (t-1)-th time step t-1 Taking, for example, a matrix of size 1×4, a first weight matrix of size 2×4, and a second weight matrix of size 4×4, the linear operator of the update gate of the GRU network at the t-th time step is described for the operation unit 211 (i.e., the operator x in Equation 1 t ×W xz +h t-1 ×W hz ), and the output result z of the update gate at the t-th time step is generated t ′ (a matrix of size 1×4).
[0092] For example, as Figure 3 shown, the input data x at the t-th time step t is a matrix of size 1×2, where the first row data of the input data x t includes 1 and 2. The first weight matrix W at the t-th time step xz is a matrix of size 2×4, where the first row data of the first weight matrix W xz includes 3, 2, 5, and 4, and the second row data of the first weight matrix W xz includes 4, 1, 1, and 1. The output data h at the (t-1)-th time step t-1 is a matrix of size 1×4, where the first row data of the output data h t-1 includes 1, 2, 1, and 3. The second weight matrix W at the t-th time step hz is a matrix of size 4×4, where the first row data of the second weight matrix W hz includes 3, 2, 1, and 0, the second row data of the second weight matrix W hz includes 2, 1, 0, and 3, the third row data includes 1, 0, 1, and 2, and the fourth row data includes 1, 3, 2, and 1. The operation unit 211 runs the linear operator of the update gate, and the generated output result z t ′ is a matrix of size 1×4, where the first row data of the output result z t ′ includes 22, 17, 15, and 17.
[0093] Next, taking the input data x at the t-th time step t as a matrix of size 1×2, the output data h at the (t-1)-th time step t-1 as a matrix of size 1×4, a third weight matrix W xr as a matrix of size 2×4, and a fourth weight matrix W hr as a matrix of size 4×4 as an example, the linear operator of the reset gate of the GRU network at the t-th time step is described for the operation unit 211 (i.e., the operator x in Equation 2t ×W xr +h t-1 ×W hr ) to generate the output result r of the reset gate at the t-th time step t ′ (a matrix of size 1×4).
[0094] For example, as Figure 4 shown, the input data x at the t-th time step t is a matrix of size 1×2, where the first row data of the input data x t includes 1 and 2. The third weight matrix W at the t-th time step xr is a matrix of size 2×4, where the first row data of the third weight matrix W xr includes 3, 2, 0, 2, and the second row data of the third weight matrix W xr includes 2, 1, 1, 1. The output data h at the (t - 1)-th time step t-1 is a matrix of size 1×4, where the first row data of the output data h t-1 includes 1, 2, 1, 3. The fourth weight matrix W at the t-th time step hr is a matrix of size 4×4, where the first row data of the fourth weight matrix W hr includes 0, 2, 1, 0, the second row data of the fourth weight matrix W hr includes 2, 1, 0, 0, the third row data of the fourth weight matrix W hr includes 1, 0, 1, 2, and the fourth row data of the fourth weight matrix W hr includes 1, 0, 2, 1. The arithmetic unit 211 runs the linear operator of the reset gate to generate the output result z t ′ of the reset gate, which is a matrix of size 1×4, where the first row data of the output result z t ′ includes 15, 8, 10, 9.
[0095] From Figures 3 to 4 the description, it can be seen that the arithmetic unit 211 needs to first read the input data x at the t-th time step from the storage unit 212 t , the output data h at the (t - 1)-th time step t-1 , the first weight matrix W xz and the second weight matrix W hz , and generate the output result z t ′ of the update gate by running the linear operator of the update gate. Then, the arithmetic unit 211 reads the input data x at the t-th time step from the storage unit 212 again t , the output data h at the (t - 1)-th time step t-1 , the third weight matrix W xr and the fourth weight matrix Whr By running the linear operator of the reset gate, the output result r of the reset gate is generated t ′.
[0096] It is not difficult to see that in the process of the arithmetic unit 211 inferring the input data of one time step to obtain the text corresponding to the speech data of this time step, it is necessary to read the data twice from the storage unit 212 (that is, the input data x of the t-th time step read for the first time mentioned above t , the output data h of the (t - 1)-th time step t-1 , the first weight matrix W xz and the second weight matrix W hz y and the input data x of the t-th time step read for the second time t , the output data h of the (t - 1)-th time step t-1 , the third weight matrix W xr and the fourth weight matrix W hr ), and perform matrix operations on the data read each time to generate the output result z t ′ of the update gate of the t-th time step and the output result r t ′ of the reset gate of the reset gate, which reduces the speed of the arithmetic unit to infer the input data.
[0097] It is not difficult to understand that the GRU model 10 includes GRU networks of n time steps. In the GRU network of each time step of the GRU model 10, the arithmetic unit 211 needs to read the data twice from the storage unit 212 to calculate the gate value of the update gate and the gate value of the reset gate of the GRU network of each time step. Therefore, when the GRU model 10 processes a piece of speech, the arithmetic unit 211 needs to read the data from the storage unit 212 at least 2 * n times. When the GRU model 10 processes 10 pieces of speech, the arithmetic unit 211 needs to read the data from the storage unit 212 at least 20 * n times. In this way, the arithmetic unit 211 needs to read the data from the storage unit 212 multiple times when running the GRU model 10 for data processing (for example, speech recognition), which affects the speed of the processor 21 for speech processing, and further affects the speed of speech recognition of the electronic device 20, resulting in a long speech recognition time and affecting the user experience.
[0098] It can be seen from the above description that the matrix operation structures of the update gate operator and the reset gate operator are the same, that is, both are matrix operations of X×H1 + Y×H2, that is, the arithmetic unit 211 has an arithmetic logic unit corresponding to the matrix operation of X×H1 + Y×H2. However, the weight coefficients of the two input data of the matrix operations of the update gate operator and the reset gate operator are different, that is, the first weight matrix is different from the third weight matrix and the second weight matrix is different from the fourth weight matrix.
[0099] To reduce the number of times the arithmetic unit 211 reads data from the storage unit 212, an embodiment of the present application provides another method for running a GRU model. The arithmetic unit 211 reads the input data x at the t-th time step from the storage unit 212 once t , the first weight matrix W xz , the second weight matrix W hz , the third weight matrix W xr , the fourth weight matrix W hr , and the output data h at the (t - 1)-th time step t-1 . Then, the first weight matrix W xz and the third weight matrix W xr are concatenated to obtain the first concatenated weight matrix. The second weight matrix W hz and the fourth weight matrix W hr are concatenated to obtain the second concatenated weight matrix. The first concatenated weight matrix, the second concatenated weight matrix, the input data x at the t-th time step t , and the output data h at the (t - 1)-th time step t-1 are subjected to matrix operations of X×H1 + Y×H2 to obtain the output results z t ′ of the update gate and the output results r t ′ at the t-th time step. In this way, the number of times the arithmetic unit 211 reads data from the storage unit 212 is reduced, thereby improving the speed of the arithmetic unit for running the neural network model
[0100] For example, if the arithmetic logic unit corresponding to the X×H1 + Y×H2 matrix operation of the arithmetic unit 211 can process a data volume greater than the data volume for generating the output results of the update gate at the t-th time step and generating the output results of the reset gate at the t-th time step at one time, the arithmetic unit 211 can read the input data x at the t-th time step from the storage unit 212 once t , the first weight matrix W xz , the second weight matrix W hz , the third weight matrix W xr , the fourth weight matrix W hr , and the output data h at the (t - 1)-th time step t-1 . Then, the first weight matrix W xz and the third weight matrix W xr are concatenated along the width direction to generate the first concatenated weight matrix. The second weight matrix W xz and the fourth weight matrix W hrStitch along the width direction to generate a second stitching weight matrix. The operation unit 211 passes through the arithmetic logic unit corresponding to the matrix operation of X×H1 + Y×H2, that is, takes the first stitching weight matrix as H1 and the second stitching weight matrix as H2 and inputs them into the above arithmetic logic unit, so that the operation unit 211 can determine the output result z of the update gate at the t-th time step by reading data from the storage unit 212 once. t ' and the output result r t ' of the reset gate. In this way, the number of times the operation unit 211 reads data from the storage unit 212 is reduced. When the GRU model 10 processes a piece of speech to be recognized, the operation unit 211 only needs to read data from the storage unit 212 n times (halving the amount of data read) to generate the text corresponding to the speech to be recognized, thereby improving the speed of the processor 21 for speech processing, further improving the speed of speech recognition of the electronic device 20, shortening the speech recognition time, and enhancing the user experience.
[0101] Specifically, for Figure 3 and Figure 4 the input data x at the t-th time step shown t , the output data h at the (t - 1)-th time step t-1 , the first weight matrix W xz , the second weight matrix W hz matrix operation and the input data x at the t-th time step t , the output data h at the (t - 1)-th time step t-1 , the third weight matrix W xr and the fourth weight matrix W hr matrix operation, the operation unit 211 can simultaneously read the input data x at the t-th time step t , the output data h at the (t - 1)-th time step t-1 , the first weight matrix W xz , the second weight matrix W hz , the third weight matrix W xr and the fourth weight matrix W hr from the storage unit 212, and splice the first weight matrix W xz and the third weight matrix W xr into the Figure 5 shown first stitching weight matrix W xzr , and splice the second weight matrix W hz and the fourth weight matrix W hr into the Figure 6 shown second stitching weight matrix W hzr ; then take the first stitching weight matrix W xzr , the second stitching weight matrix W hzrand the input data x at the t-th time step t , the output data h at the (t - 1)-th time step t-1 perform the matrix operation of X×H1 + Y×H2 to obtain Figure 7 the first output data zr at the t-th time step shown t '; refer to Figure 8 , and then the operation unit 211 will Figure 7 obtained Figure 3 and Figure 4 the output result z t ' of the update gate and the output result r t ' of the reset gate at the t-th time step are split to obtain the output result z t ' of the update gate and the output result r t ' of the reset gate at the t-th time step. The specific calculation process will be described below and will not be elaborated here.
[0102] Next, in combination with the hardware structure of the electronic device 20, the process of the electronic device 20 executing the GRU model 10 operation method provided in the embodiments of the present application will be introduced in detail.
[0103] Figure 9 According to some embodiments of the present application, a flowchart of a GRU model 10 operation method is shown. Next, taking the GRU network (GRUt) at the t-th time step in the GRU model 10 as an example, the embodiments of the present application will be introduced. The operation method includes:
[0104] S901: The operation unit 211 reads the input data, the first weight matrix, the second weight matrix, and the output data at the (t - 1)-th time step from the storage unit 212. Among them, the first weight matrix is the weight coefficient matrix W in formula (1) xz , and the second weight matrix is the weight coefficient matrix W in formula (1) hz .
[0105] S902: The operation unit 211 determines the output result of the update gate at the t-th time step by running the linear operator of the update gate of the GRU network at the t-th time step according to the input data, the first weight matrix, the second weight matrix, and the output data at the (t - 1)-th time step. Among them, the linear operator of the update gate of the GRU network at the t-th time step can be the operator x t ×W xz +h t-1 ×W hz in formula (1).
[0106] S903: The operation unit 211 determines the gate value of the update gate of the GRU network at the t-th time step by running the non-linear operator of the update gate of the GRU network at the t-th time step according to the output result of the update gate. Among them, the non-linear operator of the update gate of the GRU network can be the σ() (i.e., sigmoid function) operator in formula (1).
[0107] S904: The operation unit 211 reads the input data, the third weight matrix, the fourth weight matrix, and the output data of the (t - 1)-th time step at the t-th time step from the storage unit 212. Among them, the third weight matrix is the weight coefficient matrix W in formula (2). xr and the second weight matrix is the weight coefficient matrix W in formula (2). hr 。
[0108] S905: The operation unit 211 determines the output result of the reset gate by the linear operator of the reset gate of the GRU network according to the input data, the third weight matrix, the fourth weight matrix, and the output data of the (t - 1)-th time step at the t-th time step.
[0109] S906: The operation unit 211 determines the gate value of the reset gate of the GRU network at the t-th time step by running the non-linear operator of the reset gate of the GRU network at the t-th time step according to the output result of the reset gate. Among them, the non-linear operator of the reset gate of the GRU network can be the σ() (i.e., sigmoid function) operator in formula (2).
[0110] S907: The operation unit 211 determines the output data at the t-th time step according to the gate value of the reset gate at the t-th time step, the gate value of the update gate at the t-th time step, the input data at the t-th time step, and the output data of the (t - 1)-th time step.
[0111] In some embodiments, the operation unit 211 determines the output data at the t-th time step according to the gate value of the reset gate at the t-th time step, the gate value of the update gate at the t-th time step, the input data at the t-th time step, and the output data of the (t - 1)-th time step. Specifically, the output data at the t-th time step can be calculated by formula (3) and formula (4). For the specific calculation process, refer to the descriptions of formula (3) and formula (4), which will not be elaborated here.
[0112] It is not difficult to see that in steps S901 to S906, when the arithmetic unit 211 calculates the gate values of the reset gate and the update gate of the GRU network at the t-th time step, it needs to read data once at step S901, that is, read the input data at the t-th time step, the first weight matrix, the second weight matrix, and the output data at the (t - 1)-th time step, and read data again at step S903, that is, read the input data at the t-th time step, the third weight matrix, the fourth weight matrix, and the output data at the (t - 1)-th time step.
[0113] Therefore, in order to reduce the number of times the arithmetic unit 211 reads data from the storage unit 212 during the process of calculating the gate values of the reset gate and the update gate at the t-th time step. The present application also provides a method for running the GRU model 10, and the arithmetic unit 211 only needs to read the input data at the t-th time step, the first weight matrix, the second weight matrix, the third weight matrix, the fourth weight matrix, and the output data at the (t - 1)-th time step from the storage unit 212 once. By splicing the first weight matrix and the third weight matrix, and splicing the second weight matrix and the fourth weight matrix, and then the arithmetic unit 211 generates the gate value of the reset gate at the t-th time step and the gate value of the update gate at the t-th time step according to the spliced weight matrices, the input data at the t-th time step, and the output data at the (t - 1)-th time step.
[0114] Figure 10 According to some embodiments of the present application, a schematic flowchart of another method for running the GRU model 10 is shown. Taking the GRU network (GRUt) at the t-th time step in the GRU model 10 as an example below, the embodiments of the present application will be introduced. The running method includes:
[0115] S1001: The arithmetic unit 211 reads the input data at the t-th time step, the first weight matrix, the second weight matrix, the third weight matrix, the fourth weight matrix, and the output data at the (t - 1)-th time step from the storage unit 212.
[0116] For example, as Figures 1A to 1C shown, taking the GRU network (GRUt) at the t-th time step in the GRU model 10 as an example, in order to calculate the gate value of the update gate and the gate value of the reset gate, the arithmetic unit 211 can read the input data at the t-th time step, the first weight matrix, the second weight matrix, the third weight matrix, the fourth weight matrix, and the output data at the (t - 1)-th time step from the storage unit 212, where the first weight matrix is the weight coefficient matrix W in formula (1) xz , the second weight matrix is the weight coefficient matrix W in formula (1) hz , and the second weight matrix is the weight coefficient matrix W in formula (2) xr, the second weight matrix is the weight coefficient matrix \(W\) in formula (2). hr .
[0117] In some embodiments, the input data \(x\) at the \(t\)-th time step t is a matrix of size \(H1\times W1\), where \(H1\) represents the number of rows of the input data \(x\) at the \(t\)-th time step t , and \(W1\) represents the number of columns of the input data \(x\) at the \(t\)-th time step t . The first weight matrix \(W\) at the \(t\)-th time step xz and the third weight matrix \(W\) xr are both matrices of size \(H2\times W2\), and \(H2\) represents the number of rows of the first weight matrix \(W\) xz and the third weight matrix \(W\) xr . The second weight matrix \(W\) at the \(t\)-th time step hz and the fourth weight matrix \(W\) hr are both matrices of size \(H3\times W3\), and \(H3\) represents the number of rows of the second weight matrix \(W\) hz and the fourth weight matrix \(W\) hr . \(W3\) represents the number of columns of the second weight matrix \(W\) hz and the fourth weight matrix \(W\) hr . The output data \(h\) at the \((t - 1)\)-th time step t-1 is a matrix of size \(H4\times W4\), and \(H4\) represents the number of rows of the output data \(h\) at the \((t - 1)\)-th time step t-1 , and \(W4\) represents the number of columns of the output data \(h\) at the \((t - 1)\)-th time step t-1 .
[0118] It is not difficult to understand that in order to ensure the linear operator of the update gate of the GRU network at the \(t\)-th time step, the matrix multiplication of the input data \(x\) t and the first weight matrix \(W\) xz , the matrix multiplication of the output data \(h\) t-1 and the second weight matrix \(W\) hz , and the result of the matrix multiplication of the input data \(x\) t and the first weight matrix \(W\) xz and the result of the matrix multiplication of the output data \(h\) t-1 and the second weight matrix \(W\) hzThe result of matrix multiplication is used to perform matrix addition. The value of the number of rows H1 of the input data at the t-th time step is equal to the value of the number of rows H4 of the output data at the (t - 1)-th time step. The value of the number of rows H2 of the first weight matrix is equal to the value of the number of columns W1 of the input data at the t-th time step. The value of the number of rows H3 of the second weight matrix is equal to the value of the number of columns W4 of the output data at the (t - 1)-th time step. The number of columns W2 of the first weight matrix is equal to the number of columns W3 of the second weight matrix. Similarly, for the corresponding relationship between the number of rows and columns of each matrix in the linear operator of the reset gate of the GRU network at the t-th time step, refer to the corresponding relationship between the number of rows and columns of each matrix in the linear operator of the update gate of the GRU network at the t-th time step, which will not be elaborated here.
[0119] In some embodiments, the data in the input data at the t-th time step, the first weight matrix, the second weight matrix, the third weight matrix, the fourth weight matrix, and the output data at the (t - 1)-th time step can be integers or floating-point numbers. According to the actual application, the present application does not specifically limit the data types of the input data at the t-th time step, the first weight matrix, the second weight matrix, and the output data at the (t - 1)-th time step.
[0120] S1002: The operation unit 211 splices the first weight matrix at the t-th time step and the third weight matrix at the t-th time step along the width direction of the matrix to generate a spliced first spliced weight matrix.
[0121] In some embodiments, the first weight matrix W at the t-th time step xz is a matrix of size H2×W2, where H2 represents both the number of rows of W xz and the height of the matrix of W xz , and W2 represents both the number of columns of W xz and the width of the matrix of W xz . The third weight matrix W at the t-th time step xr is a matrix of size H2×W2, where H2 represents both the number of rows of W xr and the height of the matrix of W xr , and W2 represents both the number of columns of W xr and the width of the matrix of W xr .
[0122] It is not difficult to see that the height values and width values of the first weight matrix and the third weight matrix at the t-th time step are respectively equal, but the data in the first weight matrix and the third weight matrix at the t-th time step are different. Therefore, the operation unit 211 can splice the first weight matrix at the t-th time step and the third weight matrix at the t-th time step along the width direction of the matrix to generate a spliced first spliced weight matrix.
[0123] In some embodiments, the first weight matrix at the t-th time step is a matrix of size H2×W2, and the third weight matrix at the t-th time step is a matrix of size H2×W2. The operation unit 211 splices the first weight matrix at the t-th time step and the third weight matrix at the t-th time step along the width direction of the matrix to generate a spliced first spliced weight matrix. The spliced first spliced weight matrix is a matrix of size H2×2*W2.
[0124] For example, as Figure 5 shown, the first weight matrix W at the t-th time step xz and the third weight matrix W at the t-th time step xr are both matrices of size 2×4. The first row data of the first weight matrix W xz include 3, 2, 5, 4, and the second row data of the first weight matrix W xz include 4, 1, 1, 1. The first row data of the third weight matrix W xr include 3, 2, 0, 2, and the second row data of the third weight matrix W xr include 2, 1, 1, 1. Then the operation unit 211 splices the first weight matrix W xz and the third weight matrix W xr along the width direction of the matrix to generate a spliced first spliced weight matrix W xzr of size 2×8. As Figure 5 shown, the first row data of the first spliced weight matrix W xzr include: 3, 2, 5, 4, 3, 2, 0, 2, and the second row data of the first spliced weight matrix W xzr include: 4, 1, 1, 1, 2, 1, 1, 1.
[0125] S1003: The operation unit 211 splices the second weight matrix at the t-th time step and the fourth weight matrix at the t-th time step along the width direction of the matrix to generate a spliced second spliced weight matrix.
[0126] In some embodiments, the second weight matrix W at the t-th time step hz is a matrix of size H3×W3. H3 represents both the number of rows of W hz and the height of the matrix of W hz , and W3 represents both the number of columns of W hz and the width of the matrix of W hz . The third weight matrix W at the t-th time step hr is a matrix of size H2×W2. H2 represents both the number of rows of W hr and the height of the matrix of W hr and W2 represents both the number of columns of Whr represents the number of columns and also represents W hr the width of the matrix of
[0127] It is not difficult to see that the width values and height values of the second weight matrix and the fourth weight matrix at the t-th time step are respectively equal, but the data in the second weight matrix and the fourth weight matrix at the t-th time step are different. The operation unit 211 can splice the second weight matrix and the fourth weight matrix at the t-th time step along the height direction of the matrix to generate a spliced second spliced weight matrix.
[0128] In some embodiments, both the second weight matrix and the fourth weight matrix at the t-th time step are matrices of size H3×W3. Therefore, the operation unit 211 can splice the second weight matrix and the fourth weight matrix at the t-th time step along the width direction of the matrix to generate a spliced second spliced weight matrix. The second spliced weight matrix is a matrix of size H3×2*W3.
[0129] For example, as Figure 6 shown, the second weight matrix W hz and the fourth weight matrix W xr at the t-th time step are both matrices of size 4×4. Among them, the first row data of the second weight matrix W hz include: 3, 2, 1, 0, the second row data of the second weight matrix W hz include: 2, 1, 0, 3, the third row data include: 1, 0, 1, 2, and the fourth row data include 1, 3, 2, 1. The first row data of the fourth weight matrix W hr include 0, 2, 1, 0, the second row data of the fourth weight matrix W hr include: 2, 1, 0, 0, the third row data of the fourth weight matrix W hr include: 1, 0, 1, 2, and the fourth row data of the fourth weight matrix W hr include: 1, 0, 2, 1. The operation unit 211 splices the second weight matrix W hz and the third weight matrix W hr along the width direction of the matrix to generate a spliced second spliced weight matrix W hzr of size 4×8. Among them, the first row data of the second spliced weight matrix W hzr include: 3, 2, 1, 0, 0, 2, 1, 0, the second row data of the second weight matrix W hz include: 2, 1, 0, 3, 2, 1, 0, 0, the second row data of the second weight matrix W hz include: 1, 0, 1, 2, 1, 0, 1, 2, the second weight matrix Whz The fourth row data of
[0130] S1004: The operation unit 211 determines the first operator of the GRU network at the t-th time step according to the linear operators of the reset gate and the update gate of the GRU network at the t-th time step, and runs the first operator of the GRU network at the t-th time step according to the input data at the t-th time step, the first splicing weight matrix, the second splicing weight matrix, and the output data at the (t - 1)-th time step to generate the first output data of the GRU network at the t-th time step.
[0131] In some embodiments, the operation unit 211 determines the first operator at the t-th time step according to the linear operators of the reset gate and the update gate of the GRU network. The operation unit 211 runs the first operator at the t-th time step according to the input data at the t-th time step, the first splicing weight matrix, the second splicing weight matrix, and the output data at the (t - 1)-th time step to generate the first output data at the t-th time step. Specifically, the first output data zr t ′ can be represented by the following formula (5):
[0132] zr t ′ = x t ×W xzr +h t-1 ×W hzr (5)
[0133] In formula (5), x t ×W xzr +h t-1 ×W hzr represents the first operator of the GRU network at the t-th time step, x t represents the input data at the t-th time step, h t-1 represents the output data at the (t - 1)-th time step, W xzr represents the first splicing weight matrix at the t-th time step, W hzr represents the second splicing weight matrix at the t-th time step, × represents matrix multiplication, and + represents matrix addition.
[0134] For example, as Figure 7 shown, the input data x t at the t-th time step is a matrix of size 1×2, where the first row data of the input data x t includes 1, 2. The output data h t-1 at the (t - 1)-th time step is a 1×4 matrix, where the first row data of the output data h t-1 includes 1, 2, 1, 3. The first splicing weight matrix W xzris a matrix of size 2×8. The first splicing weight matrix W xzr The specific data reference in Figure 5 description. The second splicing weight matrix W hzr is a matrix of size 4×8. The second splicing weight matrix W hzr The specific data reference in Figure 6 description.
[0135] Specifically, the operation unit 211 runs the first operator of the GRU network at the t-th time step based on the input data x t at the t-th time step, the first splicing weight matrix W xzr the second splicing weight matrix W hzr and the output data h t-1 at the (t - 1)-th time step to generate the first output data zr t ' of the GRU network at the t-th time step. As Figure 7 shown, the first output data zr t ' of the GRU network at the t-th time step is a matrix of size 1×8. The first row data of the first output data zr t ' of the GRU network at the t-th time step includes: 22, 17, 15, 17, 15, 8, 10, 9.
[0136] S1005: The operation unit 211 determines the output results of the update gate and the reset gate in the first output data of the GRU network at the t-th time step according to the width values of the first weight matrix and the third weight matrix, or the width values of the second weight matrix and the fourth weight matrix.
[0137] For example, as Figure 8 shown, the width values of the first weight matrix W xz and the third weight matrix W xr are both 4. Therefore, the operation unit 211 uses the data from the 1st column to the 4th column of the first output data zr t ' as the output result z t ' of the update gate, and uses the data from the 5th column to the 8th column of the first output data as the output result r t ' of the reset gate. As Figure 8 shown, the output result z t ' of the update gate of the GRU network at the t-th time step is a matrix of size 1×4. The first row data of the output result z t ' of the update gate includes 22, 17, 15, 17. The output result r t ' of the reset gate is a matrix of size 1×4. The first row data of the output result z t ' of the reset gate includes 15, 8, 10, 9.
[0138] S1006: The operation unit 211 generates the gate value of the reset gate and the gate value of the update gate by running the non-linear operator of the GRU network at the t-th time step according to the output result of the reset gate and the output result of the update gate.
[0139] In some embodiments, the operation unit 211 determines the gate value of the update gate of the GRU network at the t-th time step by running the non-linear operator of the update gate of the GRU network at the t-th time step according to the output result of the update gate. Specifically, the gate value z of the update gate of the GRU network at the t-th time step t can be calculated by the following formula (6):
[0140] z t = σ(z' t ) (6)
[0141] In formula (6), σ(z' t ) represents the non-linear operator of the update gate of the GRU network at the t-th time step, σ() represents the sigmoid function, and the calculation formula of the sigmoid function can be expressed as
[0142] In some embodiments, the operation unit 211 determines the gate value of the reset gate of the GRU network at the t-th time step by running the non-linear operator of the reset gate of the GRU network at the t-th time step according to the output result of the reset gate. Specifically, the gate value z of the reset gate of the GRU network at the t-th time step t can be calculated by the following formula (7):
[0143] r t = σ(r t ') (7)
[0144] In formula (7), σ(r t ') represents the non-linear operator of the reset gate of the GRU network at the t-th time step, σ() represents the sigmoid function, and the calculation formula of the sigmoid function can be expressed as
[0145] S1007: The operation unit 211 determines the output data at the t-th time step according to the gate value of the reset gate at the t-th time step, the gate value of the update gate at the t-th time step, the input data at the t-th time step, and the output data at the (t - 1)-th time step.
[0146] In some embodiments, the arithmetic unit 211 determines the output data at the t-th time step based on the gate value of the reset gate, the gate value of the update gate, the input data at the t-th time step, and the output data at the (t - 1)-th time step at the t-th time step. Specifically, the output data at the t-th time step can be calculated by formulas (3) and (4). For the specific calculation process, refer to the descriptions of formulas (3) and (4), which will not be elaborated here.
[0147] From Figure 10 the process of the method for running the GRU model 10, it can be seen that during the process of the arithmetic unit 211 calculating the gate value of the reset gate and the gate value of the update gate of the GRU network at the t-th time step, the arithmetic unit 211 splices the first weight matrix and the third weight matrix along the width direction to generate the first spliced weight matrix. The second weight matrix and the fourth weight matrix are spliced along the width direction to generate the second spliced weight matrix. The arithmetic unit 211 runs the first operator, uses the first spliced weight matrix as the weight parameter of the input data at the t-th time step, and uses the second spliced weight matrix as the weight parameter of the output data at the (t - 1)-th time step for matrix operations, so that the arithmetic unit 211 can determine the gate value of the update gate and the gate value of the reset gate at the t-th time step by reading data from the storage unit 212 once. Thus, when the GRU model 10 processes a piece of speech to be recognized, compared with Figure 9 the arithmetic unit 211 in which the arithmetic unit 211 needs to read data from the storage unit 212 at least 2n times to recognize the text corresponding to the speech, the arithmetic unit 211 only needs to read data from the storage unit 212 n times to recognize the text corresponding to the speech, thereby reducing the data reading amount by half, improving the speed of the processor 21 for speech processing, and further improving the speed of speech recognition of the electronic device 20. The time for speech recognition becomes shorter, enhancing the user experience.
[0148] The embodiments disclosed in the present application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device.
[0149] The program code can be applied to the input instructions to execute the various functions described in the present application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of the present application, the processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0150] The program code can be implemented in a high-level procedural language or an object-oriented programming language to communicate with the processing system. When needed, the program code can also be implemented in assembly language or machine language. In fact, the mechanisms described in this application are not limited to the scope of any specific programming language. In any case, the language can be a compiled language or an interpreted language.
[0151] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments can also be implemented as instructions carried or stored on one or more transient or non-transitory machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, the instructions can be distributed via a network or via other computer-readable media. Thus, machine-readable media can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to, floppy disks, optical disks, optical discs, compact discs read-only memory (CD-ROMs), magneto-optical discs, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) in electrical, optical, acoustic, or other forms via the Internet. Thus, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).
[0152] In the drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or ordering may not be required. Rather, in some embodiments, these features may be arranged in a different manner and / or order than shown in the illustrative drawings. Additionally, including a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.
[0153] It should be noted that each unit / module mentioned in the device embodiments of the present application is a logical unit / module. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or can be implemented by a combination of multiple physical units / module. The physical implementation manner of these logical units / modules themselves is not the most important. The combination of the functions implemented by these logical units / modules is the key to solving the technical problems proposed by the present application. In addition, in order to highlight the innovative part of the present application, the above device embodiments of the present application do not introduce units / modules that are not closely related to solving the technical problems proposed by the present application. This does not mean that there are no other units / modules in the above device embodiments.
[0154] It should be noted that in the examples and the description of the present patent, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one" does not exclude the presence of another identical element in the process, method, article or device including the said element.
[0155] Although the present application has been illustrated and described by referring to some preferred embodiments of the present application, those of ordinary skill in the art should understand that various changes can be made to it in form and detail without departing from the spirit and scope of the present application.
Claims
1. A method for running a neural network model, applied to an electronic device, characterized in that, The neural network model is used for speech recognition, text translation, and image classification. Among them, the neural network model includes a first operation and a second operation; And the method includes: Obtain the data matrix to be operated on for the first operation or the second operation. Among them, the number of operation factors in the first operation and the second operation is the same, and for each operation factor in the first operation, there is a corresponding operation factor in the second operation with the same data matrix to be operated on but different operation coefficient matrices; Perform a third operation on the data matrix to be operated on to generate a result matrix of the third operation. Among them, the third operation is an operation method obtained by combining the operation coefficient matrices of the corresponding operation factors in the first operation and the second operation; Split the result matrix of the third operation to obtain the result matrix of the first operation and the result matrix of the second operation respectively.
2. The method according to claim 1, characterized in that, The operation coefficient matrices of the first operation and the second operation include two data dimensions: height and width. Among them, the height and width of the operation coefficient matrices of the corresponding operation factors in the first operation and the second operation are correspondingly equal; The third operation is an operation method obtained by combining the operation coefficient matrices of the corresponding operation factors in the first operation and the second operation along any data dimension direction.
3. The method according to claim 2, wherein The third operation is an operation method obtained by combining the operation coefficient matrices of the corresponding operation factors in the first operation and the second operation along the width direction.
4. The method according to claim 1, characterized in that The splitting of the result matrix of the third operation to obtain the result matrix of the first operation and the result matrix of the second operation respectively includes: The result matrix of the third operation includes two data dimensions: height and width; Split the result matrix of the third operation along any data dimension direction to obtain the result matrix of the first operation and the result matrix of the second operation respectively.
5. The method according to claim 4, characterized in that Split the result matrix of the third operation along the width direction to obtain the result matrix of the first operation and the result matrix of the second operation respectively.
6. The method according to claim 1, wherein The third operation being an operation method obtained by combining the operation coefficient matrices of the corresponding operation factors in the first operation and the second operation includes: The data matrix to be operated on includes a first input data matrix and a second input data matrix. The operation factors of the first operation and the second operation include the matrix product of the first input data matrix of the first operation and the second operation and the corresponding operation coefficient matrix, and the matrix product of the second input data matrix of the first operation and the second operation and the corresponding operation coefficient matrix; The third operation is an operation method obtained by combining the operation coefficient matrices corresponding to the first input data matrix of the first operation and the second operation, and combining the operation coefficient matrices corresponding to the second input data matrix of the first operation and the second operation.
7. The method according to claim 1, characterized in that The neural network model is a recurrent neural network model, and the first operation and the second operation are operations of the fully connected layer of the recurrent neural network model.
8. The method according to claim 7, characterized in that The neural network model includes at least one of the following: gated recurrent unit model, long short-term memory model.
9. An electronic device, characterized in that, The electronic device includes an arithmetic unit and a storage unit. Among them, the arithmetic unit runs a first matrix operation circuit; The arithmetic unit obtains a data matrix to be operated for the first operation or the second operation from the storage unit. Among them, the number of operation factors in the first operation and the second operation is the same, and for each operation factor in the first operation, there is a corresponding operation factor in the second operation with the same data matrix to be operated but different operation coefficient matrices; The arithmetic unit performs a third operation on the data matrix to be operated by running the first matrix operation circuit to generate a result matrix of the third operation. Among them, the third operation is an operation method obtained by combining the operation coefficient matrices of the corresponding operation factors in the first operation and the second operation. The first matrix operation circuit is used to generate the result matrix of the first operation and the result matrix of the second operation; The arithmetic unit splits the result matrix of the third operation to obtain the result matrix of the first operation and the result matrix of the second operation respectively.
10. An electronic device, characterized in that, Comprising: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the one or more processors of the electronic device, for executing the neural network model running method according to any one of claims 1 to 8.
11. A readable medium, characterized in that, Instructions are stored on a readable medium of the electronic device, and when the instructions are executed, the electronic device executes the neural network model running method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Data processing method and system
CN109766407A
Distributed placement of linear operators for accelerated deep learning
WO2021084506A1