Method for processing voice data, electronic device, and storage medium
By using weight common factors to round the weight transformation matrix during winograd transformation and inverse transformation, the problem of data calculation overflow in the quantization model of winograd fast convolution algorithm is solved, and the speed of speech data processing and the deployment efficiency of deep learning models are improved.
Patent Information
- Application Number
- CN202210761612.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-06-29
AI Technical Summary
In the prior art, winograd fast convolution algorithm is prone to data calculation overflow when quantizing the model to process speech data, affecting the deployment of deep learning models and the efficiency of speech data processing.
By obtaining the input quantization characteristics and weight transformation matrix of speech data, the weight transformation matrix is rounded using the weight common factor, the dynamic range of the integer data of the quantization model is adjusted, and corresponding processing is performed during winograd transformation and inverse transformation to prevent data calculation overflow.
It effectively prevents data calculation overflow, improves the speed of voice data processing and the deployment efficiency of deep learning models, and ensures calculation accuracy and processing speed.
Smart Images

Figure CN115171663B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method for processing voice data, an electronic device, and a storage medium. Background Art
[0002] Edge devices equipped with a deep learning model for processing voice data can perform tasks such as speech recognition or voiceprint recognition, realizing functions such as instant communication between machines and humans and identity recognition. During the deployment process of the deep learning model for processing voice data, a relatively high computing power cost is required, and problems such as poor inference real-time performance often occur on edge devices with insufficient computing power.
[0003] Currently, there is a problem that the winograd fast convolution algorithm is used to accelerate the convolution operation that occupies most of the time when processing voice data. By replacing the reduced number of multiplication operations with an increased number of addition operations, the time required for convolution calculation is shortened.
[0004] However, when the winograd fast convolution algorithm processes voice data in some quantization models, data calculation overflow often occurs, affecting the deployment of the deep learning model and the processing of voice data. Summary of the Invention
[0005] This application aims to at least solve one of the technical problems existing in the prior art. For this purpose, this application proposes a method for processing voice data to prevent data calculation overflow in the winograd fast convolution of voice data.
[0006] The method for processing voice data according to the first aspect embodiment of this application includes:
[0007] Obtain the input quantization features and weight transformation matrix of the voice data to be processed;
[0008] Obtain the common factor of the weights of the weight transformation matrix;
[0009] Based on the common factor of the weights, perform rounding processing on the weights in the weight transformation matrix to obtain a target weight transformation matrix;
[0010] Perform winograd transformation and matrix multiplication processing on the input quantization features and the target weight transformation matrix to obtain a target quantization matrix;
[0011] Perform winograd inverse transformation on the target quantization matrix;
[0012] Based on the common factor of the weights, process the output result of the winograd inverse transformation of the target quantization matrix to obtain the output quantization features of the voice data.
[0013] According to the method for processing voice data in an embodiment of the present application, the weight transformation matrix before Winograd transformation and the output result of Winograd inverse transformation are processed through a weight common factor to adjust the dynamic range of integer data in the quantization model, effectively prevent data calculation overflow, improve the processing speed of voice data, and accelerate the deployment of deep learning models related to voice data processing.
[0014] According to an embodiment of the present application, the process of rounding the weights in the weight transformation matrix based on the weight common factor to obtain a target weight transformation matrix includes:
[0015] Multiplying the weights in the weight transformation matrix by the weight common factor to obtain the target weight transformation matrix.
[0016] According to an embodiment of the present application, the process of processing the output result of Winograd inverse transformation of the target quantization matrix based on the weight common factor to obtain the output quantization feature of the voice data includes:
[0017] Dividing the output results of Winograd inverse transformation of the last column and the last row in the target quantization matrix by the weight common factor to obtain the output quantization feature.
[0018] According to an embodiment of the present application, the output result of Winograd inverse transformation of the target quantization matrix is stored in a data format of twice the integer type of the output result of Winograd inverse transformation.
[0019] According to an embodiment of the present application, the process of performing Winograd transformation and matrix multiplication on the input quantization feature and the target weight transformation matrix to obtain a target quantization matrix includes:
[0020] Performing Winograd transformation on the input quantization feature and the target weight transformation matrix respectively to obtain a transformed feature matrix and a transformed weight matrix;
[0021] Multiplying the transformed feature matrix and the transformed weight matrix according to batch matrix multiplication to obtain the target quantization matrix.
[0022] According to an embodiment of the present application, the process of multiplying the transformed feature matrix and the transformed weight matrix according to batch matrix multiplication to obtain the target quantization matrix includes:
[0023] Inputting the transformed feature matrix and the transformed weight matrix into a batch matrix multiplication function interface with a first parameter added to obtain the target quantization matrix output by the batch matrix multiplication function interface.
[0024] According to an embodiment of the present application, performing Winograd transforms on the input quantization feature and the target weight transform matrix respectively to obtain a transformed feature matrix and a transformed weight matrix includes:
[0025] Performing a Winograd transform based on the input quantization feature and the feature transform matrix of the input quantization feature to obtain the transformed feature matrix;
[0026] Performing a Winograd transform based on the transformed weight matrix and the quantization convolution kernel to obtain the transformed weight matrix.
[0027] A processing device for voice data according to an embodiment of the second aspect of the present application includes:
[0028] A first acquisition module, configured to acquire an input quantization feature and a weight transform matrix of voice data to be processed;
[0029] A second acquisition module, configured to acquire a weight common factor of the weight transform matrix;
[0030] A first processing module, configured to perform rounding processing on the weights in the weight transform matrix based on the weight common factor to obtain a target weight transform matrix;
[0031] A second processing module, configured to perform a Winograd transform and a matrix multiplication process on the input quantization feature and the target weight transform matrix to obtain a target quantization matrix;
[0032] A third processing module, configured to perform an inverse Winograd transform on the target quantization matrix;
[0033] A fourth processing module, configured to process the output result of the inverse Winograd transform of the target quantization matrix based on the weight common factor to obtain an output quantization feature of the voice data.
[0034] An electronic device according to an embodiment of the third aspect of the present application includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for processing voice data as described in any one of the above is implemented.
[0035] A non-transitory computer-readable storage medium according to an embodiment of the fourth aspect of the present application, on which a computer program is stored. When the computer program is executed by a processor, the method for processing voice data as described in any one of the above is implemented.
[0036] A computer program product according to an embodiment of the fifth aspect of the present application includes a computer program. When the computer program is executed by a processor, the method for processing voice data as described in any one of the above is implemented.
[0037] One or more of the above technical solutions in the embodiments of the present application have at least one of the following technical effects:
[0038] By using the weight common factor to process the weight transformation matrix before the winograd transformation and the output result of the winograd inverse transformation, the dynamic range of the integer data of the quantization model is adjusted, effectively preventing data calculation overflow, improving the processing speed of voice data, and accelerating the deployment of deep learning models related to voice data processing.
[0039] Furthermore, divide the output result of the winograd inverse transformation of the last column and the last row in the target quantization matrix by the weight common factor to adjust the dynamic range of the integer data of the target quantization matrix and prevent data calculation overflow in the output of the winograd inverse transformation.
[0040] Even further, the output result of the winograd inverse transformation of the target quantization matrix is stored in a double integer data format of the output result of the winograd inverse transformation, which can maintain the correctness of the calculation in the process of weight matrix conversion.
[0041] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0043] Figure 1 is one of the flow diagrams of the method for processing voice data provided by the embodiments of the present application;
[0044] Figure 2 is another flow diagram of the method for processing voice data provided by the embodiments of the present application;
[0045] Figure 3 is yet another flow diagram of the method for processing voice data provided by the embodiments of the present application;
[0046] Figure 4 is the structural diagram of the device for processing voice data provided by the embodiments of the present application;
[0047] Figure 5 is the structural diagram of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The following further describes in detail the implementation manners of the present application in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but cannot be used to limit the scope of the present application.
[0049] In the description of the embodiments of the present application, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0050] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0051] The following in conjunction with Figures 1 to 3 describes the processing method of the voice data of the embodiments of the present application. The execution subject of this method can be the controller of the terminal device, or the cloud, or the edge server.
[0052] As Figure 1 shown, the processing method of the voice data of the embodiments of the present application includes steps 110 to 160.
[0053] Step 110, obtain the input quantization feature and the weight transformation matrix of the voice data to be processed.
[0054] In this step, the voice data to be processed can be input into the quantization convolutional layer of the voice processing model for feature extraction to obtain the input quantization feature and the weight transformation matrix corresponding to the voice data to be processed.
[0055] It should be noted that generally, the calculations inside the model use floating-point calculations, and floating-point calculations consume relatively large computing resources. In this embodiment, the voice processing model can be a quantization model for processing voice data, and uses integer-type (int) data calculations, which greatly reduces the computing resources consumed.
[0056] In this embodiment, the input quantization feature of the voice data to be processed is integer-type feature data. In actual execution, the input quantization feature can be 8-bit integer-type data, that is, the input quantization feature is the data after int8 quantization.
[0057] It can be understood that the weight transformation matrix used for Winograd fast convolution processing of the input quantization features can also be data of 8-bit integer type.
[0058] Step 120: Obtain the common factor of the weights of the weight transformation matrix.
[0059] It can be understood that multiple weights required for convolution processing of the input quantization features are arranged according to the corresponding rows and columns to obtain the weight transformation matrix.
[0060] In this step, obtaining the common factor of the weights of the weight transformation matrix means obtaining the common factor of multiple weights in the weight transformation matrix.
[0061] It should be noted that the weights of the weight transformation matrix can be weights with decimals.
[0062] In actual implementation, methods for extracting common factors such as the maximum likelihood method, the principal axis iteration method, the weighted least squares method, the generalized weighted least squares method, and the minimum residual method can be used to obtain the common factor of the weights of the weight transformation matrix.
[0063] For example, the weight transformation matrix G is as follows
[0064]
[0065] In this embodiment, the common factor of the weights of the obtained weight transformation matrix G is 24.
[0066] Step 130: Based on the common factor of the weights, perform rounding processing on the weights in the weight transformation matrix to obtain the target weight transformation matrix.
[0067] In this step, according to the common factor of the weights of the weight transformation matrix, perform rounding processing on the weight transformation matrix to obtain a new weight transformation matrix, that is, the target weight transformation matrix.
[0068] It should be noted that when performing rounding processing on the weights of the weight transformation matrix using the common factor of the weights, the same processing is performed on all the weights in the weight transformation matrix, and only the decimal weights of the original weight transformation matrix are processed into integer weights, and the weight distribution of the weight transformation matrix remains unchanged.
[0069] For example, using the common factor of weights 24 to perform rounding processing on the weight transformation matrix G, the target weight transformation matrix G' is obtained as follows:
[0070]
[0071] Among them, both the weight transformation matrix G and the target weight transformation matrix G' are matrices with 6 rows and 3 columns, and the weight distribution of the weight transformation matrix G and the target weight transformation matrix G' remains unchanged.
[0072] In actual execution, before performing winograd fast convolution processing on the input quantization features and the weight transformation matrix in the quantization convolution layer of the speech processing model, a data preprocessing module is set up to perform rounding processing on the weights in the weight transformation matrix using the weight common factor.
[0073] Step 140: Perform winograd transformation and matrix multiplication processing on the input quantization features and the target weight transformation matrix to obtain the target quantization matrix.
[0074] In this step, perform winograd transformation on the input quantization features, perform matrix transformation and data rearrangement to implement the input transform in the winograd fast convolution algorithm; perform winograd transformation on the target weight transformation matrix, perform matrix transformation and data rearrangement to implement the filter transform in the winograd fast convolution algorithm.
[0075] Perform matrix multiplication processing on the matrices obtained by performing winograd transformation on the input quantization features and the target weight transformation matrix. Matrix multiplication can be performed through batched matrix multiplication (Batched-GEMM) to obtain the target quantization matrix.
[0076] Step 150: Perform winograd inverse transformation on the target quantization matrix.
[0077] In this step, perform winograd inverse transformation on the target quantization matrix, perform data rearrangement and matrix transformation to implement the output transform in the winograd fast convolution algorithm.
[0078] Step 160: Based on the weight common factor, process the output result of the winograd inverse transformation of the target quantization matrix to obtain the output quantization features of the speech data.
[0079] In actual execution, after performing batched matrix multiplication on the input quantization features and the target weight transformation matrix in the quantization convolution layer of the speech processing model, a data postprocessing module is set up to prevent data calculation overflow when implementing the output transform in the winograd fast convolution algorithm using the weight common factor.
[0080] For example, the input quantization features can be data of 8-bit integer type (int8), and the target weight transformation matrix is also data of 8-bit integer type (int8).
[0081] Perform matrix multiplication on the matrix obtained by performing the Winograd transform on the input quantization feature and the target weight transform matrix, execute int16 * int16 calculation, and perform the Winograd inverse transform on the target quantization matrix, which is a process of converting the data in the target quantization matrix into int16 data for output.
[0082] A voice signal is an analog signal whose amplitude changes continuously over time. Voice data is a digital signal that is discrete in both time and amplitude and is obtained by converting the voice signal. Operating on voice data quickly and efficiently to obtain the expected results is the key to voice data processing.
[0083] In actual execution, through integer-based data calculations, the computational resources consumed can be reduced and the processing speed of voice data can be accelerated.
[0084] In related technologies, when using the Winograd fast convolution algorithm to process voice data in a quantization model, the range of change of the integer data corresponding to time and amplitude in the voice data is large, and the dynamic range of the integer data of the quantization model is limited. Data calculation overflow often occurs, affecting the deployment of the deep learning model and the processing of voice data. In this embodiment, through the weight common factor, the int16 data output of the Winograd inverse transform of the target quantization matrix is processed to adjust the dynamic range of the integer data of the quantization model, effectively preventing data calculation overflow, and obtaining the output quantization feature of the finally output voice data. By extracting the common factor, a Winograd fast convolution algorithm for preventing data overflow is realized, which can be effectively directly deployed in a voice processing model in an actual production environment. The specific steps for accelerating the deployment are as follows:
[0085] As Figure 3 shown, Step 1: Deploy and load the voice processing model.
[0086] Step 2: Determine whether the convolutional layer of the voice processing model is a quantization convolutional layer, that is, determine whether the convolutional layer can process integer-based feature data.
[0087] Step 3: Determine whether the acceleration conditions of the Winograd fast convolution algorithm are met.
[0088] Step 4: According to the width and height of the input quantization feature, determine the feature transform matrix of the input quantization feature of the Winograd fast convolution.
[0089] Step 5: According to the width and height of the input quantization feature, determine the number of output results, and select the corresponding F(4x4, 3x3) Winograd int8 fast convolution or F(2x2, 3x3) Winograd int8 fast convolution.
[0090] Among them, F(4x4, 3x3) winograd int8 means that 8-bit integer data uses the winograd fast convolution algorithm, 4x4 means the output result is 16, and 3x3 means the convolution kernel of the convolution layer is a 3x3 convolution kernel.
[0091] F(2x2, 3x3) winograd int8 means that 8-bit integer data uses the winograd fast convolution algorithm, 2x2 means the output result is 4, and 3x3 means the convolution kernel of the convolution layer is a 3x3 convolution kernel.
[0092] The following introduces a specific embodiment.
[0093] A data preprocessing module and a data postprocessing module are set on the quantization convolution layer of the speech processing model for realizing voice wake-up, and are deployed on a terminal device.
[0094] Performing fast convolution processing of winograd can effectively solve the problem that the inference time of voice wake-up is too long and voice data frames are lost due to inability to process in real time. For example, from the issuance of the wake-up keyword instruction "Xiaomei Xiaomei" to the terminal device's feedback of "You said", the overall voice link delay is reduced from 18 milliseconds to within 9 milliseconds, and the overall inference performance of the speech processing model is doubled. According to the method for processing speech data provided by the embodiments of the present application, the weight transformation matrix before winograd transformation and the output result of winograd inverse transformation are processed through the weight common factor, the dynamic range of the integer data of the quantization model is adjusted, data calculation overflow is effectively prevented, the processing speed of speech data is improved, and the deployment of deep learning models related to speech data processing is accelerated.
[0095] Step 130 of the above method for processing speech data includes:
[0096] Multiply the weights in the weight transformation matrix by the weight common factor to obtain a target weight transformation matrix.
[0097] In this embodiment, all the weights in the weight transformation matrix are multiplied by the weight common factor of the weight transformation matrix to obtain a target weight transformation matrix, and all the weights in the target weight transformation matrix are integers.
[0098] For example, the weight transformation matrix G is as follows
[0099]
[0100] In this embodiment, the weight common factor of the obtained weight transformation matrix G is 24.
[0101] Multiply all the weights of the weight transformation matrix G by the weight common factor 24 to obtain a target weight transformation matrix G', as shown below:
[0102]
[0103] Among them, after the 0 weight value is multiplied by the weight common factor, it is still the 0 weight value.
[0104] In the related art, by rounding the fractional weight values of the weight transformation matrix before the Winograd transformation to integer weight values, this method will change the weight distribution of the weight transformation matrix and reduce the calculation accuracy of the Winograd fast convolution processing.
[0105] In this embodiment, before performing the Winograd transformation, the non-divisible fractional weight values in the weight transformation matrix are preprocessed into corresponding integer weight values by multiplying by the weight common factor. All the weight values in the weight transformation matrix are multiplied by the weight common factor, and the weight distribution of the target weight transformation matrix and the weight transformation matrix remains unchanged, ensuring the calculation accuracy of the Winograd fast convolution processing.
[0106] Step 160 of the above method for processing voice data includes:
[0107] Dividing the output results of the Winograd inverse transformation of the last column and the last row of the target quantization matrix by the weight common factor to obtain the output quantization features.
[0108] The input quantization features and the target weight transformation matrix are multiplied in batches to obtain the target quantization matrix, and the data in the target quantization matrix is subjected to the Winograd inverse transformation to obtain the corresponding output results of the Winograd inverse transformation.
[0109] Before the output results of the Winograd inverse transformation of the last column and the last row in the target quantization matrix are output, dividing the output results of the Winograd inverse transformation of the last column and the last row of the target quantization matrix by the weight common factor to adjust the dynamic range of the integer data in the target quantization matrix and prevent data calculation overflow in the output of the Winograd inverse transformation.
[0110] For example, the input quantization features can be 8-bit integer data (int8), and the target weight transformation matrix is also 8-bit integer data (int8).
[0111] The matrix obtained by performing the Winograd transformation on the input quantization features and the target weight transformation matrix is multiplied, and the int16*int16 calculation is performed. The Winograd inverse transformation of the target quantization matrix is a process of converting the data in the target quantization matrix into int16 data for output.
[0112] Before the data in the last column and the last row of the target quantization matrix are output as int16 data, divide the data in the last column and the last row by the weight common factor to prevent data calculation overflow and obtain the output quantization features of the voice data.
[0113] In this embodiment, the output result of the winograd inverse transform of the target quantization matrix is stored in a data format of twice the integer type of the output result of the winograd inverse transform.
[0114] For example, the input quantization features and the matrix obtained by performing the winograd transform on the target weight transform matrix are multiplied through matrix multiplication, performing int16*int16 calculations, and the output result of the winograd inverse transform of the target quantization matrix is stored in the int32 data format.
[0115] Among them, int32 is a data format of twice the integer type of int16.
[0116] In this embodiment, the output result of the winograd inverse transform of the target quantization matrix is the calculation result of int16, and the calculation result of int16 is stored in the int32 data format, which can maintain the correctness of the calculation during the conversion process of the weight matrix.
[0117] Step 140 of the above voice data processing method includes:
[0118] Perform winograd transforms on the input quantization features and the target weight transform matrix respectively to obtain the transformed feature matrix and the transformed weight matrix;
[0119] Multiply the transformed feature matrix and the transformed weight matrix according to the batch matrix multiplication to obtain the target quantization matrix.
[0120] In this embodiment, perform a winograd transform on the input quantization features, perform matrix transformation and data rearrangement to obtain the transformed feature matrix, and implement the input transformation in the winograd fast convolution algorithm.
[0121] In actual execution, perform a winograd transform based on the input quantization features and the feature transform matrix of the input quantization features to obtain the transformed feature matrix.
[0122] In this embodiment, perform a winograd transform on the target weight transform matrix, perform matrix transformation and data rearrangement to obtain the transformed weight matrix, and implement the weight transformation in the winograd fast convolution algorithm.
[0123] In actual execution, perform a winograd transform based on the transformed weight matrix and the quantization convolution kernel to obtain the transformed weight matrix.
[0124] The calculation formula of the Winograd fast convolution algorithm is as follows:
[0125] Y = A T [[GgG T ⊙ [B T dB]]A
[0126] Among them, d is the input quantization feature, g is the quantization convolution kernel, B is the feature transformation matrix, G is the weight transformation matrix, G is the output feature inverse transformation matrix, and Y is the output quantization feature after Winograd fast convolution calculation.
[0127] ⊙ is the Hadamard product, indicating that the corresponding positions of the matrices are multiplied. The batch matrix multiplication is realized by the matrix operation method of the Hadamard product for the transformed feature matrix and the transformed weight matrix.
[0128] In this embodiment, [GgG T represents the transformed weight matrix, [B T dB] represents the transformed feature matrix; A is the output feature inverse transformation matrix, [[GgG T ⊙ [B T dB]] and A and A T The calculation process of is the output process of the Winograd inverse transformation.
[0129] The following introduces a specific embodiment.
[0130] As Figure 2 shown, d is the input quantization feature, g is the quantization convolution kernel, B is the feature transformation matrix, and G is the weight transformation matrix.
[0131] By extracting the common factor of the weights of the weight transformation matrix G and performing matrix parameter preprocessing, the non-divisible decimal weights in the weight transformation matrix G are preprocessed into corresponding integer weights by multiplying by the common factor of the weights, and the target weight transformation matrix G' is obtained.
[0132] Based on the input quantization feature d and the feature transformation matrix B of the input quantization feature d, perform Winograd transformation, perform matrix transformation and data rearrangement to obtain the transformed feature matrix.
[0133] Based on the target weight transformation matrix G' and the quantization convolution kernel g, perform Winograd transformation, perform matrix transformation and data rearrangement to obtain the transformed weight matrix.
[0134] Perform batch matrix multiplication on the transformed feature matrix and the transformed weight matrix to obtain the target quantization matrix.
[0135] Perform the Winograd inverse transform on the target quantization matrix, perform data rearrangement, divide the output results of the Winograd inverse transform in the last column and the last row of the target quantization matrix by the weight common factor, adjust the integer data dynamic range of the target quantization matrix, and perform matrix transformation to obtain the output quantization feature Y.
[0136] In some embodiments, input the transformed feature matrix and the transformed weight matrix into the batch matrix multiplication function interface with the addition of a first parameter to obtain the target quantization matrix output by the batch matrix multiplication function interface.
[0137] In this embodiment, when performing the batch matrix multiplication process of the transformed feature matrix and the transformed weight matrix, add a first parameter to the commonly used batch matrix multiplication function interface. That is, inside the function of batched general matrix multiplication (Batched-GEMM), before saving the calculation result of multiplying the corresponding positions of the transformed feature matrix and the transformed weight matrix, multiply by the first parameter first.
[0138] It can be understood that at the external call of the batch matrix multiplication function, the matrix index can be judged through the first parameter. Among them, when judging the matrix index, the first parameter can be passed as 1 or the weight common factor, and the matrix coefficient removed by the weight common factor during the preprocessing of rounding the weight transformation matrix is added back by performing an inverse operation during the batch matrix multiplication calculation.
[0139] For example, the input quantization feature can be 8-bit integer data (int8), and the target weight transformation matrix is also 8-bit integer data (int8).
[0140] Perform matrix multiplication on the matrices obtained by performing the Winograd transform on the input quantization feature and the target weight transformation matrix, execute int16*int16 calculation, add a first parameter to the standard int16*int16 batch matrix multiplication function interface, and in the batch matrix multiplication function, multiply by the first parameter before saving the int32 calculation result.
[0141] Next, the processing device for voice data provided by the embodiments of the present application will be described. The processing device for voice data described below can be correspondingly referred to the processing method for voice data described above.
[0142] As Figure 4 shown, the processing device for voice data provided by the embodiments of the present application includes:
[0143] The first acquisition module 410 is configured to acquire the input quantization feature and the weight transformation matrix of the voice data to be processed;
[0144] The second acquisition module 420 is configured to acquire the weight common factor of the weight transformation matrix;
[0145] The first processing module 430 is configured to perform rounding processing on the weights in the weight transformation matrix based on the weight common factor to obtain a target weight transformation matrix;
[0146] The second processing module 440 is configured to perform winograd transformation and matrix multiplication processing on the input quantization feature and the target weight transformation matrix to obtain a target quantization matrix;
[0147] The third processing module 450 is configured to perform winograd inverse transformation on the target quantization matrix;
[0148] The fourth processing module 460 is configured to process the output result of the winograd inverse transformation of the target quantization matrix based on the weight common factor to obtain the output quantization feature of the voice data.
[0149] According to the voice data processing device provided by the embodiments of the present application, the weight transformation matrix before winograd transformation and the output result of winograd inverse transformation are processed by the weight common factor to adjust the dynamic range of the integer data of the quantization model, effectively prevent data calculation overflow, improve the processing speed of voice data, and accelerate the deployment of deep learning models related to voice data processing.
[0150] In some embodiments, the first processing module 430 is configured to multiply the weights in the weight transformation matrix by the weight common factor to obtain a target weight transformation matrix.
[0151] In some embodiments, the fourth processing module 460 is configured to divide the output result of the winograd inverse transformation of the last column and the last row in the target quantization matrix by the weight common factor to obtain the output quantization feature.
[0152] In some embodiments, the output result of the winograd inverse transformation of the target quantization matrix is stored in a double integer data format of the output result of the winograd inverse transformation.
[0153] In some embodiments, the second processing module 440 is configured to perform winograd transformation on the input quantization feature and the target weight transformation matrix respectively to obtain a transformed feature matrix and a transformed weight matrix;
[0154] Multiply the transformed feature matrix and the transformed weight matrix according to batch matrix multiplication to obtain a target quantization matrix.
[0155] In some embodiments, the second processing module 440 is configured to input the transformed feature matrix and the transformed weight matrix into the batch matrix multiplication function interface with the first parameter added to obtain the target quantization matrix output by the batch matrix multiplication function interface.
[0156] In some embodiments, the second processing module 440 is configured to perform a Winograd transform based on the input quantization feature and the feature transformation matrix of the input quantization feature to obtain a transformed feature matrix;
[0157] Perform a Winograd transform based on the transformed weight matrix and the quantization convolution kernel to obtain a transformed weight matrix.
[0158] Figure 5 An example of the physical structure diagram of an electronic device is shown as Figure 5 shown. The electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call the logical instructions in the memory 530 to execute the processing method of voice data. The method includes: obtaining the input quantization feature and the weight transformation matrix of the voice data to be processed; obtaining the common factor of the weights of the weight transformation matrix; based on the common factor of the weights, performing rounding processing on the weights in the weight transformation matrix to obtain a target weight transformation matrix; performing a Winograd transform and a matrix multiplication process on the input quantization feature and the target weight transformation matrix to obtain a target quantization matrix; performing an inverse Winograd transform on the target quantization matrix; based on the common factor of the weights, processing the output result of the inverse Winograd transform of the target quantization matrix to obtain the output quantization feature of the voice data.
[0159] In addition, when the logical instructions in the above-mentioned memory 530 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0160] Further, the present application also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the voice data processing method provided in each of the above method embodiments. The method includes: obtaining an input quantization feature and a weight transformation matrix of the voice data to be processed; obtaining a weight common factor of the weight transformation matrix; based on the weight common factor, performing rounding processing on the weights in the weight transformation matrix to obtain a target weight transformation matrix; performing winograd transformation and matrix multiplication processing on the input quantization feature and the target weight transformation matrix to obtain a target quantization matrix; performing winograd inverse transformation on the target quantization matrix; and based on the weight common factor, processing the output result of the winograd inverse transformation of the target quantization matrix to obtain an output quantization feature of the voice data.
[0161] On the other hand, an embodiment of the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the voice data processing method provided in each of the above embodiments. The method includes: obtaining an input quantization feature and a weight transformation matrix of the voice data to be processed; obtaining a weight common factor of the weight transformation matrix; based on the weight common factor, performing rounding processing on the weights in the weight transformation matrix to obtain a target weight transformation matrix; performing winograd transformation and matrix multiplication processing on the input quantization feature and the target weight transformation matrix to obtain a target quantization matrix; performing winograd inverse transformation on the target quantization matrix; and based on the weight common factor, processing the output result of the winograd inverse transformation of the target quantization matrix to obtain an output quantization feature of the voice data.
[0162] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0163] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, and this computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
[0165] The above embodiments are only used to illustrate the present application, rather than to limit the present application. Although the present application has been described in detail with reference to the embodiments, those of ordinary skill in the art should understand that various combinations, modifications or equivalent replacements of the technical solutions of the present application do not deviate from the spirit and scope of the technical solutions of the present application, and should all be covered by the scope of the claims of the present application.
Claims
1. A method for processing voice data, characterized in that, Including: Obtaining an input quantization feature of speech data to be processed and a weight transformation matrix; Obtaining a weight common factor of the weight transformation matrix; Multiplying the weights in the weight transformation matrix by the weight common factor to obtain a target weight transformation matrix; wherein, all weights in the target weight transformation matrix are integers; the weight distribution of the target weight transformation matrix is the same as that of the weight transformation matrix; Performing a Winograd transform and a matrix multiplication process on the input quantization feature and the target weight transformation matrix to obtain a target quantization matrix; Performing an inverse Winograd transform on the target quantization matrix; Dividing the output results of the inverse Winograd transform of the last column and the last row in the target quantization matrix by the weight common factor to obtain an output quantization feature.
2. The method for processing voice data according to claim 1, wherein The output result of the inverse Winograd transform of the target quantization matrix is stored in a double-type data format of the output result of the inverse Winograd transform.
3. The method for processing voice data according to claim 1 or 2, characterized in that, The performing a Winograd transform and a matrix multiplication process on the input quantization feature and the target weight transformation matrix to obtain a target quantization matrix includes: Performing a Winograd transform on the input quantization feature and the target weight transformation matrix respectively to obtain a transformed feature matrix and a transformed weight matrix; Multiplying the transformed feature matrix and the transformed weight matrix according to a batch matrix multiplication to obtain the target quantization matrix.
4. The method for processing voice data according to claim 3, wherein The multiplying the transformed feature matrix and the transformed weight matrix according to a batch matrix multiplication to obtain the target quantization matrix includes: Inputting the transformed feature matrix and the transformed weight matrix into a batch matrix multiplication function interface with a first parameter added to obtain the target quantization matrix output by the batch matrix multiplication function interface.
5. The method for processing voice data according to claim 3, wherein The performing a Winograd transform on the input quantization feature and the target weight transformation matrix respectively to obtain a transformed feature matrix and a transformed weight matrix includes: Performing a Winograd transform based on the input quantization feature and a feature transformation matrix of the input quantization feature to obtain the transformed feature matrix; Performing a Winograd transform based on the transformed weight matrix and a quantization convolution kernel to obtain the transformed weight matrix.
6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for processing speech data according to any one of claims 1 to 5.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for processing speech data according to any one of claims 1 to 5.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for processing speech data according to any one of claims 1 to 5.
Citation Information
Patent Citations
Convolutional neural network calculation method and device
CN111260020A
Convolutional neural network processing method and device, equipment and storage medium
CN111382854A