Speech Recognition Method, Device, Electronic Device, Storage Medium and Program Product
By introducing semi-supervised training and knowledge distillation technology into the speech recognition model, the unlabeled and marked speech data sets are used to optimize the model, and the contradiction between speech recognition efficiency and accuracy is solved, and efficient and high-performance speech recognition effects are achieved.
Patent Information
- Application Number
- CN202210443950.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-26
AI Technical Summary
While improving efficiency, the existing speech recognition models have large accuracy losses and are difficult to meet high performance needs.
By introducing the unlabeled first sample speech dataset and the marked second sample speech dataset, the semi-supervised training method is used to initialize the original model for unsupervised training, and supervised training is performed after pruning, and combining knowledge distillation technology, the speech recognition model is gradually optimized.
While improving the efficiency of speech recognition, it significantly improves the accuracy of speech recognition, realizes the reduction of model size and the reduction of deployment cost, and improves the overall performance of the model.
Smart Images

Figure CN115132181B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of speech recognition, and particularly to a speech recognition method, device, electronic device, storage medium and program product. Background Art
[0002] With the rapid development of Internet technology, automatic speech recognition technology has been widely applied in people's daily life and work, such as instant messaging speech recognition, smart home speech recognition, in-vehicle system speech recognition, and so on. Automatic speech recognition technology generally realizes based on a speech recognition model. With the improvement of the accuracy requirement of speech recognition, the performance requirements for speech recognition in various application scenarios are getting higher and higher, and the volume of the speech recognition model also increases accordingly, thus reducing the speech recognition efficiency.
[0003] In order to improve the speech recognition efficiency, knowledge distillation can be used to train the speech recognition model, aiming to reduce the volume of the speech recognition model and improve the speech recognition efficiency. However, for the speech recognition model obtained by using the distillation training method in the related technology, although the effect of reducing the model volume can be achieved, there is still a certain gap in performance between the speech recognition model obtained after distillation training and the original speech recognition model, thus reducing the speech recognition accuracy. Therefore, how to improve the speech recognition accuracy while improving the speech recognition efficiency is an urgent problem to be solved currently. Summary of the Invention
[0004] The following is an overview of the subject matter described in detail in this document. This overview is not intended to limit the scope of protection of the claims.
[0005] Embodiments of the present application provide a speech recognition method, device, electronic device, storage medium and program product, which can improve the speech recognition accuracy while improving the speech recognition efficiency.
[0006] On the one hand, embodiments of the present application provide a speech recognition method, including:
[0007] Obtaining an unlabeled first sample speech data set and a labeled second sample speech data set;
[0008] Initializing an original model, and performing unsupervised training on the original model based on the first sample speech data set to obtain a basic speech processing model; wherein, the basic speech processing model includes a plurality of sequentially connected basic encoding layers;
[0009] Pruning the basic encoding layers after a preset encoding layer among the plurality of basic encoding layers, and performing supervised training on the pruned basic speech processing model based on the second sample speech data set to obtain a first speech recognition model;
[0010] Initialize the second speech recognition model, and based on the first speech recognition model as a reference, perform distillation training on the second speech recognition model based on the first sample speech data set to obtain a target speech recognition model;
[0011] Perform speech recognition on the target speech data based on the target speech recognition model to obtain a target recognition result corresponding to the target speech data.
[0012] On the other hand, an embodiment of the present application also provides a speech recognition device, including:
[0013] A sample data acquisition module, configured to acquire an unlabeled first sample speech data set and a labeled second sample speech data set;
[0014] A first training module, configured to initialize an original model, and perform unsupervised training on the original model based on the first sample speech data set to obtain a basic speech processing model; wherein, the basic speech processing model includes a plurality of sequentially connected basic encoding layers;
[0015] A second training module, configured to prune the basic encoding layers located after a preset encoding layer among the plurality of basic encoding layers, and perform supervised training on the pruned basic speech processing model based on the second sample speech data set to obtain a first speech recognition model;
[0016] A third training module, configured to initialize a second speech recognition model, and based on the first speech recognition model as a reference, perform distillation training on the second speech recognition model based on the first sample speech data set to obtain a target speech recognition model;
[0017] A speech recognition module, configured to perform speech recognition on the target speech data based on the target speech recognition model to obtain a target recognition result corresponding to the target speech data.
[0018] Further, the first speech recognition model includes a first encoding network and a first output layer connected to each other, and the first encoding network includes a plurality of sequentially connected first encoding layers; the second speech recognition model includes a second encoding network and a second output layer connected to each other, and the second encoding network includes a plurality of sequentially connected second encoding layers, and the number of the second encoding layers is less than the number of the first encoding layers. Specifically, the third training module is configured to:
[0019] Randomly initialize the encoding parameters of each of the second encoding layers;
[0020] Use the output parameters of the first output layer as the output parameters of the second output layer;
[0021] Initialize the second speech recognition model according to the encoding parameters of each of the second encoding layers and the output parameters of each of the second output layers.
[0022] Further, the first speech recognition model further includes a first convolutional network connected to the first encoding network, and the first convolutional network includes a plurality of first convolutional layers connected in sequence; the second speech recognition model further includes a second convolutional network connected to the second encoding network, and the second convolutional network includes a plurality of second convolutional layers connected in sequence. The number of the second convolutional layers is equal to the number of the first convolutional layers. The second convolutional layer before the preset convolutional layer is the target convolutional layer, and the feature dimension of the target convolutional layer is smaller than the feature dimension of the corresponding first convolutional layer of the target convolutional layer. The specific operations of the third training module are as follows:
[0023] Randomly initialize the convolutional parameters of the preset convolutional layer and the convolutional parameters of the target convolutional layer;
[0024] Use the convolutional parameters of the corresponding first convolutional layer of the remaining convolutional layers as the convolutional parameters of the remaining convolutional layers; wherein, the remaining convolutional layers are the second convolutional layers other than the preset convolutional layer and the target convolutional layer among the plurality of second convolutional layers;
[0025] Initialize the second speech recognition model according to the convolutional parameters of each of the second convolutional layers, the encoding parameters of each of the second encoding layers, and the output parameters of each of the second output layers.
[0026] Further, the specific operations of the third training module are as follows:
[0027] Input the first sample speech data set into the first speech recognition model, obtain the first convolutional features output by the first convolutional layer corresponding to the preset convolutional layer, and obtain the second convolutional features output by the last first convolutional layer;
[0028] Input the first sample speech data set into the second speech recognition model, obtain the third convolutional features output by the preset convolutional layer, and obtain the fourth convolutional features output by the last second convolutional layer;
[0029] Determine a first convolutional loss value according to the first convolutional features and the third convolutional features, and determine a second convolutional loss value according to the second convolutional features and the fourth convolutional features;
[0030] Determine a target convolutional loss value according to the first convolutional loss value and the second convolutional loss value, and perform distillation training on the second convolutional network according to the target convolutional loss value.
[0031] Further, the specific operations of the third training module are as follows:
[0032] Input the first sample voice dataset into the first voice recognition model, determine the same number of benchmark coding layers as the second coding layer from the first voice recognition model, and obtain the first coding features output by each of the benchmark coding layers;
[0033] Input the first sample voice dataset into the second voice recognition model, and obtain the second coding features output by each of the second coding layers;
[0034] Determine the distillation fitting parameters of each of the second coding layers, and adjust the feature dimensions of the corresponding second coding features according to the distillation fitting parameters;
[0035] According to the second coding features with adjusted feature dimensions and the corresponding first coding features, determine the coding layer loss values corresponding to each of the second coding layers, determine the target coding loss value according to each of the coding layer loss values, and perform distillation training on the second coding network according to the target coding loss value.
[0036] Furthermore, the above-mentioned third training module is specifically used for:
[0037] Obtain the first sample recognition result output by the first output layer and the second sample recognition result output by the second output layer;
[0038] Determine the target output loss value according to the first sample recognition result and the second sample recognition result;
[0039] Perform distillation training on the second coding network according to the target coding loss value and the target output loss value.
[0040] Furthermore, the above-mentioned third training module is specifically used for:
[0041] Perform distillation training on the second coding network according to the target coding loss value, and perform distillation training on the second coding network again according to the target output loss value;
[0042] Alternatively, weight the target coding loss value and the target output loss value to obtain a target model loss value, and perform distillation training on the second coding network according to the target model loss value.
[0043] Furthermore, the above-mentioned third training module is specifically used for:
[0044] Adjust the feature dimensions of the second output layer in the target voice recognition model according to the distillation fitting parameters of the last second coding layer;
[0045] Perform voice recognition on the target voice data based on the target voice recognition model with adjusted feature dimensions.
[0046] Further, the third training module is specifically configured to:
[0047] Prune the distillation fitting parameters of each target encoding layer in the target speech recognition model, and perform speech recognition on target speech data based on the pruned target speech recognition model; wherein, the target encoding layer is the second encoding layer except for the last second encoding layer;
[0048] Alternatively, prune the distillation fitting parameters of each second encoding layer in the target speech recognition model, and perform speech recognition on target speech data based on the pruned target speech recognition model.
[0049] Further, the original model includes an original convolutional network and an original encoding network connected in sequence, and the first training module is specifically configured to:
[0050] Input the first sample speech data set into the original model, perform a masking operation on the original convolutional features output by the original convolutional network to obtain masked convolutional features;
[0051] Perform product quantization operation on the original convolutional features to obtain quantized convolutional features;
[0052] Obtain the masked encoding features output by the original encoding network after processing the masked convolutional features;
[0053] Determine a first original loss value according to the masked encoding features and the quantized convolutional features;
[0054] Perform unsupervised training on the original model according to the first original loss value.
[0055] Further, the first training module is specifically configured to:
[0056] Obtain a first quantity of the quantization codebooks and a second quantity of the cluster centers in each quantization codebook when performing the product quantization operation;
[0057] Determine the probability distribution of any one of the cluster centers in each quantization codebook being selected;
[0058] Determine a second original loss value according to the first quantity, the second quantity and the probability distribution;
[0059] Perform unsupervised training on the original model according to the first original loss value and the second original loss value.
[0060] On the other hand, an embodiment of the present application further provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned speech recognition method is implemented.
[0061] On the other hand, an embodiment of the present application further provides a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned speech recognition method is implemented.
[0062] On the other hand, an embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes to implement the above-mentioned speech recognition method.
[0063] The embodiments of the present application at least include the following beneficial effects: By introducing an unlabeled first sample speech data set and a labeled second sample speech data set, and using different types of sample speech data sets at different training stages, the effect of semi-supervised training can be achieved. Moreover, since the first sample speech data set does not need to be labeled, the cost of data labeling can be reduced, the training efficiency of the model can be improved, and it is beneficial to improve the performance of the target speech recognition model. And, since the features output by the basic coding layer closer to the end are more likely to fit the training task and will affect the calculation of the loss value, pruning the basic coding layer behind the preset coding layer in the basic speech processing model can improve the training effect of the basic speech processing model, improve the performance of the basic speech processing model, and at the same time can also reduce the volume of the basic speech processing model, and correspondingly reduce the volume of the first speech recognition model. Subsequently, when distilling and training the second speech recognition model based on the first speech recognition model, the volume of the target speech recognition model can be further reduced, and the deployment cost of the target speech recognition model can be reduced. And, since the performance of the basic speech processing model is improved, the performance of the target speech recognition model can also be further improved. It can be seen that the target speech recognition model provided by the embodiments of the present application has the advantages of small volume and high performance. Therefore, using the target speech recognition model to perform speech recognition on target speech data can improve the speech recognition accuracy while improving the speech recognition efficiency.
[0064] Other features and advantages of the present application will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] The accompanying drawings are used to provide a further understanding of the technical solutions of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solutions of the present application, and do not constitute a limitation to the technical solutions of the present application.
[0066] Figure 1 A schematic diagram of an implementation environment provided for an embodiment of the present application;
[0067] Figure 2 A schematic diagram of another implementation environment provided for an embodiment of the present application;
[0068] Figure 3 A schematic flowchart of a speech recognition method provided for an embodiment of the present application;
[0069] Figure 4 An exemplary structural diagram of a basic speech processing model provided for an embodiment of the present application;
[0070] Figure 5 An exemplary structural diagram of a basic encoding layer provided for an embodiment of the present application;
[0071] Figure 6 A schematic flowchart of unsupervised training of the original model provided for an embodiment of the present application;
[0072] Figure 7 An exemplary structural diagram of a first speech recognition model provided for an embodiment of the present application;
[0073] Figure 8 An exemplary structural diagram of a second speech recognition model provided for an embodiment of the present application;
[0074] Figure 9 A schematic diagram of distillation training of a second convolutional network provided for an embodiment of the present application;
[0075] Figure 10 A schematic diagram of distillation training of a second encoding network provided for an embodiment of the present application;
[0076] Figure 11 A general training flowchart of a target speech recognition model provided for an embodiment of the present application;
[0077] Figure 12 A detailed training flowchart of a target speech recognition model provided for an embodiment of the present application;
[0078] Figure 13 A structural diagram of a speech recognition device provided for an embodiment of the present application;
[0079] Figure 14 A partial structural block diagram of a terminal provided for an embodiment of the present application;
[0080] Figure 15 This is a partial structural block diagram of the server provided by the embodiment of the present application. Detailed implementation manners
[0081] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0082] To facilitate the understanding of the technical solutions provided by the embodiments of the present application, some key terms used in the embodiments of the present application will be explained here first:
[0083] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0084] Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.
[0085] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0086] Knowledge Distillation (KD): Also known as dark knowledge extraction, it refers to the process of guiding the training of a relatively simple and computationally less-intensive student network through a teacher network with a complex structure and large computational requirements but excellent performance, in order to improve the performance of the student network and achieve knowledge transfer. Knowledge distillation can make the model lightweight (facilitating deployment) while minimizing performance loss.
[0087] Automatic Speech Recognition (ASR): It is an active research topic in the field of artificial intelligence. The purpose of automatic speech recognition is to convert speech signals into corresponding text representations.
[0088] Currently, automatic speech recognition technology is generally implemented based on speech recognition models. With the increasing demand for the accuracy of speech recognition, the performance requirements for speech recognition in various application scenarios are getting higher and higher, and the volume of speech recognition models is also increasing, thus reducing the speech recognition efficiency. To improve the speech recognition efficiency, knowledge distillation can be used to train the speech recognition model, aiming to reduce the volume of the speech recognition model and improve the speech recognition efficiency. Among them, the training method of knowledge distillation can transfer the knowledge embedded in a large teacher network to a small student network, and then train the small student network to reproduce the behavior of the teacher network, thereby improving the convenience of model deployment. However, the number of parameters of the student network is smaller than that of the teacher network, so there is a problem of model accuracy loss, that is, the speech recognition model obtained by using the distillation training method in related technologies, although it can achieve the effect of reducing the model volume, there is still a certain gap in performance between the speech recognition model obtained after distillation training and the original speech recognition model, thus reducing the accuracy of speech recognition.
[0089] Based on this, the embodiments of the present application provide a speech recognition method, device, electronic device, storage medium and program product, which can improve the speech recognition accuracy while improving the speech recognition efficiency.
[0090] Refer to Figure 1 , Figure 1A schematic diagram of an implementation environment provided by an embodiment of this application. This implementation environment includes a first terminal 101. Among them, the first terminal 101 can obtain an unlabeled first sample speech data set and a labeled second sample speech data set, initialize an original model, perform unsupervised training on the original model based on the first sample speech data set to obtain a basic speech processing model, prune the basic coding layers after a preset coding layer among the multiple basic coding layers of the basic speech processing model, perform supervised training on the pruned basic speech processing model based on the second sample speech data set to obtain a first speech recognition model, initialize a second speech recognition model, and perform distillation training on the second speech recognition model based on the first sample speech data set with the first speech recognition model as a benchmark to obtain a target speech recognition model. Then, the first terminal 101 collects target speech data to be recognized and calls the pre-deployed target speech recognition model to perform speech recognition on the target speech data to obtain the target recognition result of the target speech data.
[0091] Refer to Figure 2 , Figure 2 A schematic diagram of another implementation environment provided by an embodiment of this application. This implementation environment includes a second terminal 201 and a server 202. Among them, the server 202 can obtain an unlabeled first sample speech data set and a labeled second sample speech data set, initialize an original model, perform unsupervised training on the original model based on the first sample speech data set to obtain a basic speech processing model, prune the basic coding layers after a preset coding layer among the multiple basic coding layers of the basic speech processing model, perform supervised training on the pruned basic speech processing model based on the second sample speech data set to obtain a first speech recognition model, initialize a second speech recognition model, and perform distillation training on the second speech recognition model based on the first sample speech data set with the first speech recognition model as a benchmark to obtain a target speech recognition model. Then, the second terminal 201 sends the target speech data to be recognized to the server 202, and the server 202 calls the pre-deployed target speech recognition model to perform speech recognition on the target speech data and sends the target recognition result of the target speech data to the second terminal 201.
[0092] The above server 202 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0093] In addition, the server 202 can also be a node server in a blockchain network. Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Essentially, blockchain is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. Blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.
[0094] The first terminal 101 and the second terminal 201 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but are not limited thereto. The second terminal 201 and the server 202 can be directly or indirectly connected through a wired or wireless communication method, and the embodiments of the present application do not limit this here.
[0095] The method provided by the embodiments of the present application can be applied to various technical fields, including but not limited to technical fields such as cloud technology, artificial intelligence, and speech recognition.
[0096] Refer to Figure 3 , Figure 3 is a schematic flowchart of the speech recognition method provided by the embodiments of the present application. The speech recognition method can be executed by the server, or by the terminal, or by the cooperation of the terminal and the server. The speech recognition method includes but is not limited to the following steps 301 to step 305.
[0097] Step 301: Obtain an unlabeled first sample speech data set and a labeled second sample speech data set.
[0098] Among them, the first sample speech data set can include multiple unlabeled speech data, for example, it can be instant messaging speech data, oral exam speech data, smart home control speech data, vehicle-mounted control speech data, etc. The first sample speech data set does not contain the labeled text corresponding to each speech data; the second sample speech data set can include multiple labeled speech data. Similarly, the speech data in the second sample speech data set can also be instant messaging speech data, oral exam speech data, smart home control speech data, vehicle-mounted control speech data, etc. The second sample speech data set contains the labeled text corresponding to each speech data. Of course, the speech data in the second sample speech data set has the same type as the speech data in the first sample speech data set to improve the overall training effect of the speech recognition model.
[0099] In a possible implementation, the first sample speech dataset and the second sample speech dataset can be obtained from the local space or from other platforms, and the other platforms can refer to devices dedicated to data storage. The speech data in the above-mentioned first sample speech dataset and second sample speech dataset can come from the same object, that is, the same object can input different contents as samples of speech data. For example, object A inputs speech content S1 as a sample of speech data, and object A inputs speech content S2 as a sample of speech data. Or, the speech data in the above-mentioned first sample speech dataset and second sample speech dataset can also come from different objects, that is, different objects can input different contents as samples of speech data; for example, object A inputs speech content S1 as a sample of speech data, and object B inputs speech content S2 as a sample of speech data.
[0100] Step 302: Initialize the original model, and perform unsupervised training on the original model based on the first sample speech dataset to obtain a basic speech processing model;
[0101] In a possible implementation, referring to Figure 4 , Figure 4 FIG.
[0102] is an exemplary structural schematic diagram of the basic speech processing model provided by the embodiment of the present application. The basic speech processing model includes a basic convolutional network, a linear layer, a basic encoding network, and a basic output layer connected in sequence. The basic convolutional network includes a plurality of basic convolutional layers connected in sequence, and the basic encoding network includes a plurality of basic encoding layers connected in sequence; the original model is the model in the initial state corresponding to the basic speech processing model, that is, the structure of the original model is the same as that of the basic speech processing model, and the parameters of the original model can be randomly initialized.
[0103] In a possible implementation, referring to Figure 5 , Figure 5FIG. 0 is an exemplary structural diagram of a basic encoding layer provided by an embodiment of the present application. Specifically, the basic encoding layer mainly includes a multi-head attention unit and a feed-forward network unit. The features obtained after convolution processing by the basic convolutional layer are input into the basic encoding layer. First, the multi-head attention unit is used to extract features to obtain high-level features. The high-level features output by the multi-head attention unit are then subjected to feature mapping through the feed-forward network unit to obtain the output features of the current basic encoding layer, and then sent to the next basic encoding layer for processing until the last basic encoding layer is processed and then sent to the basic output layer for processing. Moreover, the features output by the multi-head attention unit and the features output by the feed-forward network unit can be further subjected to residual connection processing and normalization processing. The role of residual connection processing is to enable the features output by the multi-head attention unit or the features output by the feed-forward network unit to carry more information, and it can also make the backpropagation during training more stable. The role of normalization processing is to accelerate the convergence speed of the model during training and improve the training efficiency.
[0104] Based on Figure 5 the structure of the basic encoding layer shown, the structural parameters of the basic encoding layer in the basic speech processing model may include the feature dimension of the multi-head attention unit and the feature dimension of the feed-forward network unit. For example, the feature dimension of the multi-head attention unit may be 1024, and the feature dimension of the feed-forward network unit may be 4096.
[0105] It can be understood that Figure 4 the basic speech processing model shown in [FIG. 10] has seven basic convolutional layers and twenty-four basic encoding layers. In fact, the number of basic convolutional layers and the number of basic encoding layers of the basic speech processing model can be determined according to actual needs, and the embodiments of the present application do not make any limitations. In addition, the structural parameters of each basic convolutional layer and the structural parameters of each basic encoding layer can also be determined according to actual needs, and the embodiments of the present application do not make any limitations.
[0106] In addition, the role of the linear layer is to adjust the feature dimension of the features output by the basic convolutional network so that the feature dimension of the basic convolutional network matches the feature dimension of the basic encoding network. The role of the basic output layer is mainly to perform classification processing based on the features output by the basic encoding network to obtain the final output result. In a possible implementation manner, both the linear layer and the basic output layer can be implemented based on a fully connected layer.
[0107] Among them, since the first sample speech dataset is unlabeled data, the training based on the first sample speech dataset is unsupervised training. Correspondingly, after completing the unsupervised training of the original model, the obtained basic speech processing model can be made to have the function of extracting speech features. Therefore, by first performing unsupervised training on the original model based on the first sample speech dataset, the training can be specifically carried out for the speech feature extraction function, achieving the purpose of refined training, which is beneficial to improving the accuracy of speech feature extraction of the basic speech processing model. Subsequently, when further training is carried out based on the basic speech processing model, the training effect can be effectively improved.
[0108] Step 303: Prune the basic encoding layers after the preset encoding layer among the multiple basic encoding layers, and perform supervised training on the pruned basic speech processing model based on the second sample speech dataset to obtain the first speech recognition model.
[0109] Among them, since the second sample speech dataset is labeled data, the training based on the second sample speech dataset is supervised training. Correspondingly, after completing the supervised training of the basic speech processing model, the obtained first speech recognition model can be made to have the speech recognition function.
[0110] In a possible implementation manner, based on Figure 4 the shown basic speech processing model, the preset encoding layer can be the twentieth encoding layer. At this time, prune the basic encoding layers after the twentieth among the twenty-four basic encoding layers, that is, prune the twenty-first basic encoding layer, the twenty-second basic encoding layer, the twenty-third basic encoding layer, and the twenty-fourth basic encoding layer. Correspondingly, the pruned basic speech processing model includes twenty basic encoding layers. Since the features output by the basic encoding layers closer to the end are more likely to fit the training task and will affect the calculation of the loss value. Among them, being more likely to fit the training task means that the loss value is more likely to converge, but actually the performance of the model may not meet the requirements, thus affecting the training effect of the model. Therefore, pruning the basic encoding layers after the preset encoding layer in the basic speech processing model can improve the training effect of the basic speech processing model, improve the performance of the basic speech processing model, and at the same time can reduce the volume of the basic speech processing model, and correspondingly reduce the volume of the first speech recognition model.
[0111] Step 304: Initialize the second speech recognition model, and based on the first sample speech dataset, perform distillation training on the second speech recognition model with the first speech recognition model as the benchmark to obtain the target speech recognition model.
[0112] In a possible implementation, the second speech recognition model is the model in the initial state corresponding to the target speech recognition model, that is, the second speech recognition model has the same structure as the target speech recognition model. The first speech recognition model serves as the teacher model for distillation training, and the second speech recognition model serves as the student model for distillation training. Among them, the number of parameters of the second speech recognition model is less than that of the first speech recognition model, that is, the second speech recognition model has a smaller model size and a higher degree of model lightweight.
[0113] By taking the first speech recognition model as a benchmark and performing distillation training on the second speech recognition model based on the first sample speech dataset, the second speech recognition model can learn the knowledge of the first speech recognition model. Since the first speech recognition model is trained by the pruned basic speech processing model, the first speech recognition model can inherit the performance optimization effect of the pruned basic speech processing model. When using the first speech recognition model as the teacher model for distillation training, it is beneficial to improve the performance of the obtained target speech recognition model.
[0114] Moreover, in the process of training the target speech recognition model provided by the embodiments of the present application, it is divided into three training stages: unsupervised training of the original model, supervised training of the pruned basic speech processing model, and distillation training of the second speech recognition model. By introducing the unlabeled first sample speech dataset and the labeled second sample speech dataset, and corresponding to different types of sample speech datasets in different training stages, the effect of semi-supervised training can be achieved. And because the first sample speech dataset does not need to be labeled, it can reduce the cost of data labeling, improve the training effect of the model, and is beneficial to improving the performance of the target speech recognition model.
[0115] Therefore, the speech recognition method provided by the embodiments of the present application can be applied to scenarios where the acquisition speed of labeled sample speech data is slow, the labeling cost is high, and the amount of effective data is small. For example, in the oral examination scenario of intelligent education, since the examination is held regularly and the examination questions for each candidate are based on a limited number of examination question types and questions, the acquisition of speech data is difficult. Based on this, by corresponding to different types of sample speech datasets in different training stages in the embodiments of the present application, the usage requirement of the labeled speech data can be reduced, and the data labeling cost can be reduced. Correspondingly, the above-mentioned first sample speech dataset and the second sample speech dataset include speech data in the oral examination.
[0116] In summary, in the speech recognition method provided by the embodiments of the present application, by means of semi-supervised training combined with pruning the basic speech processing model, the performance of the target speech recognition model can be effectively improved. Moreover, by means of pruning the basic speech processing model combined with distillation training, the volume of the target speech recognition model can be effectively reduced, so that the target speech recognition model has the advantages of small volume and high performance.
[0117] Step 305: Perform speech recognition on the target speech data based on the target speech recognition model to obtain a target recognition result corresponding to the target speech data.
[0118] Among them, the target speech data is the speech data to be recognized, which can be the speech data to be recognized in the oral examination scenario, the speech data to be recognized in the instant messaging scenario, the speech data to be recognized in the smart home control scenario, the speech data to be recognized in the vehicle-mounted system control scenario, etc. The embodiments of the present application do not make limitations. Since the target speech recognition model provided by the embodiments of the present application has the advantages of small volume and high performance, therefore, using the target speech recognition model to perform speech recognition on the target speech data can improve the speech recognition accuracy while improving the speech recognition efficiency.
[0119] In a possible implementation manner, referring to Figure 6 , Figure 6 is a schematic flowchart of unsupervised training of the original model provided by the embodiments of the present application. The original model includes an original convolutional network and an original encoding network connected in sequence (for simplicity of description, the linear layer and the original output layer are not shown). When performing unsupervised training on the original model based on the first sample speech data set, the first sample speech data set can be input into the original model, a masking operation is performed on the original convolutional features output by the original convolutional network to obtain masked convolutional features, a product quantization operation is performed on the original convolutional features to obtain quantized convolutional features, the masked encoding features output after the original encoding network processes the masked convolutional features are obtained, a first original loss value is determined according to the masked encoding features and the quantized convolutional features, and the original model is trained unsupervised according to the first original loss value.
[0120] Specifically, the form of the original convolutional features output by the original convolutional network is Z1......Z r , where Z1 is the feature representation of the first speech frame, Z rIt is the feature representation of the r-th speech frame. After the original convolutional network outputs the original convolutional features, on the one hand, the original convolutional features are input to the original encoding network after a masking operation. Among them, one or more of the r speech frames can be selected for the masking operation, and the feature values corresponding to the masked speech frames are set to zero. On the other hand, a product quantization operation is performed on the original convolutional features to obtain quantized convolutional features. The product quantization operation is to decompose the original vector space into the Cartesian product of several low-dimensional vector spaces (quantization codebooks), and perform clustering processing in the quantization codebooks obtained by the decomposition to obtain the cluster centers of each quantization codebook and the features of the cluster centers. Then, the features of the cluster centers of each quantization codebook are used to replace other features, so that the original infinite feature expression space collapses into a finite offline space, making the features more robust and having a higher feature expression ability.
[0121] After obtaining the masked encoded features and the quantized convolutional features, the first original loss value can be determined according to the masked encoded features and the quantized convolutional features. The first original loss value is used to make the features corresponding to the masked speech frames in the masked encoded features as similar as possible to the features of the corresponding speech frames in the quantized convolutional features, while the features corresponding to the masked speech frames in the masked encoded features are as dissimilar as possible to the features of the remaining speech frames in the quantized convolutional features. Even if the input to the original encoding network is the masked convolutional features obtained after masking processing, the original encoding network can still capture the feature information well, thereby improving the performance of the original encoding network. Among them, when specifically calculating the first original loss value, the first similarity between the feature values corresponding to the masked speech frames in the masked encoded features and the feature values of the corresponding speech frames (positive samples) in the quantized convolutional features can be calculated, and the second similarity between the feature values corresponding to the masked speech frames in the masked encoded features and the feature values of the remaining speech frames (negative samples) in the quantized convolutional features can be calculated. The first original loss value can be obtained according to the quotient between the first similarity and the second similarity.
[0122] In a possible implementation manner, a second original loss value can be further introduced to perform unsupervised training on the original model. The second original loss value is used to supervise the product quantization operation to make the distances between the cluster centers as far as possible, thereby improving the rationality of the product quantization operation. Among them, the first quantity of the quantization codebooks when performing the product quantization operation and the second quantity of the cluster centers in each quantization codebook can be obtained, the probability distribution of any one of the cluster centers in each quantization codebook being selected can be determined, and the second original loss value can be determined according to the first quantity, the second quantity, and the probability distribution. The original model is subjected to unsupervised training according to the first original loss value and the second original loss value.
[0123] Specifically, the above probability distribution can be divided by the product of the first quantity and the second quantity to obtain a second original loss value. After obtaining the second original loss value, the first original loss value and the second original loss value can be weighted to obtain a target original loss value, and then the original model can be unsupervised trained according to the target original loss value.
[0124] In a possible implementation, referring to Figure 7 , Figure 7 is an exemplary structural schematic diagram of the first speech recognition model provided by the embodiments of the present application, corresponding to the structure of the Figure 4 shown basic speech processing model. Figure 7 The first speech recognition model shown includes a first convolutional network, a linear layer, a first encoding network, and a first output layer connected in sequence. The first convolutional network includes a plurality of first convolutional layers connected in sequence. The first encoding network includes a plurality of first encoding layers connected in sequence. Since the first speech recognition model is trained based on the pruned basic speech processing model, the structure of the first convolutional network is similar to the structure of the basic convolutional network, that is, the number of the first convolutional networks is seven, and the structural parameters of each first convolutional network are set as (512, 10, 5), (512, 3, 2), (512, 3, 2), (512, 3, 2), (512, 3, 2), (512, 2, 2), (512, 2, 2) in sequence. The number of the first encoding layers is twenty, the feature dimension of the multi-head attention unit in the first encoding layer is 1024, and the feature dimension of the feed-forward network unit is 4096. The functions of the linear layer and the first output layer can be referred to the explanations in the basic speech processing model and will not be elaborated here.
[0125] To reduce the model size, the structure of the first speech recognition model is compressed in the embodiments of the present application to obtain a second speech recognition model. For example, the feature dimension of the first convolutional network can be reduced, the feature dimension of the first encoding network can be reduced, or the number of the first encoding layers can be reduced, etc. Among them, one or a combination of the above several compression methods can be selected for execution.
[0126] In a possible implementation, the feature dimension of the first convolutional layer, the feature dimension of the first encoding layer, and the number of the first encoding layers can be reduced simultaneously, so as to improve the volume compression effect on the second speech recognition model.
[0127] Among them, when reducing the feature dimension of the first convolutional network, the feature dimension of the first N first convolutional layers in the first convolutional network can be reduced, where N is a positive integer. Because reducing the feature dimension of the first convolutional layer closer to the input end can more significantly improve the processing efficiency of the model. For example, N can be 2, that is, at this time, the feature dimension of the first two first convolutional layers in the first convolutional network is reduced. Based on Figure 4For the structure shown, the feature dimension of the first two first convolutional layers can be reduced from the original 512 to 256. It should be noted that the computational load of the first convolutional network is generally concentrated in the first two first convolutional layers. Therefore, when N is 2, the efficiency improvement effect of the second speech recognition model can be more obvious.
[0128] Among them, to reduce the feature dimension of the first encoding network, based on Figure 4 For the structure shown, the feature dimension of the multi-head attention unit in the first encoding layer can be reduced from the original 1024 to 384, and the feature dimension of the feed-forward network unit can be reduced from the original 4096 to 1536.
[0129] Among them, to reduce the number of layers of the first encoding layer, based on Figure 4 For the structure shown, the number of the first encoding layers can be reduced from the original twenty to ten.
[0130] It can be understood that the specific values of reducing the feature dimension of the first convolutional network, reducing the feature dimension of the first encoding network, and reducing the number of layers of the first encoding layer can be determined according to the actual model volume compression requirements, and the embodiments of the present application do not make limitations.
[0131] After reducing the feature dimension of the first convolutional network, reducing the feature dimension of the first encoding network, and reducing the number of layers of the first encoding layer, the obtained second speech recognition model includes a second convolutional network, a linear layer, a second encoding network, and a second output layer connected in sequence. The second convolutional network includes a plurality of second convolutional layers connected in sequence. The second encoding network includes a plurality of second encoding layers connected in sequence. And the number of the second encoding layers is less than the number of the first encoding layers. The number of the second convolutional layers is equal to the number of the first convolutional layers. The second convolutional layer located before the preset convolutional layer in the second speech recognition model is the target convolutional layer, and the feature dimension of the target convolutional layer is less than the feature dimension of the corresponding first convolutional layer of the target convolutional layer.
[0132] Among them, the target convolutional layer is the second convolutional layer obtained after reducing the feature dimension, and the preset convolutional layer is the first second convolutional layer after the last target convolutional layer. For example, based on the above example, assuming that when reducing the feature dimension of the first two first convolutional layers in the first convolutional network, the target convolutional layers are the first second convolutional layer and the second second convolutional layer. Correspondingly, the preset convolutional layer is the third second convolutional layer. It can be understood that when reducing the feature dimension of the first three first convolutional layers in the first convolutional network, the target convolutional layers are the first second convolutional layer, the second second convolutional layer, and the third second convolutional layer. Correspondingly, the preset convolutional layer is the fourth second convolutional layer, and so on.
[0133] For example, referring to Figure 8 , Figure 8FIG. 0 is an exemplary structural diagram of a second speech recognition model provided by an embodiment of the present application. The number of second convolutional networks is seven, and the structural parameters of each second convolutional network are set in sequence as (256, 10, 5), (256, 3, 2), (512, 3, 2), (512, 3, 2), (512, 3, 2), (512, 2, 2), (512, 2, 2). The number of second encoding layers is ten. The feature dimension of the multi-head attention unit in the second encoding layer is 384, and the feature dimension of the feed-forward network unit is 1536. The functions of the linear layer and the second output layer can be referred to the explanations in the basic speech processing model and will not be elaborated here.
[0134] Based on this, when initializing the second speech recognition model, specifically, the encoding parameters of each second encoding layer can be randomly initialized, the output parameters of the first output layer are used as the output parameters of the second output layer, the convolution parameters of the preset convolutional layer and the convolution parameters of the target convolutional layer are randomly initialized, the convolution parameters of the corresponding first convolutional layer of the remaining convolutional layers are used as the convolution parameters of the remaining convolutional layers, and the second speech recognition model is initialized according to the convolution parameters of each second convolutional layer, the encoding parameters of each second encoding layer, and the output parameters of each second output layer.
[0135] Among them, the remaining convolutional layers are the other second convolutional layers except the preset convolutional layer and the target convolutional layer in the multiple second convolutional layers. For example, if the target convolutional layer is the first second convolutional layer and the second second convolutional layer, and the preset convolutional layer is the third second convolutional layer, then the remaining convolutional layers are the fourth second convolutional layer, the fifth second convolutional layer, the sixth second convolutional layer, and the seventh second convolutional layer. If the remaining convolutional layer is the fourth second convolutional layer, then the corresponding first convolutional layer of this remaining convolutional layer is the fourth first convolutional layer, and so on.
[0136] Among them, the convolution parameters of the second convolutional layer are the parameters that need to be learned when training the second speech recognition model, such as convolution neuron weights, convolution neuron biases, and other parameters. The encoding parameters of the second encoding layer are the parameters that the second encoding layer needs to learn when training the second speech recognition model, such as attention weights, feed-forward network neuron weights, and other parameters. The output parameters of the second output layer are the parameters that the second output layer needs to learn when training the second speech recognition model, such as fully connected neuron weights, fully connected neuron biases, and other parameters. The convolution parameters, encoding parameters, and output parameters will not be listed one by one here.
[0137] By using the output parameters of the first output layer as the output parameters of the second output layer, and using the convolution parameters of the first convolutional layer corresponding to the remaining convolutional layers as the convolution parameters of the remaining convolutional layers, that is, the output parameters of the second output layer can be obtained by directly copying the output parameters of the first output layer, and the convolution parameters of the remaining convolutional layers can be obtained by directly copying the output parameters of the corresponding first convolutional layer. Thus, the acquisition efficiency of the second speech recognition model can be improved. Moreover, since the first speech recognition model has been trained, the output parameters of the first output layer and the convolution parameters of the first convolutional layer are effective. Therefore, the performance of the second speech recognition model during initialization can be improved to a certain extent, the parameter adjustment cost during subsequent training of the second speech recognition model can be reduced, and the parameter adjustment efficiency can be improved.
[0138] In a possible implementation manner, based on the structure of the second speech recognition model described above, when performing distillation training on the second speech recognition model based on the first sample speech dataset, the first sample speech dataset can be input into the first speech recognition model to obtain the first convolutional features output by the first convolutional layer corresponding to the preset convolutional layer, and obtain the second convolutional features output by the last first convolutional layer. Then, the first sample speech dataset is input into the second speech recognition model to obtain the third convolutional features output by the preset convolutional layer, and obtain the fourth convolutional features output by the last second convolutional layer. The first convolutional loss value is determined according to the first convolutional features and the third convolutional features, the second convolutional loss value is determined according to the second convolutional features and the fourth convolutional features, the target convolutional loss value is determined according to the first convolutional loss value and the second convolutional loss value, and the second convolutional network is trained by distillation according to the target convolutional loss value. Among them, training the second convolutional network by distillation can be to adjust model parameters related to the convolution parameters of the second convolutional network, etc.
[0139] In a possible implementation manner, the target convolutional loss value can be the sum of the first convolutional loss value and the second convolutional loss value, or alternatively, the target convolutional loss value can be obtained by weighting the first convolutional loss value and the second convolutional loss value. The embodiments of the present application do not make any limitations.
[0140] Among them, referring to Figure 9 , Figure 9Schematic diagram of the distillation training of the second convolutional network provided by the embodiments of the present application. The embodiments of the present application mainly determine the loss value of the distillation training by comparing the output results of the first speech recognition model and the second speech recognition model after processing the first sample speech data, and then train the second speech recognition model. Specifically, the target convolutional loss value is used to make the performance of the second convolutional network closer to the performance of the first convolutional network. By introducing the first convolutional loss value and the second convolutional loss value, the first convolutional loss value can be used to measure the impact on the second convolutional layer before the preset convolutional layer after reducing the feature dimension, and the second convolutional loss value can be used to measure the impact on the overall second speech recognition model after reducing the feature dimension. Therefore, using the combination of the first convolutional loss value and the second convolutional loss value to perform distillation training on the second convolutional network can make the overall loss value more reasonable, improve the accuracy of the distillation training of the second convolutional network, and make the performance of the second convolutional network closer to the performance of the first convolutional network.
[0141] The above target convolutional loss value can be expressed as:
[0142]
[0143] Among them,
[0144] Among them, Loss cnn-distil represents the target convolutional loss value, represents the first convolutional feature output by the third first convolutional layer, represents the second convolutional feature output by the seventh first convolutional layer, represents the third convolutional feature output by the preset convolutional layer, represents the fourth convolutional feature output by the seventh second convolutional layer, t3 represents the number of speech frames output by the third first convolutional layer or the preset convolutional layer, t7 represents the number of speech frames output by the seventh first convolutional layer or the seventh second convolutional layer, represents the feature encoder of the third first convolutional layer, represents the feature encoder of the seventh first convolutional layer, represents the feature encoder of the preset convolutional layer, represents the feature encoder of the seventh second convolutional layer, X represents the input of the corresponding convolutional layer, and MSE represents the calculation of the mean square error.
[0145] In addition, the last target convolutional layer (i.e., Figure 9Calculate the first convolution loss value based on the features output by the second second convolution layer in []. It should be noted that, compared with calculating the first convolution loss value using the features output by the last target convolution layer, when using a preset convolution layer (i.e., Figure 9 the third second convolution layer in []) to calculate the first convolution loss value, on the one hand, since the feature dimension of the target convolution layer is smaller than that of the corresponding first convolution layer, if the first convolution loss value is calculated based on the features output by the last target convolution layer, it is necessary to first perform a conversion of the feature dimension, which reduces the calculation efficiency of the first convolution loss value; on the other hand, since the preset convolution layer is the first second convolution layer after the last target convolution layer, that is, the preset convolution layer is the second convolution layer closest to the target convolution layer, the reduction of the feature dimension will have a certain impact on the output of the preset convolution layer. Therefore, using the third convolution features output by the preset convolution layer to calculate the first convolution loss value can make the first convolution loss value more reasonable and improve the accuracy of the first convolution loss value.
[0146] Moreover, since the convolution parameters of the remaining convolution layers have not changed compared with the convolution parameters of the corresponding first convolution layer, therefore, calculating the target convolution loss value only through the first convolution loss value and the second convolution loss value can avoid introducing other training parameters additionally.
[0147] In addition, since the preset convolution layer is the second convolution layer closest to the target convolution layer, if the convolution parameters of the first convolution layer corresponding to the preset convolution layer are used as the convolution parameters of the preset convolution layer when initializing the second speech recognition model, then the convolution parameters of the preset convolution layer are actually adapted to the first convolution layer of the first speech recognition model at this time. If the third convolution features output by the preset convolution layer are used to calculate the first convolution loss value at this time, it cannot accurately reflect the impact on the second convolution layer before the preset convolution layer after reducing the feature dimension. Therefore, by randomly initializing the convolution parameters of the preset convolution layer and the target convolution layer, the calculation of the first convolution loss value can be effective and reliable.
[0148] It should be noted that in the above example, the feature dimension of the target convolutional layer is smaller than that of the first convolutional layer (basic convolutional layer) corresponding to the target convolutional layer. This is because when the feature dimension of the first convolutional layer (basic convolutional layer) corresponding to the target convolutional layer is set to be relatively large, the performance of the first speech recognition model can be improved. Subsequently, when training the second speech recognition model based on the first speech recognition model, the training effect can be improved, making the performance of the obtained second speech recognition model better. Of course, when training the basic speech processing model, the feature dimension of the basic convolutional network can also be directly set to the compressed feature dimension. For example, in the above example, the feature dimension of the basic convolutional network can be set to 256, that is, the structural parameters of each basic convolutional layer of the basic speech processing model can be (256, 10, 5), (256, 3, 2), (512, 3, 2), (512, 3, 2), (512, 3, 2), (512, 2, 2), (512, 2, 2) in sequence. Subsequently, it is not necessary to perform separate distillation training on the second convolutional layer of the second speech recognition model to achieve the simplified effect of distillation training on the second speech recognition model.
[0149] The principle of distillation training for the second encoding network and the second output layer will be described in detail below.
[0150] In a possible implementation manner, when performing distillation training on the second speech recognition model based on the first sample speech dataset, the first sample speech dataset can be input into the first speech recognition model, the same number of benchmark encoding layers as the second encoding layer can be determined from the first speech recognition model, the first encoding features output by each benchmark encoding layer can be obtained, the first sample speech dataset can be input into the second speech recognition model, the second encoding features output by each second encoding layer can be obtained, the distillation fitting parameters of each second encoding layer can be determined, the feature dimension of the corresponding second encoding feature can be adjusted according to the distillation fitting parameters, the encoding layer loss value corresponding to each second encoding layer can be determined according to the second encoding feature with the adjusted feature dimension and the corresponding first encoding feature, the target encoding loss value can be determined according to each encoding layer loss value, and the second encoding network can be distilled and trained according to the target encoding loss value. Among them, distilling and training the second encoding network can be to adjust the model parameters related to the encoding parameters of the second encoding network.
[0151] In a possible implementation manner, the target encoding loss value can be the sum of the encoding layer loss values, or the target encoding loss value can also be obtained by weighting the encoding layer loss values according to different second encoding layers. The embodiments of the present application do not make any limitations.
[0152] Refer to Figure 10 , Figure 10Schematic diagram of the distillation training of the second encoding network provided by the embodiment of the present application. Specifically, the reference encoding layer is the first encoding layer corresponding to the second encoding layer, and the number of reference encoding layers is the same as the number of second encoding layers. Correspondingly, the number of reference encoding layers is ten. There are various ways to determine the reference encoding layer. For example, the first ten first encoding layers in the first encoding network can be used as the reference encoding layer; or the last ten first encoding layers in the first encoding network can be used as the reference encoding layer; or, the 2m-th first encoding layer in the first encoding network can be used as the reference encoding layer corresponding to the m-th second encoding layer in the second encoding network, where m is a positive integer. That is, the second first encoding layer in the first encoding network corresponds to the first second encoding layer in the second encoding network, the fourth first encoding layer in the first encoding network corresponds to the second second encoding layer in the second encoding network,..., and the 20th first encoding layer in the first encoding network corresponds to the tenth second encoding layer in the second encoding network.
[0153] The above-mentioned target encoding loss value can be expressed as:
[0154]
[0155] Among them, Loss decode-distil represents the target encoding loss value, represents the output of the m-th second encoding layer, t represents the number of speech frames output by the m-th second encoding layer, represents the distillation fitting parameter corresponding to the m-th second encoding layer, represents the output of the 2m-th first encoding layer (i.e., the reference encoding layer), t represents the number of speech frames output by the 2m-th first encoding layer, MSE represents the calculation of the mean square error, and m is a positive integer.
[0156] By calculating the encoding layer loss value from the first encoding features output by the reference encoding layer and the second encoding features output by the corresponding second encoding layer, the performance of the second encoding network can be made closer to the performance of the first encoding network. Moreover, by using the 2m-th first encoding layer in the first encoding network as the reference encoding layer corresponding to the m-th second encoding layer in the second encoding network, the distribution of the reference encoding layer can be made more uniform. Subsequently, when calculating the encoding layer loss value, the encoding layer loss value can be made more reasonable, improving the accuracy of the encoding layer loss value.
[0157] Moreover, the embodiments of the present application introduce distillation fitting parameters. Since the feature dimensions of the reference encoding layer and the second encoding layer are different, by adjusting the feature dimensions of the corresponding second encoded features according to the distillation fitting parameters, the first encoded features output by the reference encoding layer can be fitted to the second encoded features output by the second encoding layer, thereby improving the accuracy of the encoding layer loss value.
[0158] Among them, different second encoding layers can correspond to different distillation fitting parameters to improve the fitting effect of features. Moreover, when performing distillation training on the second encoding network according to the target encoding loss value, the distillation fitting parameters corresponding to each second encoding layer can also be adjusted according to the target encoding loss value, making the distillation fitting parameters more accurate and reasonable.
[0159] In addition, since the purpose of performing distillation training on the second encoding network is to make the performance of the second encoding network closer to the performance of the first encoding network, and the output of the first speech recognition model can be used as the label for distillation training, an unlabeled first sample speech data set can be used during the distillation training of the second encoding network, thereby reducing the labeling cost.
[0160] In a possible implementation manner, a target output loss value corresponding to the second output layer can be further introduced to perform distillation training on the second encoding network. Specifically, when performing distillation training on the second encoding network according to the target encoding loss value, the first sample recognition result output by the first output layer and the second sample recognition result output by the second output layer can be obtained, the target output loss value can be determined according to the first sample recognition result and the second sample recognition result, and the second encoding network can be trained by distillation according to the target encoding loss value and the target output loss value.
[0161] Among them, the first sample recognition result is the speech recognition result of the first speech recognition model, and the second sample recognition result is the speech recognition result of the second speech recognition model. Therefore, the target output loss value can be determined according to the first sample recognition result and the second sample recognition result, and the obtained target output loss value is used to make the overall performance of the second speech recognition model closer to the overall performance of the first speech recognition model.
[0162] The above target output loss value can be expressed as:
[0163] Loss logit =MSE(O s ,O T )
[0164] Among them, Loss logit represents the target output loss value, O T represents the first sample recognition result, O s represents the second sample recognition result, OT ∈R 1024*C ,O s ∈R 1024*C , C represents the number of categories in the vocabulary of the speech recognition task, and MSE represents the calculation of the mean square error.
[0165] Among them, since the output parameters of the second output layer are obtained by copying the output parameters of the first output layer, the feature dimension of the second output layer does not match the feature dimension of the second network. Therefore, when the features output by the second encoding network are input to the second output layer for processing, the feature dimension needs to be transformed first. Based on this, taking the number of the second encoding layers as ten as an example, the above O s can be expressed as:
[0166]
[0167] Among them, is the feature output by the tenth second encoding layer, is the distillation fitting parameter corresponding to the tenth second encoding layer, logit represents the output function of the second output layer. After transforming the feature output by the tenth second encoding layer through the corresponding distillation fitting parameter and inputting it to the second output layer, the feature dimensions of the second output layer and the first output layer can be made to match, thereby improving the accuracy of the target output loss value.
[0168] After introducing the target output loss value, there are two ways to perform distillation training on the second encoding network. One way is to perform distillation training on the second encoding network according to the target encoding loss value, and then perform distillation training on the second encoding network again according to the target output loss value, that is, adjusting the model parameters once using the target encoding loss value, and then adjusting the model parameters once using the target output loss value. Adjusting the parameters according to the target encoding loss value and the target encoding loss value are independent of each other; another way is to weight the target encoding loss value and the target output loss value to obtain the target model loss value, and perform distillation training on the second encoding network according to the target model loss value.
[0169] The above target model loss value can be expressed as:
[0170]
[0171] Among them, Loss model-distil represents the target model loss value, γ is the weight for performing distillation training on the second encoding network, and γ can be determined according to the actual situation. For example, it can be 0.5, 0.7, etc., and the embodiments of the present application do not make limitations.
[0172] In a possible implementation, after introducing the target output loss value, in addition to adjusting the encoding parameters of the second encoding network, the output parameters of the second output layer can also be adjusted according to the target output loss value. And since the features input to the second encoding network are processed by the second convolutional network, after calculating the target encoding loss value, the convolutional parameters of the second convolutional network can also be adjusted according to the target encoding loss value. Similarly, after calculating the target output loss value, the convolutional parameters of the second convolutional network can also be adjusted according to the target encoding loss value.
[0173] After completing the overall training of the target speech recognition model, when performing speech recognition on the target speech data based on the target speech recognition model, the feature dimension of the second output layer in the target speech recognition model can be adjusted according to the distillation fitting parameters of the last second encoding layer, and speech recognition is performed on the target speech data based on the target speech recognition model with the adjusted feature dimension.
[0174] For example, based on the above example, the feature dimension of the second output layer is O s ∈R 1024*C , the last second encoding layer is the tenth encoding layer, and the feature dimension of the corresponding distillation fitting parameters is Then the adjusted feature dimension of the second output layer is 384*C, which has a certain degree of decrease compared to the feature dimension of the first output layer of 1024*C.
[0175] By using the distillation fitting parameters to adjust the feature dimension of the second output layer in the target speech recognition model, the optimization effect of the feature dimension of the second output layer can be achieved, the processing efficiency of the target speech recognition model can be improved, and the speech recognition efficiency can be improved when performing speech recognition on the target speech data based on the target speech recognition model with the adjusted feature dimension in the future.
[0176] In addition, when performing speech recognition on the target speech data based on the target speech recognition model, the distillation fitting parameters of each target encoding layer in the target speech recognition model can also be pruned, and speech recognition is performed on the target speech data based on the target speech recognition model after pruning.
[0177] Among them, the target encoding layer is the second encoding layer other than the last second encoding layer. Since the distillation fitting parameters are the parameters introduced during the distillation training of the second encoding network, when using the target speech recognition model to perform speech recognition on the target speech data, the distillation fitting parameters are actually not needed. Therefore, by pruning the distillation fitting parameters, the volume of the target speech recognition model can be reduced, the processing efficiency of the target speech recognition model can be improved, and the speech recognition efficiency can be improved when performing speech recognition on the target speech data based on the target speech recognition model with the adjusted feature dimension in the future.
[0178] It can be understood that since the distillation fitting parameters of the last second coding layer are used to adjust the feature dimension of the second output layer, the distillation fitting parameters corresponding to the target coding layer are pruned. In this case, it is equivalent to combining the optimization of the feature dimension of the second output layer and the pruning of the distillation fitting parameters of the target coding layer, which can make the improvement effect of the processing efficiency of the target speech recognition model better.
[0179] In addition, if the distillation fitting parameters are not used to adjust the feature dimension of the second output layer in the target speech recognition model, the distillation fitting parameters of each second coding layer in the target speech recognition model can be pruned, and the target speech data is subjected to speech recognition based on the target speech recognition model after pruning.
[0180] The following uses a practical example to illustrate in detail the complete process of training the target speech recognition model.
[0181] Refer to Figure 11 , Figure 11 which is the overall training flowchart of the target speech recognition model provided by the embodiment of the present application. The training of the target speech recognition model mainly includes a pre-training stage, a fine-tuning stage, and a distillation stage. In the pre-training stage and the distillation stage, unsupervised data (i.e., unlabeled data) is used for training. In the fine-tuning stage, supervised data (i.e., labeled data) is used for training, so as to achieve the effect of semi-supervised training.
[0182] Specifically, refer to Figure 12 , Figure 12 which is the detailed training flowchart of the target speech recognition model provided by the embodiment of the present application. In the pre-training stage, an original model is first initialized. The original model includes a convolutional network, a linear layer, an encoding network, and an output layer connected in sequence. The convolutional network is a seven-layer convolutional feature encoder f: X → Z, which takes speech data X as input and outputs corresponding latent speech features Z1......Z r . The structure of the convolutional network is (512, 10, 5), (512, 3, 2), (512, 3, 2), (512, 3, 2), (512, 3, 2), (512, 2, 2), (512, 2, 2) in sequence. The speech features Z1......Z output by the convolutional network r are sent to a twenty-four-layer encoding network g: Z → C after feature mapping by the linear layer to further capture feature information C1......C r , and the feature dimension of the attention unit of each encoding layer in the encoding network is 1024, and the feature dimension of the feed-forward network unit is 4096. After pre-training the original model based on unsupervised data, a basic speech processing model is obtained.
[0183] In the fine-tuning stage, first prune the basic speech processing model. Specifically, prune the last four of the twenty-four encoding layers, that is, the number of remaining encoding layers is twenty. Then randomly initialize the parameters of the output layer, and fine-tune the basic speech processing model based on the supervised data. Assume that the number of categories in the classification vocabulary of the speech recognition task is C( Figure 12 exemplarily shown in Figure 12 as C = 5176), then the feature dimension of the output layer is 1024 * C. After fine-tuning, the first speech recognition model is obtained.
[0184] Table 1 Statistical Table of Processing Time Consumption of the Second Speech Recognition Model
[0185]
[0186] In the distillation stage, first initialize a second speech recognition model. Referring to Table 1 above, Table 1 is the statistical table of processing time consumption of the second speech recognition model provided by the embodiments of the present application. It can be seen that the main factors affecting the real-time rate of model processing are the parameters of the first two layers of the convolutional network and the number of encoding network layers. If d1 and d2 are respectively set to 256, the processing time of the model's real-time rate on 625-millisecond audio is reduced by 25%, and the real-time rate is reduced from 0.18 to 0.13. If the audio duration is 6s, when the parameters of the first two layers of the convolutional network are further decreased and the number of encoding network layers is further reduced to 10 layers, the real-time rate of the model is only 0.08.
[0187] Therefore, for the convolutional network of the second speech recognition model, randomly initialize the parameters of the first three convolutional layers, and copy the corresponding parameters of the first speech recognition model to the last four convolutional layers; for the convolutional network of the second speech recognition model, randomly initialize a ten-layer encoding network. The feature dimension of the multi-head attention unit of the convolutional network of the second speech recognition model is 384, and the feature dimension of the feed-forward network unit of the convolutional network of the second speech recognition model is 1536; for the output layer of the second speech recognition model, copy the parameters of the output layer of the first speech recognition model. The feature dimension of the output layer of the second speech recognition model is 384.
[0188] After the initialization of the second speech recognition model is completed, it can be trained in a step-by-step distillation manner, which is divided into a convolutional network distillation stage, an encoding network distillation stage, and an output layer distillation stage:
[0189] In the convolutional network distillation stage, mainly compare the outputs of the third convolutional layer and the seventh convolutional layer in the first speech recognition model and the second speech recognition model to determine the loss value of the convolutional network. The specific calculation of the loss value of the convolutional network can refer to the previous explanation and will not be elaborated here. In the convolutional network distillation stage, mainly adjust the parameters of the convolutional network.
[0190] In the encoding network distillation stage, the main comparison is between the output of the 2m-th encoding layer in the first speech recognition model and the output of the m-th encoding layer in the second speech recognition model to determine the loss value of the encoding network. Moreover, for each encoding layer in the second speech recognition model, corresponding distillation fitting parameters are set for feature fitting. The specific calculation of the loss value of the encoding network can be referred to the previous explanation and will not be elaborated here. In the encoding network distillation stage, the main adjustment is the parameters of the encoding network.
[0191] In the output layer distillation stage, the main comparison is between the output of the output layer of the first speech recognition model and the output of the output layer of the second speech recognition model to determine the loss value of the output layer. Moreover, since the parameters of the output layer of the second speech recognition model are obtained by copying the parameters of the output layer in the first speech recognition model, it is necessary to calculate the output of the output layer of the second speech recognition model using the distillation fitting parameters corresponding to the tenth encoding layer of the second speech recognition model. The specific calculation of the loss value of the output layer can be referred to the previous explanation and will not be elaborated here. In the output layer distillation stage, the main adjustment is the parameters of the entire second speech recognition model.
[0192] After the above distillation training, prune the distillation fitting parameters corresponding to the first nine encoding layers in the encoding network of the second speech recognition model. In addition, multiply the distillation fitting parameters corresponding to the tenth encoding layer in the encoding network of the second speech recognition model by the parameters corresponding to the output layer. That is, the feature dimension of the output layer of the pruned second speech recognition model is 384*C. The size of the second speech recognition model obtained through the above processing is only 1 / 13 of the size of the first speech recognition model.
[0193] After testing, the single-core real-time rate of the target speech recognition model provided by the embodiment of the present application is 0.353, which is lower than 0.374 of the traditional hybird model.
[0194] Table 2 Test results of the character error rate of the target speech recognition model
[0195] Test set / Model Traditional hybird model Target speech recognition model Directly train a target speech recognition model of the same size Test set 1 20.90 14.56 17.52 Test set 2 29.34 22.45 25.33
[0196] In addition, referring to Table 2 above, Table 2 shows the test results of the character error rate of the target speech recognition model provided by the embodiment of the present application. The test sets 1 and 2 are speech data in the oral examination scenario. It can be seen that the model performance of the target speech recognition model obtained through the distillation training provided by the embodiment of the present application is improved by 31% compared with the traditional hybird model and by 24.5% compared with the target speech recognition model obtained by direct training, and it can be used for actual deployment.
[0197] It can be understood that in the above example, the parameter adjustments in the three training stages are independent. In fact, it is also possible to calculate the loss values of each training stage and then perform unified parameter adjustment.
[0198] After the training of the target speech recognition model is completed, it can be used for subsequent speech recognition processing. There are at least the following application scenarios:
[0199] Scenario 1
[0200] The target speech data to be recognized can be instant messaging speech data. For example, instant messaging user A sends instant messaging speech data to instant messaging user B. Instant messaging user B uses the speech-to-text function to convert the instant messaging speech data into the corresponding text. Then, the terminal used by instant messaging user B can call the pre-trained target speech recognition model, perform speech recognition on the instant messaging speech data based on the target speech recognition model, obtain the corresponding recognition result and display it on the screen.
[0201] Scenario 2
[0202] The target speech data to be recognized can be oral exam speech data. For example, the exam terminal collects the oral exam speech data of the examinee during the oral exam, calls the pre-trained target speech recognition model, performs speech recognition on the oral exam speech data based on the target speech recognition model, obtains the corresponding recognition result and grades according to the recognition result.
[0203] Scenario 3
[0204] The target speech data to be recognized can be smart home control speech data. For example, smart home appliances collect the smart home control speech data of the user, call the pre-trained target speech recognition model, perform speech recognition on the smart home control speech data based on the target speech recognition model, obtain the corresponding recognition result, determine the corresponding control instruction according to the recognition result and execute it.
[0205] Scenario 4
[0206] The target speech data to be recognized can be vehicle system control speech data. For example, vehicle-mounted devices collect the vehicle system control speech data of the driver, call the pre-trained target speech recognition model, perform speech recognition on the vehicle system control speech data based on the target speech recognition model, obtain the corresponding recognition result, determine the corresponding control instruction according to the recognition result and execute it.
[0207] It can be understood that although the steps in the above-mentioned various flowcharts are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above-mentioned flowchart may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0208] It should be noted that in each specific implementation manner of this application, when it comes to performing relevant processing based on data related to the characteristics of the target object, such as target object attribute information or attribute information sets, etc., the permission or consent of the target object will be obtained first. Moreover, the collection, use, and processing of these data will comply with the relevant laws, regulations, and standards of the relevant countries and regions. In addition, when the embodiment of this application needs to obtain the target object attribute information, it will obtain the separate permission or separate consent of the target object through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary target object-related data for the normal operation of the embodiment of this application will be obtained.
[0209] Refer to Figure 13 , Figure 13 , which is a schematic structural diagram of the voice recognition device provided by the embodiment of this application. The voice recognition device 1300 includes:
[0210] A sample data acquisition module 1301, configured to acquire an unlabeled first sample voice data set and a labeled second sample voice data set;
[0211] A first training module 1302, configured to initialize an original model, and perform unsupervised training on the original model based on the first sample voice data set to obtain a basic voice processing model; wherein, the basic voice processing model includes a plurality of sequentially connected basic encoding layers;
[0212] A second training module 1303, configured to prune the basic encoding layers after a preset encoding layer among the plurality of basic encoding layers, and perform supervised training on the pruned basic voice processing model based on the second sample voice data set to obtain a first voice recognition model;
[0213] A third training module 1304, configured to initialize a second voice recognition model, and perform distillation training on the second voice recognition model based on the first sample voice data set with the first voice recognition model as a benchmark to obtain a target voice recognition model;
[0214] A voice recognition module 1305, configured to perform voice recognition on target voice data based on a target voice recognition model to obtain a target recognition result corresponding to the target voice data.
[0215] Furthermore, the first voice recognition model includes a first encoding network and a first output layer connected to each other, and the first encoding network includes a plurality of first encoding layers connected in sequence; the second voice recognition model includes a second encoding network and a second output layer connected to each other, and the second encoding network includes a plurality of second encoding layers connected in sequence, and the number of the second encoding layers is less than the number of the first encoding layers. Specifically, the third training module 1304 is configured to:
[0216] Randomly initialize the encoding parameters of each second encoding layer;
[0217] Use the output parameters of the first output layer as the output parameters of the second output layer;
[0218] Initialize the second voice recognition model according to the encoding parameters of each second encoding layer and the output parameters of each second output layer.
[0219] Furthermore, the first voice recognition model further includes a first convolutional network connected to the first encoding network, and the first convolutional network includes a plurality of first convolutional layers connected in sequence; the second voice recognition model further includes a second convolutional network connected to the second encoding network, and the second convolutional network includes a plurality of second convolutional layers connected in sequence, and the number of the second convolutional layers is equal to the number of the first convolutional layers. The second convolutional layer before a preset convolutional layer is a target convolutional layer, and the feature dimension of the target convolutional layer is less than the feature dimension of the corresponding first convolutional layer of the target convolutional layer. Specifically, the third training module 1304 is configured to:
[0220] Randomly initialize the convolutional parameters of the preset convolutional layer and the convolutional parameters of the target convolutional layer;
[0221] Use the convolutional parameters of the corresponding first convolutional layer of the remaining convolutional layers as the convolutional parameters of the remaining convolutional layers; wherein, the remaining convolutional layers are the remaining second convolutional layers among the plurality of second convolutional layers except the preset convolutional layer and the target convolutional layer;
[0222] Initialize the second voice recognition model according to the convolutional parameters of each second convolutional layer, the encoding parameters of each second encoding layer, and the output parameters of each second output layer.
[0223] Furthermore, specifically, the third training module 1304 is configured to:
[0224] Input the first sample voice data set into the first voice recognition model, obtain first convolutional features output by the first convolutional layer corresponding to the preset convolutional layer, and obtain second convolutional features output by the last first convolutional layer;
[0225] Input the first sample voice dataset into the second speech recognition model, obtain the third convolutional feature output by the preset convolutional layer, and obtain the fourth convolutional feature output by the last second convolutional layer;
[0226] Determine the first convolutional loss value according to the first convolutional feature and the third convolutional feature, and determine the second convolutional loss value according to the second convolutional feature and the fourth convolutional feature;
[0227] Determine the target convolutional loss value according to the first convolutional loss value and the second convolutional loss value, and perform distillation training on the second convolutional network according to the target convolutional loss value.
[0228] Furthermore, the above-mentioned third training module 1304 is specifically used for:
[0229] Input the first sample voice dataset into the first speech recognition model, determine the benchmark coding layers with the same number as the second coding layer from the first speech recognition model, and obtain the first coding features output by each benchmark coding layer;
[0230] Input the first sample voice dataset into the second speech recognition model, and obtain the second coding features output by each second coding layer;
[0231] Determine the distillation fitting parameters of each second coding layer, and adjust the feature dimension of the corresponding second coding feature according to the distillation fitting parameters;
[0232] Determine the coding layer loss value corresponding to each second coding layer according to the second coding feature with adjusted feature dimension and the corresponding first coding feature, determine the target coding loss value according to each coding layer loss value, and perform distillation training on the second coding network according to the target coding loss value.
[0233] Furthermore, the above-mentioned third training module 1304 is specifically used for:
[0234] Obtain the first sample recognition result output by the first output layer and the second sample recognition result output by the second output layer;
[0235] Determine the target output loss value according to the first sample recognition result and the second sample recognition result;
[0236] Perform distillation training on the second coding network according to the target coding loss value and the target output loss value.
[0237] Furthermore, the above-mentioned third training module 1304 is specifically used for:
[0238] Perform distillation training on the second coding network according to the target coding loss value, and perform distillation training on the second coding network again according to the target output loss value;
[0239] Alternatively, the target encoding loss value and the target output loss value are weighted to obtain a target model loss value, and the second encoding network is distilled and trained according to the target model loss value.
[0240] Further, the third training module 1304 is specifically configured to:
[0241] Adjust the feature dimension of the second output layer in the target speech recognition model according to the distillation fitting parameters of the last second encoding layer;
[0242] Perform speech recognition on the target speech data based on the target speech recognition model with the adjusted feature dimension.
[0243] Further, the third training module 1304 is specifically configured to:
[0244] Prune the distillation fitting parameters of each target encoding layer in the target speech recognition model, and perform speech recognition on the target speech data based on the pruned target speech recognition model; wherein, the target encoding layer is the second encoding layer other than the last second encoding layer;
[0245] Alternatively, prune the distillation fitting parameters of each second encoding layer in the target speech recognition model, and perform speech recognition on the target speech data based on the pruned target speech recognition model.
[0246] Further, the original model includes an original convolutional network and an original encoding network connected in sequence, and the first training module 1302 is specifically configured to:
[0247] Input the first sample speech data set into the original model, perform a masking operation on the original convolutional features output by the original convolutional network to obtain masked convolutional features;
[0248] Perform a product quantization operation on the original convolutional features to obtain quantized convolutional features;
[0249] Obtain the masked encoding features output by the original encoding network after processing the masked convolutional features;
[0250] Determine a first original loss value according to the masked encoding features and the quantized convolutional features;
[0251] Perform unsupervised training on the original model according to the first original loss value.
[0252] Further, the first training module 1302 is specifically configured to:
[0253] Obtain the first quantity of the quantization codebooks and the second quantity of the cluster centers in each quantization codebook when performing the product quantization operation;
[0254] Determine the probability distribution of any one cluster center in each quantization codebook being selected;
[0255] Determine a second original loss value according to a first quantity, a second quantity, and a probability distribution;
[0256] Perform unsupervised training on the original model according to the first original loss value and the second original loss value.
[0257] The above-mentioned speech recognition device 1300 and the foregoing speech recognition method are based on the same inventive concept. Therefore, using the speech recognition device 1300 to perform speech recognition on target speech data can improve both the speech recognition efficiency and the speech recognition accuracy.
[0258] The electronic device for executing the above-mentioned speech recognition method provided by the embodiments of the present application may be a terminal. Refer to Figure 14 , Figure 14 which is a partial structural block diagram of the terminal provided by the embodiments of the present application. The terminal includes components such as a Radio Frequency (RF) circuit 1410, a memory 1420, an input unit 1430, a display unit 1440, a sensor 1450, an audio circuit 1460, a wireless fidelity (WiFi) module 1470, a processor 1480, and a power supply 1490. Those skilled in the art can understand that Figure 14 the terminal structure shown in
[0259] does not limit the terminal. It may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0260] The RF circuit 1410 can be used for receiving and sending signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 1480 for processing; in addition, the data designed for the uplink is sent to the base station.
[0261] The input unit 1430 can be used to receive input digital or character information, and generate key signal inputs related to the settings and function controls of the terminal. Specifically, the input unit 1430 may include a touch panel 1431 and other input devices 1432.
[0262] The display unit 1440 can be used to display the input information or the provided information and various menus of the terminal. The display unit 1440 may include a display panel 1441.
[0263] The audio circuit 1460, the speaker 1461, and the microphone 1462 can provide an audio interface.
[0264] In this embodiment, the processor 1480 included in the terminal can execute the speech recognition method of the previous embodiment.
[0265] The electronic device provided in the embodiment of the present application for executing the above speech recognition method can also be a server. Refer to Figure 15 , Figure 15 which is a partial structural block diagram of the server provided in the embodiment of the present application. The server 1500 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 1522 (for example, one or more processors) and a memory 1532, and one or more storage media 1530 (for example, one or more mass storage devices) for storing application programs 1542 or data 1544. Among them, the memory 1532 and the storage media 1530 can be transient storage or persistent storage. The program stored in the storage media 1530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 1500. Further, the central processor 1522 can be set to communicate with the storage media 1530 and execute a series of instruction operations in the storage media 1530 on the server 1500.
[0266] The server 1500 may further include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558, and / or one or more operating systems 1541, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.
[0267] The processor in the server 1500 can be used to execute the speech recognition method.
[0268] The embodiment of the present application also provides a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the speech recognition methods of the foregoing embodiments.
[0269] The embodiment of the present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the speech recognition method described above.
[0270] In the description of the present application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0271] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the relationship between related objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0272] It should be understood that in the description of the embodiments of the present application, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number.
[0273] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0274] The unit described as a separation component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0275] In addition, each functional unit in various embodiments of the present application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0276] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0277] It should also be understood that the various implementation manners provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.
[0278] The above has specifically described the preferred embodiments of the present application, but the present application is not limited to the above implementation manners. Those skilled in the art can also make various equivalent deformations or substitutions without violating the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.
Claims
1. A voice recognition method, characterized in that, Including: Obtaining a first sample speech data set without annotation and a second sample speech data set with annotation; Initializing an original model, and performing unsupervised training on the original model based on the first sample speech data set to obtain a basic speech processing model; wherein, the basic speech processing model includes a plurality of sequentially connected basic encoding layers; Pruning the basic encoding layers after a preset encoding layer among the plurality of basic encoding layers, and performing supervised training on the pruned basic speech processing model based on the second sample speech data set to obtain a first speech recognition model. Pruning the basic encoding layers after a preset encoding layer among the plurality of basic encoding layers means deleting the basic encoding layers after a preset encoding layer among the plurality of basic encoding layers; Initializing a second speech recognition model, and performing distillation training on the second speech recognition model based on the first sample speech data set with the first speech recognition model as a benchmark to obtain a target speech recognition model; Performing speech recognition on target speech data based on the target speech recognition model to obtain a target recognition result corresponding to the target speech data.
2. The voice recognition method according to claim 1, wherein, The first speech recognition model includes a first encoding network and a first output layer connected to each other, and the first encoding network includes a plurality of sequentially connected first encoding layers; the second speech recognition model includes a second encoding network and a second output layer connected to each other, and the second encoding network includes a plurality of sequentially connected second encoding layers, and the number of the second encoding layers is less than the number of the first encoding layers. Initializing the second speech recognition model includes: Randomly initializing the encoding parameters of each of the second encoding layers; Taking the output parameters of the first output layer as the output parameters of the second output layer; Initializing the second speech recognition model according to the encoding parameters of each of the second encoding layers and the output parameters of each of the second output layers.
3. The speech recognition method according to claim 2, characterized in that, The first speech recognition model further includes a first convolutional network connected to the first encoding network, and the first convolutional network includes a plurality of sequentially connected first convolutional layers; the second speech recognition model further includes a second convolutional network connected to the second encoding network, and the second convolutional network includes a plurality of sequentially connected second convolutional layers, and the number of the second convolutional layers is equal to the number of the first convolutional layers. The second convolutional layer before a preset convolutional layer is a target convolutional layer, and the feature dimension of the target convolutional layer is less than the feature dimension of the corresponding first convolutional layer of the target convolutional layer. Initializing the second speech recognition model according to the encoding parameters of each of the second encoding layers and the output parameters of each of the second output layers includes: Randomly initializing the convolutional parameters of the preset convolutional layer and the convolutional parameters of the target convolutional layer; Taking the convolutional parameters of the corresponding first convolutional layers of the remaining convolutional layers as the convolutional parameters of the remaining convolutional layers; wherein, the remaining convolutional layers are the remaining second convolutional layers among the plurality of second convolutional layers except the preset convolutional layer and the target convolutional layer. Initialize the second speech recognition model according to the convolution parameters of each of the second convolution layers, the encoding parameters of each of the second encoding layers, and the output parameters of each of the second output layers.
4. The speech recognition method according to claim 3, wherein The distillation training of the second speech recognition model based on the first sample speech dataset includes: Input the first sample speech dataset into the first speech recognition model, obtain the first convolution features output by the first convolution layer corresponding to the preset convolution layer, and obtain the second convolution features output by the last first convolution layer; Input the first sample speech dataset into the second speech recognition model, obtain the third convolution features output by the preset convolution layer, and obtain the fourth convolution features output by the last second convolution layer; Determine the first convolution loss value according to the first convolution features and the third convolution features, and determine the second convolution loss value according to the second convolution features and the fourth convolution features; Determine the target convolution loss value according to the first convolution loss value and the second convolution loss value, and perform distillation training on the second convolution network according to the target convolution loss value.
5. The speech recognition method according to claim 2, wherein The distillation training of the second speech recognition model based on the first sample speech dataset includes: Input the first sample speech dataset into the first speech recognition model, determine the benchmark encoding layers with the same number as the second encoding layer from the first speech recognition model, and obtain the first encoding features output by each of the benchmark encoding layers; Input the first sample speech dataset into the second speech recognition model, and obtain the second encoding features output by each of the second encoding layers; Determine the distillation fitting parameters of each of the second encoding layers, and adjust the feature dimensions of the corresponding second encoding features according to the distillation fitting parameters; Determine the encoding layer loss values corresponding to each of the second encoding layers according to the second encoding features with adjusted feature dimensions and the corresponding first encoding features, determine the target encoding loss value according to each of the encoding layer loss values, and perform distillation training on the second encoding network according to the target encoding loss value.
6. The voice recognition method according to claim 5, characterized in that, The distillation training of the second encoding network according to the target encoding loss value includes: Obtain the first sample recognition result output by the first output layer and the second sample recognition result output by the second output layer; Determine the target output loss value according to the first sample recognition result and the second sample recognition result; Perform distillation training on the second encoding network according to the target encoding loss value and the target output loss value.
7. The voice recognition method according to claim 6, characterized in that The distillation training of the second encoding network according to the target encoding loss value and the target output loss value includes: Perform distillation training on the second encoding network according to the target encoding loss value, and perform distillation training on the second encoding network again according to the target output loss value; Alternatively, weight the target encoding loss value and the target output loss value to obtain the target model loss value, and perform distillation training on the second encoding network according to the target model loss value.
8. The voice recognition method according to any one of claims 5 to 7, characterized in that, Performing speech recognition on the target speech data based on the target speech recognition model includes: Adjusting the feature dimension of the second output layer in the target speech recognition model according to the distillation fitting parameters of the last second coding layer; Performing speech recognition on the target speech data based on the target speech recognition model with the adjusted feature dimension.
9. The voice recognition method according to any one of claims 5 to 7, characterized in that, Performing speech recognition on the target speech data based on the target speech recognition model includes: Pruning the distillation fitting parameters of each target coding layer in the target speech recognition model, and performing speech recognition on the target speech data based on the pruned target speech recognition model; wherein, the target coding layer is the second coding layer except the last second coding layer; Alternatively, pruning the distillation fitting parameters of each second coding layer in the target speech recognition model, and performing speech recognition on the target speech data based on the pruned target speech recognition model.
10. The voice recognition method according to any one of claims 1 to 7, characterized in that, The original model includes an original convolutional network and an original coding network connected in sequence. The unsupervised training of the original model based on the first sample speech dataset includes: Inputting the first sample speech dataset into the original model, performing a masking operation on the original convolutional features output by the original convolutional network to obtain masked convolutional features; Performing a product quantization operation on the original convolutional features to obtain quantized convolutional features; Obtaining the masked coding features output after the original coding network processes the masked convolutional features; Determining a first original loss value according to the masked coding features and the quantized convolutional features; Performing unsupervised training on the original model according to the first original loss value.
11. The voice recognition method according to claim 10, characterized in that, The performing unsupervised training on the original model according to the first original loss value includes: Obtaining the first quantity of the quantization codebooks and the second quantity of the cluster centers in each quantization codebook when performing the product quantization operation; Determining the probability distribution of any one of the cluster centers in each quantization codebook being selected; Determining a second original loss value according to the first quantity, the second quantity and the probability distribution; Performing unsupervised training on the original model according to the first original loss value and the second original loss value.
12. A voice recognition device, characterized in that, Including: A sample data acquisition module, configured to acquire an unlabeled first sample speech dataset and a labeled second sample speech dataset; A first training module, configured to initialize the original model, and perform unsupervised training on the original model based on the first sample speech dataset to obtain a basic speech processing model; wherein, the basic speech processing model includes a plurality of basic coding layers connected in sequence; A second training module, configured to prune the basic coding layers after a preset coding layer in the plurality of basic coding layers, and perform supervised training on the pruned basic speech processing model based on the second sample speech dataset to obtain a first speech recognition model. Pruning the basic coding layers after a preset coding layer in the plurality of basic coding layers means deleting the basic coding layers after a preset coding layer in the plurality of basic coding layers; The third training module is used to initialize the second speech recognition model, and based on the first speech recognition model as a benchmark, distill and train the second speech recognition model based on the first sample speech data set to obtain a target speech recognition model; The speech recognition module is used to perform speech recognition on target speech data based on the target speech recognition model to obtain a target recognition result corresponding to the target speech data.
13. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the speech recognition method according to any one of claims 1 to 11.
14. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech recognition method according to any one of claims 1 to 11.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech recognition method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Compressed speech recognition model optimizing method and system
CN108389576A
Speech recognition neural network model and training method thereof, and speech recognition method
CN112687263A