Speech recognition model processing method, speech recognition method and device
By introducing a decoder into the speech recognition model, using speech semantic joint features to improve speech recognition accuracy, the problem of low recognition accuracy of the existing non-autoregressive speech recognition model is solved, and more efficient speech recognition effect is achieved.
Patent Information
- Application Number
- CN202111292319.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-11-03
AI Technical Summary
The existing non-autoregressive speech recognition model has shortcomings in terms of speech recognition accuracy, and only uses the information of the speech signal at the speech level, resulting in low recognition accuracy.
By obtaining sample signals and labeled character sequences, input a speech recognition model to obtain speech features and predicted character sequences, and decoder decodes the speech features based on the semantic features of the forward character sequence to generate speech semantic joint features, and finally predict based on the joint features to improve the recognition accuracy.
By jointly training the speech recognition model and the decoder, the context information at the semantic level can be distilled into the speech recognition model, thereby significantly improving the accuracy of speech recognition.
Smart Images

Figure CN114360502B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method for processing a speech recognition model, a speech recognition method and a device. Background Art
[0002] With the development of computer technology and artificial intelligence technology, voice recognition is required in many scenarios, such as virtual robot interaction scenarios, smart device control scenarios, machine translation scenarios, text conversion scenarios of voice messages, etc. For example, the terminal receives the voice signal input by the user through the virtual robot program installed on the terminal, performs voice recognition on the voice signal to obtain a voice recognition result, and performs corresponding operations based on the voice recognition result. For another example, a voice control client is installed on the smart device, and the smart device receives the voice signal input by the user through the voice control client, performs voice recognition on the voice signal to obtain a voice recognition result, obtains a control instruction based on the voice recognition result, and then performs corresponding operations.
[0003] At present, non-autoregressive speech recognition models have been widely used due to their advantages such as fast speech recognition speed. However, non-autoregressive speech recognition models only use the information of speech signals at the speech level and have the disadvantage of low recognition accuracy. Summary of the invention
[0004] Based on this, it is necessary to provide a processing method for a speech recognition model, a speech recognition method and a speech recognition device that can improve the accuracy of speech recognition in response to the above technical problems.
[0005] A method for processing a speech recognition model, the method comprising:
[0006] Obtaining sample signals and corresponding labeled character sequences;
[0007] Inputting the sample signal into a speech recognition model to obtain speech features corresponding to the sample signal and a first predicted character sequence output based on the speech features;
[0008] Inputting a forward character sequence corresponding to the marked character sequence into a decoder, wherein the forward character sequence is generated based on the previous character of each character in the marked character sequence;
[0009] In the decoder, the speech feature is decoded according to the semantic feature corresponding to the forward character sequence to obtain the speech and semantic joint feature corresponding to the sample signal, and prediction is performed based on the speech and semantic joint feature to obtain a second predicted character sequence corresponding to the sample signal;
[0010] The speech recognition model and the decoder are jointly trained based on the speech recognition loss calculated based on the labeled character sequence and the first predicted character sequence, and the semantic recognition loss calculated based on the labeled character sequence and the second predicted character sequence.
[0011] A processing device for a speech recognition model, the device comprising:
[0012] An acquisition module, used to acquire a sample signal and a corresponding labeled character sequence;
[0013] An encoding module, used for inputting the sample signal into a speech recognition model to obtain speech features corresponding to the sample signal and a first predicted character sequence output based on the speech features;
[0014] An input module, used for inputting a forward character sequence corresponding to the marked character sequence into a decoder, wherein the forward character sequence is generated based on the previous character of each character in the marked character sequence;
[0015] A decoding module, configured to decode the speech feature in the decoder according to the semantic feature corresponding to the forward character sequence, obtain the speech and semantic joint feature corresponding to the sample signal, and make a prediction based on the speech and semantic joint feature to obtain a second predicted character sequence corresponding to the sample signal;
[0016] A training module is used to jointly train the speech recognition model and the decoder based on the speech recognition loss calculated according to the labeled character sequence and the first predicted character sequence, and the semantic recognition loss calculated according to the labeled character sequence and the second predicted character sequence.
[0017] In one embodiment, the encoding module is also used to: input the sample signal into the speech recognition model; output the speech features corresponding to the sample signal through the encoder of the speech recognition model; and output the first predicted character sequence based on the speech features through a classifier connected to the encoder in the speech recognition model.
[0018] In one embodiment, the encoder includes a feature extraction network and a self-attention-based speech context network; the encoding module is also used to: input the sample signal into the encoder to obtain a speech vector sequence corresponding to the sample signal output by the feature extraction network in the encoder; perform random masking on the speech vectors in the speech vector sequence; input the masked speech vector sequence into the speech context network to obtain contextual speech features output by the speech context network as speech features corresponding to the sample signal.
[0019] In one embodiment, the decoder includes a vectorization layer, a semantic context network based on self-attention, and a speech semantic context network based on cross-attention; the decoding module is also used to: convert the forward character sequence into a corresponding forward character vector sequence through the vectorization layer of the decoder, and input the forward character vector sequence into the semantic context network; calculate the contextual semantic features corresponding to the forward character sequence based on the forward character vector sequence through the semantic context network as the semantic features corresponding to the forward character sequence; calculate the speech semantic joint features corresponding to the sample signal based on the semantic features corresponding to the forward character sequence and the speech features through the speech semantic context network.
[0020] In one embodiment, the decoding module is further used to: input the speech-semantics joint feature into the classifier of the decoder; and output a second predicted character sequence corresponding to the sample signal based on the speech-semantics joint feature through the classifier.
[0021] In one embodiment, the speech recognition model includes an encoder and a classifier connected to the encoder; the encoder is a pre-trained encoder obtained by self-supervised training using unlabeled sample signals; the training module is also used to: perform supervised training on the decoder and the classifier of the speech recognition model according to the speech recognition loss and the semantic recognition loss; when the supervised training stop condition is met, perform supervised training on the decoder and the speech recognition model according to the speech recognition loss and the semantic recognition loss.
[0022] In one embodiment, the encoder is a pre-trained encoder obtained by self-supervised training using an unlabeled sample signal; the speech recognition model also includes a pre-training module, which is used to: obtain the unlabeled sample signal; input the unlabeled sample signal into an initial encoder to obtain a speech vector sequence corresponding to the unlabeled sample signal output by a feature extraction network in the initial encoder; perform a quantization operation on the speech vector sequence to obtain a speech quantization vector sequence; after randomly masking the speech vectors in the speech vector sequence, determine the masked speech vector; input the masked speech vector sequence into a speech context network of the initial encoder to obtain a predicted speech vector corresponding to the masked speech vector output by the speech context network; construct a self-supervised training loss based on the difference between the speech quantization vector corresponding to the masked speech vector in the speech quantization vector sequence and the predicted speech vector; after updating the network parameters of the initial encoder according to the self-supervised training loss, return to the step of obtaining the unlabeled sample signal to continue training until the training is completed, and obtain the pre-trained encoder.
[0023] In one embodiment, the training module is also used to: construct the speech recognition loss based on the difference between the labeled character sequence and the first predicted character sequence; construct the semantic recognition loss based on the difference between the labeled character sequence and the second predicted character sequence; weightedly sum the speech recognition loss and the semantic recognition loss according to a preset loss weighting coefficient to obtain a target loss; and jointly train the speech recognition model and the decoder according to the target loss.
[0024] In one embodiment, the processing device of the speech recognition model also includes a speech recognition module, which is used to: obtain a signal to be recognized; input the signal to be recognized into a trained speech recognition model to obtain speech features output by an encoder in the speech recognition model, and a speech recognition result output by a classifier in the speech recognition model based on the speech features.
[0025] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the processing method of the speech recognition model when executing the computer program.
[0026] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the processing method of the speech recognition model.
[0027] A computer program includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the processing method of the above-mentioned speech recognition model.
[0028] The processing method, device, computer equipment and storage medium of the above-mentioned speech recognition model input a sample signal into the speech recognition model to obtain speech features corresponding to the sample signal and a first predicted character sequence output based on the speech features, and input a forward character sequence corresponding to the marked character sequence into a decoder. In the decoder, the speech features are decoded according to the semantic features corresponding to the forward character sequence to obtain speech and semantic joint features corresponding to the sample signal. Since the forward character sequence is generated based on the previous character of each character in the marked character sequence, the speech features output by the encoder are decoded and encoded according to the semantic features corresponding to the forward character sequence. The speech and semantic joint features obtained carry context information at the semantic level. A second predicted character sequence corresponding to the sample signal is predicted based on the speech and semantic joint features. The semantic recognition loss constructed based on the second predicted character sequence and the marked character sequence is used to assist the speech recognition model in training, so that the context information at the semantic level can be distilled into the speech recognition model, thereby improving the recognition accuracy of the speech recognition model.
[0029] A speech recognition method, the method comprising:
[0030] Acquire a signal to be identified;
[0031] Inputting the signal to be recognized into a trained speech recognition model to obtain speech features output by an encoder in the speech recognition model, and speech recognition results output by a classifier in the speech recognition model based on the speech features;
[0032] Among them, the speech recognition model and decoder are obtained by joint training based on speech recognition loss and semantic recognition loss, the speech recognition loss is calculated based on a first predicted character sequence and a labeled character sequence corresponding to the sample signal, and the semantic recognition loss is calculated based on a second predicted character sequence and the labeled character sequence. The first predicted character sequence is obtained after classification based on the speech features output by the encoder, and the second predicted character sequence is obtained by predicting the speech and semantic joint features obtained by decoding the speech features using the semantic features corresponding to the forward character sequence corresponding to the labeled character sequence by the decoder, and the forward character sequence is generated based on the previous character of each character in the labeled character sequence.
[0033] A speech recognition device, comprising:
[0034] An acquisition module, used for acquiring a signal to be identified;
[0035] A speech recognition module, used for inputting the signal to be recognized into a trained speech recognition model, obtaining speech features output by an encoder in the speech recognition model, and a speech recognition result output by a classifier in the speech recognition model based on the speech features;
[0036] Among them, the speech recognition model and decoder are obtained by joint training based on speech recognition loss and semantic recognition loss, the speech recognition loss is calculated based on a first predicted character sequence and a labeled character sequence corresponding to the sample signal, and the semantic recognition loss is calculated based on a second predicted character sequence and the labeled character sequence. The first predicted character sequence is obtained after classification based on the speech features output by the encoder, and the second predicted character sequence is obtained by predicting the speech and semantic joint features obtained by decoding the speech features using the semantic features corresponding to the forward character sequence corresponding to the labeled character sequence by the decoder, and the forward character sequence is generated based on the previous character of each character in the labeled character sequence.
[0037] A computer device comprises a memory and a processor, wherein the memory stores a computer program and the processor implements the steps of the above-mentioned speech recognition method when executing the computer program.
[0038] A computer-readable storage medium stores a computer program, which implements the steps of the above-mentioned speech recognition method when executed by a processor.
[0039] A computer program includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the above-mentioned speech recognition method.
[0040] The above-mentioned speech recognition method, apparatus, computer equipment and storage medium input the signal to be recognized into a trained speech recognition model, obtain the speech features output by the encoder in the speech recognition model, and the speech recognition results output by the classifier in the speech recognition model based on the speech features. Since the trained speech recognition model can use contextual information at the semantic level for speech recognition, the speech recognition accuracy can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 An application environment diagram of a method for processing a speech recognition model in one embodiment;
[0042] Figure 2 is a schematic diagram of a speech recognition scenario in one embodiment;
[0043] Figure 3 is a flowchart of a method for processing a speech recognition model in one embodiment;
[0044] Figure 4 A schematic diagram of a decoder-assisted training of a speech recognition model in one embodiment;
[0045] Figure 5 A schematic diagram of obtaining speech features corresponding to a sample signal through an encoder in one embodiment;
[0046] Figure 6 A schematic diagram of performing self-supervised pre-training on an initial encoder in one embodiment;
[0047] Figure 7 A schematic diagram of another embodiment of a decoder-assisted training speech recognition model;
[0048] Figure 8 is a flowchart of a method for processing a speech recognition model in one embodiment;
[0049] Fig. 9A schematic diagram of a decoder-assisted training of a speech recognition model in yet another embodiment;
[0050] Fig.10 A schematic diagram of a test result in one embodiment;
[0051] Fig.11 is a flowchart of a speech recognition method in one embodiment;
[0052] Fig.12 is a structural block diagram of a processing device for a speech recognition model in one embodiment;
[0053] Fig.13 is a structural block diagram of a speech recognition device in one embodiment;
[0054] Fig.14 is an internal structure diagram of a computer device in one embodiment;
[0055] Fig.15 FIG. 4 is a diagram showing the internal structure of a computer device in another embodiment. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0057] The processing method of the speech recognition model and the speech recognition method provided in the embodiments of the present application involve artificial intelligence (AI) technology. Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making.
[0058] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0059] The processing method of the speech recognition model provided in the embodiment of the present application mainly relates to the machine learning technology (Machine Learning, ML) of artificial intelligence. Machine learning is a multi-field interdisciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all fields of artificial intelligence. Machine learning and deep learning generally include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by formula.
[0060] For example, in an embodiment of the present application, a speech recognition model and a decoder are jointly trained based on speech recognition loss and semantic recognition loss, and finally a speech recognition model for recognizing speech signals is obtained.
[0061] The speech recognition method provided in the embodiment of the present application mainly relates to speech technology of artificial intelligence. The key technologies of speech technology include automatic speech recognition technology, speech synthesis technology and voiceprint recognition technology. Enabling computers to listen, see, speak and feel is the development direction of human-computer interaction in the future, among which speech has become one of the most promising human-computer interaction methods in the future.
[0062] For example, in an embodiment of the present application, the encoder in the trained speech recognition model outputs the speech features corresponding to the signal to be recognized, and the classifier in the trained speech recognition model outputs the speech recognition result based on the speech features.
[0063] The processing method of the speech recognition model and the speech recognition method provided in the embodiments of the present application may also involve blockchain technology. Blockchain is a new application mode of computer technologies such as distributed data storage, point-to-point transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer.
[0064] For example, in an embodiment of the present application, the server can be a blockchain node in a blockchain network, the trained speech recognition model can be stored on the blockchain, and the signal to be recognized can be uploaded to the data block of the blockchain to perform speech recognition on the signal to be recognized.
[0065] The speech recognition model processing method and speech recognition method provided in this application can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The terminal 102 can be, but is not limited to, various smart phones, tablet computers, laptops, desktop computers, portable wearable devices, smart speakers, vehicle-mounted devices, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0066] In one embodiment, terminal 102 obtains a sample signal and a corresponding annotated character sequence, and sends the sample signal and the corresponding annotated character sequence to server 104. Server 104 inputs the sample signal into a speech recognition model to obtain speech features corresponding to the sample signal, and a first predicted character sequence output based on the speech features; a forward character sequence corresponding to the annotated character sequence is input into a decoder, and the forward character sequence is generated based on the previous character of each character in the annotated character sequence; in the decoder, the speech features are decoded according to the semantic features corresponding to the forward character sequence to obtain speech and semantic joint features corresponding to the sample signal, and prediction is performed based on the speech and semantic joint features to obtain a second predicted character sequence corresponding to the sample signal; based on the speech recognition loss calculated based on the annotated character sequence and the first predicted character sequence, and the semantic recognition loss calculated based on the annotated character sequence and the second predicted character sequence, the speech recognition model and the decoder are jointly trained.
[0067] The processing method of the speech recognition model provided in the embodiment of the present application may be performed by the processing device of the speech recognition model provided in the embodiment of the present application, or by a computer device integrated with the processing device of the speech recognition model, wherein the processing device of the speech recognition model may be implemented in hardware or software. The computer device may be Figure 1 The terminal 102 or server 104 shown in .
[0068] In one embodiment, the terminal 102 obtains a signal to be recognized, and sends the signal to be recognized to the server 104. The server 104 inputs the signal to be recognized into a trained speech recognition model to obtain speech features output by the encoder in the speech recognition model, and a speech recognition result output by the classifier in the speech recognition model based on the speech features; wherein the speech recognition model and the decoder are jointly trained based on speech recognition loss and semantic recognition loss, the speech recognition loss is calculated based on a first predicted character sequence and a labeled character sequence corresponding to the sample signal, the semantic recognition loss is calculated based on a second predicted character sequence and the labeled character sequence, the first predicted character sequence is obtained by classification based on the speech features output by the encoder, the second predicted character sequence is obtained by predicting the speech and semantic joint features obtained by decoding the speech features using the semantic features corresponding to the forward character sequence corresponding to the labeled character sequence by the decoder, and the forward character sequence is generated based on the previous character of each character in the labeled character sequence.
[0069] The speech recognition method provided in the embodiment of the present application may be implemented by the speech recognition device provided in the embodiment of the present application, or by a computer device integrated with the speech recognition device, wherein the speech recognition device may be implemented in hardware or software. The computer device may be Figure 1 The terminal 102 or server 104 shown in FIG.
[0070] The speech recognition method provided in the embodiment of the present application can be applied to speech interaction scenarios, such as virtual robot interaction scenarios, smart device control scenarios, machine translation scenarios, text conversion scenarios of voice messages, etc. Speech interaction scenarios usually involve speech recognition technology and semantic recognition technology. Speech recognition technology can convert speech signals into text, and semantic recognition technology can recognize the intention of the text converted from speech signals. The speech recognition model trained in the present application is specifically applied to speech recognition technology.
[0071] For example, a virtual robot program is installed on the terminal, and the background server of the virtual robot program stores the speech recognition model trained by the present application. The terminal receives the voice signal input by the user through the virtual robot program, and the speech recognition model stored in the background server recognizes the text corresponding to the voice signal. The terminal can perform corresponding operations based on the text or the semantic recognition result of the text.
[0072] Taking the in-vehicle robot as an example, the in-vehicle robot is a social robot used in the in-vehicle smart cockpit scenario and is a type of service robot. The in-vehicle robot can respond to the user's voice input in the car and provide corresponding services, such as playing music / radio / news / e-books, navigation, checking the weather / nearby food, making calls, interactive chatting, etc.
[0073] Reference Figure 2 The speech recognition system of the vehicle-mounted robot may include an acoustic front-end module, a cloud-based speech recognition module, an offline speech recognition module, an offline / cloud-based semantic recognition module, and the like. Among them, the acoustic front-end module is used to provide functions such as speech noise reduction, sound source localization, and echo cancellation. The offline speech recognition module is used to provide functions such as fixed wake-up word wake-up, customized wake-up word wake-up, and offline speech recognition. The cloud-based speech recognition module may include a speech recognition model, and the speech recognition model is used to recognize the speech signal as text. Optionally, the speech recognition model can be divided into an acoustic model, a language model and a dictionary, and a decoder. The acoustic model is used to recognize the speech signal as phonemes, and the language model and the dictionary are used to convert the phonemes into text. The decoder is used to combine the acoustic model, the language model, and the dictionary to perform the entire search process from speech signal to text. The offline / cloud-based semantic recognition module is used to identify the intention of the text converted from the speech signal. The speech recognition model trained in this application can be applied to the cloud-based speech recognition module of the vehicle-mounted robot to improve the accuracy of the vehicle-mounted robot's speech recognition.
[0074] For another example, a voice control client is installed on a smart device, and the background server of the voice control client stores the voice recognition model trained by the present application. The smart device receives the voice signal input by the user through the voice control client, and the voice recognition model stored in the background server recognizes the text corresponding to the voice signal. The smart device can obtain control instructions based on the text or the semantic recognition result of the text, and then perform corresponding operations. Smart devices include but are not limited to smart home devices, etc.
[0075] For another example, a translation client is installed on the terminal, and the background server of the translation client stores the speech recognition model trained by the present application. The terminal receives a speech signal input by the user through the translation client, and the speech recognition model stored in the background server recognizes the text corresponding to the speech signal, translates the text or the semantic recognition result of the text, obtains the translation result, and the terminal outputs the translation result corresponding to the speech signal.
[0076] For another example, a conversation client is installed on the terminal, and the background server of the conversation client stores the speech recognition model trained by the present application. The terminal receives a voice message input by the user through the conversation client, and in response to the voice message conversion instruction, the speech recognition model stored in the background server recognizes the text corresponding to the voice message, and the terminal can display the text message corresponding to the voice message based on the text or the semantic recognition result of the text.
[0077] In one embodiment, Figure 3 As shown, a method for processing a speech recognition model is provided. This embodiment mainly applies this method to the above Figure 1 Taking the computer device (terminal 102 or server 104) in the example as an example, the following steps are included:
[0078] Step S302: Obtain a sample signal and a corresponding marked character sequence.
[0079] The sample signal is a speech signal used to train the speech recognition model, which has a time series characteristic. The sample signal can be an original analog sound signal, or a digital signal obtained by processing the original analog sound signal. The speech recognition model is an acoustic model that has speech recognition capability after training, and can specifically be a model for performing phoneme or text recognition on speech signals obtained by training using the sample signal as training data. Each sample signal has a corresponding annotated character sequence, which can be a phoneme sequence or a text sequence. For example, for the sample signal "Today's weather is good", the label data can be a phoneme sequence "bei3jing1 tian1 qi4 hao3" or a text sequence "Today's weather is good".
[0080] In one embodiment, the speech recognition model can be a non-autoregressive model based on CTC (Connectionist temporal classification). The CTC algorithm is used to solve the labeling problem of time series data. In traditional acoustic model training, for each frame of sample signal, it is necessary to know the corresponding label characters to effectively train, so it is necessary to align the sample signals before training, which is a time-consuming task. However, when the CTC loss function is used for training, there is no need to align the sample signals, and only the sample signals and the label character sequences corresponding to the sample signals need to be provided for training. The autoregressive (Autoregressive Translation, ART) model needs to predict the next word with the generated words during speech recognition. Although it has the characteristics of high recognition accuracy, the recognition speed is slow, while the non-autoregressive speech recognition model can generate predicted words at the same time within a specific number of iterations during speech recognition, and the recognition speed is fast, but the recognition accuracy is not as good as the autoregressive speech recognition model. In this application, a decoder is introduced when training a speech recognition model, and the speech recognition model and the decoder are jointly trained to help the speech recognition model "learn" the contextual information at the semantic level. During speech recognition, the decoder does not participate in the speech recognition process of the speech recognition model, thereby improving the recognition accuracy of the speech recognition model without affecting the recognition speed of the speech recognition model.
[0081] In one embodiment, a computer device obtains a sample signal and a corresponding sequence of labeled characters, and uses the sample signal and the corresponding sequence of labeled characters to train a speech recognition model.
[0082] Step S304: input the sample signal into a speech recognition model to obtain speech features corresponding to the sample signal and a first predicted character sequence output based on the speech features.
[0083] The speech feature is data describing the characteristics of the sample signal at the speech level. The speech feature can be in the form of a vector. For example, the speech signal is converted into "[1 0.2 4 0.3 0.10 0.8 0.7 0 0.7 2.1 5.2 0...]". The first predicted character sequence is the prediction result obtained by the speech recognition model based on the speech feature by performing speech recognition on the sample signal, which can be a phoneme sequence or a text sequence.
[0084] In one embodiment, a computer device inputs a sample signal into a speech recognition model; outputs speech features corresponding to the sample signal through an encoder of the speech recognition model; and outputs a first predicted character sequence based on the speech features through a classifier connected to the encoder in the speech recognition model.
[0085] In one embodiment, the speech recognition model may include an encoder and a classifier, the encoder is used to encode the sample signal to obtain speech features corresponding to the sample signal, the classifier is used to identify the characters corresponding to each time period signal in the sample signal based on the speech features, and output a first predicted character sequence corresponding to the sample signal.
[0086] For example, see Figure 4 , Figure 4 The figure is a schematic diagram of a decoder-assisted training of a speech recognition model in one embodiment. The computer device inputs a sample signal into the speech recognition model, and the encoder of the speech recognition model outputs speech features [c1 c2 c3 c4 c5] corresponding to the sample signal, and the classifier of the speech recognition model outputs a first predicted character sequence "w1w2w3w4w5" based on the speech features [c1 c2c3 c4 c5].
[0087] In one embodiment, the encoder may adopt a general encoder structure, such as CNN (Convolutional Neural Networks), RNN (Recurrent Neural Network), etc. The classifier may also adopt a general classifier structure, such as a linear classifier, etc.
[0088] Step S306: input a forward character sequence corresponding to the marked character sequence into a decoder, wherein the forward character sequence is generated based on the previous character of each character in the marked character sequence.
[0089] Among them, the forward character sequence is generated based on the previous character of each character in the labeled character sequence. For example, if the labeled character sequence L is "The weather is good today", then according to the previous character of each character in L, the forward character sequence corresponding to L can be obtained as " / The weather". Specifically, since there is no corresponding previous character for the first character "今" in L, " / " can be used to represent the previous character of "今", and thus the first character in the forward character sequence corresponding to L is obtained. Similarly, the second character in L is "日", and its previous character is "今", so the second character in the forward character sequence corresponding to L is obtained as "今". And so on, the forward character sequence corresponding to L is obtained as " / The weather".
[0090] Step S308, in the decoder, decode the speech feature according to the semantic feature corresponding to the forward character sequence, obtain the speech-semantic joint feature corresponding to the sample signal, and make a prediction based on the speech-semantic joint feature to obtain the second predicted character sequence corresponding to the sample signal.
[0091] Among them, the second predicted character sequence is the prediction result obtained by the decoder through speech recognition based on the speech-semantic joint feature, and it can be a phoneme sequence or a character sequence. The speech-semantic joint feature is the feature obtained by the decoder using the semantic information of the previous context of the labeled character sequence reflected by the forward character sequence to decode and re-encode the speech feature. As the name implies, the speech-semantic joint feature takes into account both the features of the speech signal at the speech level and the information of the labeled character sequence corresponding to the speech information at the semantic level.
[0092] In one embodiment, the decoder decodes and re-encodes the speech feature output by the encoder according to the semantic feature corresponding to the forward character sequence to obtain the speech-semantic joint feature. Thus, the speech-semantic joint feature carries context information at the semantic level. Making a prediction based on the speech-semantic joint feature to obtain the second predicted character sequence corresponding to the sample signal, and training the speech recognition model with the semantic recognition loss constructed by the second predicted character sequence and the labeled character sequence can distill the context information at the semantic level into the speech recognition model, helping the speech recognition model alleviate the deficiencies of the independence assumption and the inability to utilize the context information at the semantic level, thereby improving the recognition accuracy of the speech recognition model.
[0093] In one embodiment, the computer device inputs the forward character sequence corresponding to the labeled character sequence into the decoder. In the decoder, obtain the semantic feature corresponding to the forward character sequence, decode and re-encode the speech feature according to the semantic feature corresponding to the forward character sequence to obtain the speech-semantic joint feature, and make a prediction based on the speech-semantic joint feature to obtain the second predicted character sequence corresponding to the sample signal.
[0094] In one embodiment, the decoder may include a vectorization layer and a speech semantic context network based on cross attention. The vectorization layer is used to obtain semantic features corresponding to the forward character sequence. The feature dimensions of the semantic features corresponding to the forward character sequence are consistent with the feature dimensions of the speech features. The speech semantic context network based on cross attention is used to decode the speech features and encode them using the semantic features corresponding to the forward character sequence, so that the obtained speech semantic joint features carry context information at the semantic level.
[0095] For example, continue to refer to Figure 4 , the computer device inputs the forward character sequence " / x2x3x4x5" corresponding to the marked character sequence into the decoder 402, obtains the semantic feature [e1 e2 e3 e4 e5] corresponding to the forward character sequence through the vectorization layer of the decoder 402, inputs the semantic feature [e1 e2 e3 e4 e5] and the speech feature [c1 c2 c3 c4c5] extracted by the encoder into the speech semantic context network of the decoder 402, obtains the speech semantic joint feature [r1 r2 r3 r4 r5] through the speech semantic context network, and predicts based on the speech semantic joint feature [r1 r2 r3 r4 r5] to obtain the second predicted character sequence "y1y2y3y4y5" corresponding to the sample signal.
[0096] Step S310, jointly training a speech recognition model and a decoder based on a speech recognition loss calculated based on the labeled character sequence and the first predicted character sequence, and a semantic recognition loss calculated based on the labeled character sequence and the second predicted character sequence.
[0097] It can be understood that the universal loss function satisfies the requirements of the embodiments of the present application for speech recognition loss and semantic recognition loss, so the computer device can use the universal loss function to construct speech recognition loss and semantic recognition loss. Universal loss functions include cross entropy loss function, cosine similarity loss function, etc.
[0098] For example, continue to refer to Figure 4 , the computer device obtains the second predicted character sequence "y1y2y3y4y5" corresponding to the sample signal through the decoder 402, and the classifier of the speech recognition model outputs the first predicted character sequence "w1w2w3w4w5" corresponding to the sample signal based on the speech feature [c1 c2 c3 c4 c5]. Therefore, the computer device can calculate the speech recognition loss based on the labeled character sequence "x1x2x3x4x5" and the first predicted character sequence "w1w2w3w4w5", and calculate the semantic recognition loss based on the labeled character sequence "x1x2x3x4x5" and the second predicted character sequence "y1y2y3y4y5", and jointly train the speech recognition model and the decoder based on the speech recognition loss and the semantic recognition loss.
[0099] In one embodiment, the computer device performs weighted summation of speech recognition loss and semantic recognition loss according to a preset loss weighting coefficient to obtain a target loss; and jointly trains a speech recognition model and a decoder according to the target loss.
[0100] In one embodiment, the target loss is a comprehensive loss function composed of speech recognition loss and semantic recognition loss. The target loss can be expressed by the following formula:
[0101] L t =λ 1 L v +λ 2 L s
[0102] Among them, L t represents the target loss; L v represents the speech recognition loss, λ 1 Represents the loss weight coefficient corresponding to the speech recognition loss, for example, λ 1 It can be 0.3; L s represents the semantic recognition loss, λ 2 Represents the loss weight coefficient corresponding to the semantic recognition loss, for example, λ 2 The value can be 0.7.
[0103] In one embodiment, the computer device obtains the gradient corresponding to the current training based on the gradient descent algorithm in the direction of minimizing the target loss, and updates the network parameters of the speech recognition model and the decoder according to the gradient. The gradient descent algorithm can be a stochastic gradient descent algorithm, or an algorithm optimized based on the stochastic gradient descent algorithm, such as a stochastic gradient descent algorithm with a momentum term.
[0104] In the processing method of the above-mentioned speech recognition model, a sample signal is input into the speech recognition model to obtain speech features corresponding to the sample signal and a first predicted character sequence output based on the speech features, and a forward character sequence corresponding to the marked character sequence is input into the decoder. In the decoder, the speech features are decoded according to the semantic features corresponding to the forward character sequence to obtain the speech and semantic joint features corresponding to the sample signal. Since the forward character sequence is generated based on the previous character of each character in the marked character sequence, the speech features output by the encoder are decoded and encoded according to the semantic features corresponding to the forward character sequence. The speech and semantic joint features obtained carry context information at the semantic level. A second predicted character sequence corresponding to the sample signal is predicted based on the speech and semantic joint features. The semantic recognition loss constructed based on the second predicted character sequence and the marked character sequence is used to assist the speech recognition model in training, which can distill the context information at the semantic level into the speech recognition model, thereby improving the recognition accuracy of the speech recognition model.
[0105] In one embodiment, the encoder includes a feature extraction network and a self-attention-based speech context network; the encoder of the speech recognition model outputs speech features corresponding to the sample signal, including: inputting the sample signal into the encoder to obtain a speech vector sequence corresponding to the sample signal output by the feature extraction network in the encoder; performing random masking on the speech vectors in the speech vector sequence; inputting the masked speech vector sequence into the speech context network to obtain contextual speech features output by the speech context network as the speech features corresponding to the sample signal.
[0106] Among them, the speech vector sequence is a sequence composed of speech vectors, and the speech vector refers to the result obtained by mapping the speech signal to a high-dimensional vector space.
[0107] In one embodiment, the computer device inputs the sample signal into the encoder to obtain a speech vector sequence corresponding to the sample signal output by the feature extraction network in the encoder, and each speech vector in the speech vector sequence is a speech vector corresponding to the speech signal in each time period in the sample signal. For example, the computer device divides the sample signal into speech signals in time periods t1 to t5, inputs the sample signal into the encoder, and obtains a speech vector sequence [z1 z2 z3 z4 z5] output by the feature extraction network in the encoder, wherein the speech vector z1 is a speech vector corresponding to the speech signal in time period t1. It can be understood that the duration of each time period can be set according to the actual application, and this application does not make specific limitations.
[0108] In one embodiment, the encoder may include a feature extraction network and a self-attention-based speech context network. The feature extraction network is used to extract features from a sample signal to obtain a speech vector sequence corresponding to the sample signal. The self-attention-based speech context network is used to encode the speech vector sequence to obtain contextual speech features corresponding to the sample signal. The self-attention-based speech context network can encode the speech vector sequence using context information. At the same time, the self-attention mechanism ensures efficient parallel efficiency and direct connection to long-distance information, thereby improving the representation capability of speech features.
[0109] In one embodiment, the feature extraction network may adopt a general feature extraction network, such as CNN (Convolutional Neural Networks), RNN (Recurrent Neural Network), etc. The self-attention-based speech context network may adopt a general self-attention model, such as a Transformer model, a Conformer model, etc.
[0110] In one embodiment, the computer device performs random masking processing on the speech vectors in the speech vector sequence. It can be understood that the general masking processing method satisfies the requirements of the embodiment of the present application for masking processing, so the general masking processing method can be used to mask the speech vectors in the speech vector sequence. Optionally, the computer device can mask the speech vectors in the speech vector sequence through GELU (Gaussian Error Linerar Units).
[0111] In one embodiment, a computer device inputs a masked speech vector sequence into a speech context network, and calculates the self-attention corresponding to each speech vector in the masked speech vector sequence through the speech context network, wherein the self-attention can reflect the importance of each speech vector in the masked speech vector sequence; and outputs contextual speech features based on each speech vector and its corresponding self-attention through a feedforward neural network.
[0112] In one embodiment, the computer device calculates the similarity between each speech vector in the speech vector sequence after masking and the speech vector sequence after masking through the self-attention network in the speech context network, normalizes each similarity, and obtains the self-attention corresponding to each speech vector in the speech vector sequence after masking. Optionally, the computer device calculates the sum of the similarities corresponding to each speech vector in the speech vector sequence after masking, and calculates the ratio of the similarity corresponding to each speech vector in the speech vector sequence after masking to the sum of the similarities, as the self-attention corresponding to each speech vector in the speech vector sequence after masking.
[0113] For example, see Figure 5 , Figure 5The figure is a schematic diagram of obtaining speech features corresponding to a sample signal through an encoder in one embodiment. The computer device divides the sample signal into speech signals in time periods t1 to t5, inputs the sample signal into the encoder 502, and obtains a speech vector sequence [z1 z2 z3 z4 z5] output by the feature extraction network in the encoder 502. The computer device performs random masking on the speech vectors in the speech vector sequence [z1 z2 z3 z4 z5] to obtain a speech vector sequence [*z2*z4*] after masking. The computer device inputs the masked speech vector sequence [*z2*z4*] into the self-attention based speech context network 504, and respectively calculates the similarities s1, s2, s3, s4, s5 between each speech vector in the masked speech vector sequence [*z2*z4*] and the masked speech vector sequence [*z2*z4*] through the self-attention network in the self-attention based speech context network 504, and normalizes the similarities s1, s2, s3, s4, s5 to obtain self-attention p1, p2, p3, p4, p5. The computer device inputs the self-attention p1, p2, p3, p4, p5 and each speech vector in the masked speech vector sequence [*z2*z4*] into the feedforward neural network in the self-attention based speech context network 504 for encoding, and obtains the context speech features [c1 c2 c3 c4c5] output by the feedforward neural network.
[0114] In this embodiment, the encoder includes a speech context network based on self-attention. The speech context network based on self-attention can use context information to encode the speech vector sequence output by the feature extraction network. At the same time, the self-attention mechanism ensures efficient parallel efficiency and direct connection for long-distance information, thereby improving the representation ability of speech features.
[0115] In one embodiment, the encoder is a pre-trained encoder obtained by self-supervised training using an unlabeled sample signal; the method also includes: obtaining an unlabeled sample signal; inputting the unlabeled sample signal into an initial encoder to obtain a speech vector sequence corresponding to the unlabeled sample signal output by a feature extraction network in the initial encoder; performing a quantization operation on the speech vector sequence to obtain a speech quantization vector sequence; performing a random masking process on the speech vectors in the speech vector sequence to determine the masked speech vector; inputting the masked speech vector sequence into a speech context network of the initial encoder to obtain a predicted speech vector corresponding to the masked speech vector output by the speech context network; constructing a self-supervised training loss based on the difference between the speech quantization vector corresponding to the masked speech vector in the speech quantization vector sequence and the predicted speech vector; after updating the network parameters of the initial encoder according to the self-supervised training loss, returning to the step of obtaining the unlabeled sample signal to continue training until the training is completed, thereby obtaining a pre-trained encoder.
[0116] The unlabeled sample signal is a speech signal used for self-supervised pre-training of the encoder. The unlabeled sample signal has no corresponding labeled data. The initial encoder is the encoder to be self-supervised pre-trained.
[0117] In one embodiment, the computer device performs a quantization operation on the speech vector sequence to obtain a speech quantization vector sequence. The quantization operation can be a discretization process, such as a product quantization process, i.e., a Cartesian product. The infinite feature space is collapsed into a finite discrete space through the quantization operation, thereby enhancing the robustness of the feature and improving the representation ability of the feature.
[0118] In one embodiment, each speech vector in the speech quantization vector sequence includes a first speech vector corresponding to the speech signal in each time period of the sample signal, and the speech feature corresponding to the unlabeled sample signal also includes a second speech vector corresponding to the speech signal in each time period of the sample signal. The computer device constructs a speech vector prediction loss corresponding to the speech signal in the same time period based on the difference between the first speech vector and the second speech vector corresponding to the speech signal in the same time period, and fuses the speech vector prediction losses corresponding to the speech signals in each time period to obtain the self-supervised training loss.
[0119] In one embodiment, the speech vector prediction loss corresponding to the speech signal in time period t can be expressed by the following formula:
[0120]
[0121] Among them, L m represents the speech vector prediction loss corresponding to the speech signal in period t; q trepresents the first speech vector corresponding to the speech signal of time period t; c t represents the second speech vector corresponding to the speech signal of time period t; Q t Represents a set of candidate speech vectors, including q t and k error speech vectors; Indicates Q t Any error speech vector in ; sim(c t ,q t ) means c t With q t The correlation between Indicates c t and The correlation between them.
[0122] In one embodiment, the computer device obtains a preset loss weighting coefficient, and weighted sums the speech vector prediction losses corresponding to the speech signal in each time period according to the preset loss weighting coefficient to obtain the self-supervised training loss.
[0123] In one embodiment, the training ends when the number of training times reaches a preset number or the loss value calculated from the self-supervised training loss is less than a preset value.
[0124] For example, see Figure 6 , Figure 6The figure is a schematic diagram of self-supervised pre-training of the initial encoder in one embodiment. The computer device divides the unlabeled sample signal into speech signals of time periods t1 to t5, inputs the unlabeled sample signal into the initial encoder, and obtains a speech vector sequence [z1 z2 z3 z4z5] output by the feature extraction network in the initial encoder. The computer device performs a quantization operation on the speech vector sequence [z1 z2 z3 z4 z5] to obtain a speech quantization vector sequence [q1 q2 q3 q4 q5]. The computer device performs a random masking process on the speech vectors in the speech vector sequence [z1 z2 z3 z4 z5] to determine the masked speech vectors z1, z3, and z5. The computer device inputs the masked speech vector sequence [*z2*z4*] into the speech context network based on self-attention, and respectively calculates the similarities s1, s2, s3, s4, s5 between each speech vector in the masked speech vector sequence [*z2*z4*] and the masked speech vector sequence [*z2*z4*] through the self-attention network in the speech context network based on self-attention, and normalizes the similarities s1, s2, s3, s4, s5 to obtain self-attention p1, p2, p3, p4, p5. The computer device predicts the predicted speech vectors c1, c3, c5 corresponding to the masked speech vectors z1, z3, z5 based on the self-attention p1, p3, p5 through the feedforward neural network in the speech context network based on self-attention. The computer device trains the initial encoder based on the speech vector prediction loss constructed based on the difference between c1 and q1, the speech vector prediction loss constructed based on the difference between c3 and q3, and the speech vector prediction loss constructed based on the difference between c5 and q5.
[0125] In this embodiment, self-supervised pre-training of the encoder can improve the representation capability of the speech features output by the encoder, thereby improving subsequent training efficiency and training effect.
[0126] In one embodiment, a speech recognition model includes an encoder and a classifier connected to the encoder; the encoder is a pre-trained encoder obtained by self-supervised training using an unlabeled sample signal; based on the speech recognition loss calculated based on the labeled character sequence and the first predicted character sequence, and the semantic recognition loss calculated based on the labeled character sequence and the second predicted character sequence, the speech recognition model and the decoder are jointly trained, including: supervised training of the decoder and the classifier of the speech recognition model based on the speech recognition loss and the semantic recognition loss; when the supervised training stopping condition is met, supervised training of the decoder and the speech recognition model based on the speech recognition loss and the semantic recognition loss.
[0127] In one embodiment, a computer device performs self-supervised pre-training on an encoder in advance. After obtaining the pre-trained encoder, the network parameters of the encoder are first fixed, and the network parameters of the decoder and the classifier of the speech recognition model are updated according to the speech recognition loss and the semantic recognition loss. When the training stop condition is met, the network parameters of the decoder and the speech recognition model are updated according to the speech recognition loss and the semantic recognition loss.
[0128] In one embodiment, the decoder includes a vectorization layer, a semantic context network based on self-attention, and a speech semantic context network based on cross-attention; the speech features are decoded according to the semantic features corresponding to the forward character sequence to obtain the speech semantic joint features corresponding to the sample signal, including: through the vectorization layer of the decoder, the forward character sequence is converted into the corresponding forward character vector sequence, and the forward character vector sequence is input into the semantic context network; through the semantic context network, based on the forward character vector sequence, the contextual semantic features corresponding to the forward character sequence are calculated as the semantic features corresponding to the forward character sequence; through the speech semantic context network, based on the semantic features and speech features corresponding to the forward character sequence, the speech semantic joint features corresponding to the sample signal are calculated.
[0129] In one embodiment, the decoder may include a vectorization layer, a semantic context network based on self-attention, and a speech semantic context network based on cross-attention. The vectorization layer is used to convert the forward character sequence into a vector form, that is, a forward character vector sequence. The semantic context network based on self-attention is used to determine the attention of each forward character vector in the forward character vector sequence, that is, the importance of each forward character vector in the forward character vector sequence. The speech semantic context network based on cross-attention is used to determine the attention contribution of the previous character to the prediction of the next character, that is, how much attention the previous character needs to pay to predict the next character.
[0130] In one embodiment, the computer device inputs the forward character vector sequence into the self-attention-based semantic context network of the decoder, and through the semantic context network, respectively calculates the similarity between each forward character vector and the forward character vector sequence, normalizes each similarity, and obtains the self-attention of each forward character vector in the forward character vector sequence as the contextual semantic feature of the forward character vector sequence. Optionally, the computer device calculates the sum of each similarity, and respectively calculates the ratio of each similarity to the sum of similarities as the self-attention of each forward character vector in the forward character vector sequence.
[0131] In one embodiment, the computer device inputs the self-attention and the speech features extracted by the encoder into the speech semantic context network based on cross-attention of the decoder, and respectively calculates the similarity between the self-attention and the speech features corresponding to each forward character vector through the speech semantic context network, normalizes each similarity, obtains the cross-attention of the self-attention corresponding to each forward character vector in the speech features, and obtains the speech semantic joint features based on each cross-attention. In one embodiment, the computer device inputs the self-attention and the speech features extracted by the encoder into the speech semantic context network based on cross-attention of the decoder, and respectively calculates the similarity between the self-attention and the speech features corresponding to each forward character vector through the cross-attention network in the speech semantic context network, normalizes each similarity, and obtains the cross-attention of the self-attention corresponding to each forward character vector in the speech features. Optionally, the computer device calculates the sum of each similarity, and respectively calculates the ratio of each similarity to the sum of similarities as the cross-attention of the self-attention corresponding to each forward character vector in the speech features.
[0132] In one embodiment, the computer device encodes the cross attention corresponding to each forward character vector into the feedforward neural network in the speech semantic context network to obtain the speech semantic joint features output by the feedforward neural network. Figure 7 , Figure 7Schematic diagram of training a speech recognition model with the assistance of a decoder in an embodiment. The computer device inputs the forward character sequence " / x2x3x4x5" into the decoder 702. Through the vectorization layer of the decoder 702, the forward character sequence " / x2x3x4x5" is converted into the corresponding forward character vector sequence [e1 e2 e3 e4 e5]. The computer device inputs the forward character vector sequence [e1 e2 e3 e4 e5] into the self-attention-based semantic context network of the decoder 702. Through the semantic context network, the similarities s1, s2, s3, s4, s5 between each forward character vector e1, e2, e3, e4, e5 and the forward character vector sequence [e1 e2 e3 e4 e5] are calculated respectively. The similarities s1, s2, s3, s4, s5 are normalized to obtain the self-attention o1, o2, o3, o4, o5 of each forward character vector in the forward character vector sequence, which are used as the context semantic features of the forward character vector sequence [e1 e2 e3 e4 e5]. The computer device inputs the self-attention o1, o2, o3, o4, o5 and the speech features [c1 c2 c3 c4 c5] extracted by the encoder into the cross-attention-based speech semantic context network 704 of the decoder 702. Through the cross-attention network in the speech semantic context network 704, the similarities s1, s2, s3, s4, s5 between each self-attention and the speech features are calculated respectively. The similarities s1, s2, s3, s4, s5 are normalized to obtain the cross-attention u1, u2, u3, u4, u5 of each self-attention in the speech features [c1 c2 c3 c4 c5]. The computer device inputs the cross-attention u1, u2, u3, u4, u5 into the feed-forward neural network in the speech semantic context network 704 for encoding, and obtains the speech semantic joint features [r1 r2 r3 r4 r5] output by the feed-forward neural network.
[0133] Among them, the cross-attention u1, u2, u3, u4, u5 are used to represent the contribution of each forward character vector to predicting the next character corresponding to the forward character vector. For example, if the labeled character sequence is "The weather is good today", and the forward character sequence is " / Today's weather", then the cross-attention u2 corresponding to the forward character vector of the forward character "今" is used to represent the importance of the forward character "今" for predicting "日".
[0134] In one embodiment, based on the speech semantic joint features for prediction, to obtain the second predicted character sequence corresponding to the sample signal, it includes: inputting the speech semantic joint features into the classifier of the decoder; and outputting the second predicted character sequence corresponding to the sample signal through the classifier based on the speech semantic joint features.
[0135] In one embodiment, the decoder may include a vectorization layer, a semantic context network based on self-attention, a speech semantic context network based on cross-attention, and a classifier, wherein the classifier is used to identify the character corresponding to each time period signal in the sample signal based on the speech semantic joint feature, and output a second predicted character sequence corresponding to the sample signal. Figure 7 , the computer device inputs the speech-semantic joint feature [r1 r2 r3 r4 r5] into the classifier of the decoder 702, and the classifier outputs the second predicted character sequence "y1y2y3y4y5" corresponding to the sample signal based on the speech-semantic joint feature. The classifier of the speech recognition model outputs the first predicted character sequence "w1w2w3w4w5" corresponding to the sample signal based on the speech feature [c1 c2 c3 c4 c5]. Thus, the computer device can jointly train the speech recognition model and the decoder 702 based on the speech recognition loss calculated based on the labeled character sequence "x1x2x3x4x5" and the first predicted character sequence "w1w2w3w4w5", and the semantic recognition loss calculated based on the labeled character sequence "x1x2x3x4x5" and the second predicted character sequence "y1y2y3y4y5".
[0136] In this embodiment, the decoder includes a speech semantic context network based on cross-attention. The speech semantic context network based on cross-attention can utilize the speech features carrying the speech level output by the encoder and the context information corresponding to the forward character vector sequence input to the decoder to assist the speech recognition model in training, thereby distilling the context information at the semantic level into the speech recognition model, helping the speech recognition model to alleviate the independence assumption and the inability to utilize the context information at the semantic level, thereby improving the speech recognition accuracy.
[0137] In one embodiment, the method also includes: obtaining a signal to be recognized; inputting the signal to be recognized into a trained speech recognition model to obtain speech features output by an encoder in the speech recognition model, and a speech recognition result output by a classifier in the speech recognition model based on the speech features.
[0138] The signal to be recognized is a voice signal to be recognized by the method provided in the embodiment of the present application. The signal to be recognized can be a voice signal received in a voice interaction scenario, such as a virtual robot interaction scenario, a smart device control scenario, a machine translation scenario, a voice message text conversion scenario, etc.
[0139] In one embodiment, a computer device obtains a signal to be recognized, inputs the signal to be recognized into a trained speech recognition model, obtains speech features output by an encoder in the speech recognition model, and obtains a speech recognition result output by a classifier in the speech recognition model based on the speech features. The speech recognition result may be the phoneme or text corresponding to the signal to be recognized.
[0140] In this embodiment, since the trained speech recognition model can use context information at the semantic level for speech recognition, the speech recognition accuracy can be improved.
[0141] In one embodiment, referring to Figure 8 , provides a method for processing a speech recognition model, comprising the following steps:
[0142] Step S802, obtaining an unlabeled sample signal; inputting the unlabeled sample signal into an initial encoder to obtain a speech vector sequence corresponding to the unlabeled sample signal output by a feature extraction network in the initial encoder; performing a quantization operation on the speech vector sequence to obtain a speech quantization vector sequence; starting from the first speech vector in the speech vector sequence, masking the speech vectors in the speech vector sequence in sequence; inputting the masked speech vector sequence into a speech context network of the initial encoder in sequence to obtain contextual speech features output by the speech context network as speech features corresponding to the unlabeled sample signal; constructing a self-supervised training loss based on the difference between the speech quantization vector sequence and the speech features corresponding to the unlabeled sample signal; after updating the network parameters of the initial encoder according to the self-supervised training loss, returning to the step of obtaining the unlabeled sample signal to continue training until the training is completed, and obtaining a pre-trained encoder.
[0143] Step 804, obtain a sample signal and a corresponding labeled character sequence; input the sample signal into a pre-trained encoder in a speech recognition model to obtain a speech vector sequence corresponding to the sample signal output by a feature extraction network in the encoder; perform random masking on the speech vectors in the speech vector sequence; input the masked speech vector sequence into a speech context network to obtain contextual speech features output by the speech context network as speech features corresponding to the sample signal; and output a first predicted character sequence based on the speech features through a classifier connected to the encoder in the speech recognition model.
[0144] Step 806, input the forward character sequence corresponding to the marked character sequence into the decoder, the forward character sequence is generated based on the previous character of each character in the marked character sequence; in the decoder, the forward character sequence is converted into the corresponding forward character vector sequence through the vectorization layer of the decoder, and the forward character vector sequence is input into the semantic context network; through the semantic context network, based on the forward character vector sequence, the contextual semantic features corresponding to the forward character sequence are calculated as the semantic features corresponding to the forward character sequence; through the speech semantic context network, based on the semantic features and speech features corresponding to the forward character sequence, the speech semantic joint features corresponding to the sample signal are calculated; the speech semantic joint features are input into the classifier of the decoder; and the classifier outputs the second predicted character sequence corresponding to the sample signal based on the speech semantic joint features.
[0145] Step 808, constructing a speech recognition loss based on the difference between the labeled character sequence and the first predicted character sequence; constructing a semantic recognition loss based on the difference between the labeled character sequence and the second predicted character sequence; weighting the speech recognition loss and the semantic recognition loss according to a preset loss weighting coefficient to obtain a target loss; performing supervised training on the decoder and the classifier of the speech recognition model according to the target loss; when the supervised training stop condition is met, performing supervised training on the decoder and the speech recognition model according to the target loss.
[0146] For example, see Fig. 9 , Fig. 9The figure is a schematic diagram of a decoder-assisted training speech recognition model in one embodiment. The computer device divides the sample signal into speech signals in time periods t1 to t5, inputs the sample signal into an encoder, and obtains a speech vector sequence [z1 z2 z3 z4 z5] output by a feature extraction network in the encoder. The computer device performs random masking on the speech vectors in the speech vector sequence [z1 z2 z3 z4 z5] to obtain a masked speech vector sequence [*z2*z4*]. The computer device inputs the speech vector sequence [*z2*z4*] after masking into the speech context network based on self-attention, and respectively calculates the similarities s1, s2, s3, s4, s5 between each speech vector in the speech vector sequence [*z2*z4*] after masking and the speech vector sequence [*z2*z4*] after masking through the self-attention network in the speech context network based on self-attention, and normalizes the similarities s1, s2, s3, s4, s5 to obtain self-attention p1, p2, p3, p4, p5. The computer device inputs the self-attention p1, p2, p3, p4, p5 and each speech vector in the speech vector sequence [*z2*z4*] after masking into the feedforward neural network in the speech context network based on self-attention for encoding, and obtains the context speech features [c1 c2 c3 c4 c5] output by the feedforward neural network. The computer device inputs the forward character sequence " / x2x3x4x5" into the decoder, and converts the forward character sequence " / x2x3x4x5" into the corresponding forward character vector sequence [e1 e2 e3 e4 e5] through the vectorization layer of the decoder. The computer device inputs the forward character vector sequence [e1e2 e3 e4 e5] into the semantic context network based on self-attention of the decoder, and calculates the similarities s1, s2, s3, s4, s5 between each forward character vector e1, e2, e3, e4, e5 and the forward character vector sequence [e1 e2 e3 e4 e5] through the semantic context network, and normalizes the similarities s1, s2, s3, s4, s5 to obtain the self-attention o1, o2, o3, o4, o5 of each forward character vector in the forward character vector sequence as the contextual semantic feature of the forward character vector sequence [e1 e2 e3 e4 e5].The computer device inputs the self-attention o1, o2, o3, o4, o5 and the speech features [c1c2 c3 c4 c5] extracted by the encoder into the speech semantic context network based on cross attention of the decoder, and respectively calculates the similarities s1, s2, s3, s4, s5 between the respective attention and the speech features through the cross attention network in the speech semantic context network, and normalizes the similarities s1, s2, s3, s4, s5 to obtain the cross attention u1, u2, u3, u4, u5 of the respective attention o1, o2, o3, o4, o5 in the speech features [c1 c2 c3 c4 c5]. The computer device inputs the cross attention u1, u2, u3, u4, u5 into the feedforward neural network in the speech semantic context network for encoding, and obtains the speech semantic joint features [r1 r2 r3 r4 r5] output by the feedforward neural network.
[0147] The computer device inputs the speech-semantic joint feature [r1 r2 r3 r4 r5] into the classifier of the decoder, and the classifier outputs the second predicted character sequence "y1y2y3y4y5" corresponding to the sample signal based on the speech-semantic joint feature. The classifier of the speech recognition model outputs the first predicted character sequence "w1w2w3w4w5" corresponding to the sample signal based on the speech feature [c1 c2 c3 c4 c5]. Thus, the computer device can jointly train the speech recognition model and the decoder based on the speech recognition loss calculated based on the labeled character sequence "x1x2x3x4x5" and the first predicted character sequence "w1w2w3w4w5", and the semantic recognition loss calculated based on the labeled character sequence "x1x2x3x4x5" and the second predicted character sequence "y1y2y3y4y5".
[0148] In one embodiment, the speech context network based on self-attention may have M layers, and the structure of each layer includes: Multi-head Self Attention, Add, Norm, Feed Forward, Add, Norm, and M may be 12.
[0149] In one embodiment, the decoder may specifically include an Embedding Layer (vectorization layer), an N-layer intermediate coding layer connected to the vectorization layer, wherein the N-layer intermediate coding layer may include a semantic context network based on self-attention and a speech semantic context network based on cross-attention in sequence. The decoder may also include a classifier connected to the N-layer intermediate coding layer. Among them, the specific structure of the semantic context network based on self-attention includes Multi-headSelf_Attention (multi-head self-attention), Add (sum operation), and Norm (normalization operation) in sequence. The specific structure of the speech semantic context network based on cross-attention includes Multi-head Cross_Attention (multi-head cross attention), Add (sum operation), Norm (normalization operation), Feed Forward (feedforward neural network), Add (sum operation), and Norm (normalization operation) in sequence. N can take a value of 6.
[0150] The processing method of the above-mentioned speech recognition model inputs a sample signal into the speech recognition model to obtain speech features corresponding to the sample signal and a first predicted character sequence output based on the speech features, and inputs a forward character sequence corresponding to the marked character sequence into a decoder. In the decoder, the speech features are decoded according to the semantic features corresponding to the forward character sequence to obtain speech and semantic joint features corresponding to the sample signal. Since the forward character sequence is generated based on the previous character of each character in the marked character sequence, the speech features output by the encoder are decoded and encoded according to the semantic features corresponding to the forward character sequence. The speech and semantic joint features obtained carry context information at the semantic level. A second predicted character sequence corresponding to the sample signal is predicted based on the speech and semantic joint features. The semantic recognition loss constructed based on the second predicted character sequence and the marked character sequence is used to assist in training the speech recognition model, which can distill the context information at the semantic level into the speech recognition model, thereby improving the recognition accuracy of the speech recognition model.
[0151] In order to verify the effect of the solution provided in the embodiment of the present application, a comparative experiment was conducted. This test adopted two training methods for the speech recognition model, one is to jointly train the speech recognition model and the decoder (hereinafter referred to as the joint training method), and the other is to train the speech recognition model alone (hereinafter referred to as the single training method). The specific implementation methods of the two training methods are now introduced.
[0152] For the joint training method, the computer device first performs self-supervised pre-training on the encoder of the speech recognition model. The pre-training step of the encoder refers to the above step S802 and will not be repeated here. After obtaining the pre-trained encoder, the computer device obtains the sample signal and the corresponding labeled character sequence, inputs the sample signal into the pre-trained encoder in the speech recognition model, and obtains the speech vector sequence corresponding to the sample signal output by the feature extraction network in the encoder; performs random masking on the speech vectors in the speech vector sequence, inputs the masked speech vector sequence into the speech context network, and obtains the contextual speech features output by the speech context network as the speech features corresponding to the sample signal; outputs the first predicted character sequence based on the speech features through the classifier connected to the encoder in the speech recognition model. The computer device inputs the forward character sequence corresponding to the marked character sequence into the decoder, and the forward character sequence is generated based on the previous character of each character in the marked character sequence; in the decoder, the forward character sequence is converted into the corresponding forward character vector sequence through the vectorization layer of the decoder, and the forward character vector sequence is input into the semantic context network, and the contextual semantic features corresponding to the forward character sequence are calculated based on the forward character vector sequence through the semantic context network as the semantic features corresponding to the forward character sequence; the speech semantic context network is used to calculate the speech semantic joint features corresponding to the sample signal based on the semantic features and speech features corresponding to the forward character sequence; The speech and semantic joint features are input into the classifier of the decoder; the classifier outputs a second predicted character sequence corresponding to the sample signal based on the speech and semantic joint features; the computer device constructs a speech recognition loss based on the difference between the labeled character sequence and the first predicted character sequence, and constructs a semantic recognition loss based on the difference between the labeled character sequence and the second predicted character sequence, and weights and sums the speech recognition loss and the semantic recognition loss according to a preset loss weighting coefficient to obtain a target loss; supervised training is performed on the classifier of the decoder and the speech recognition model according to the target loss, and when the supervised training stop condition is met, supervised training is performed on the decoder and the speech recognition model according to the target loss.
[0153] For the separate training method, the computer device first performs self-supervised pre-training on the encoder of the speech recognition model. The pre-training step of the encoder refers to the above step S802 and is not repeated here. After obtaining the pre-trained encoder, the computer device obtains the sample signal and the corresponding annotated character sequence, inputs the sample signal into the pre-trained encoder in the speech recognition model, and obtains the speech vector sequence corresponding to the sample signal output by the feature extraction network in the encoder; performs random masking on the speech vectors in the speech vector sequence, inputs the masked speech vector sequence into the speech context network, and obtains the context speech features output by the speech context network as the speech features corresponding to the sample signal; outputs the first predicted character sequence based on the speech features through the classifier connected to the encoder in the speech recognition model; constructs the speech recognition loss based on the difference between the annotated character sequence and the first predicted character sequence; performs supervised training on the classifier of the speech recognition model according to the speech recognition loss, and when the supervised training stop condition is met, performs supervised training on the encoder and classifier of the speech recognition model according to the speech recognition loss.
[0154] The self-supervised training data used in the two training methods is 960 hours of librispeech data, and the supervised training data used is the open source Chinese speech recognition dataset Aishell-1. The Aishell-1 dataset includes training set, validation set and test set. The number of test items in the Aishell-1 training set is 120098, the number of test items in the Aishell-1 validation set is 14326, and the number of test items in the Aishell-1 test set is 7176. The feature dimensions of the decoder and encoder are both 768. For the joint training method, the loss weighting coefficient of speech recognition loss is 0.3, and the loss weighting coefficient of semantic recognition loss is 0.7. The M layer of the self-attention-based speech context network is 12, and the N layer of the decoder is 6.
[0155] The speech recognition models trained by the joint training method and the separate training method were tested, and the test results are as follows: Fig.10 As shown in the figure, it can be seen that compared with the speech recognition model trained by the single training method, the speech recognition model trained by the joint training method has a significantly lower word error rate, which means that the joint training method can significantly improve the model performance.
[0156] In one embodiment, Fig.11 As shown, a speech recognition method is provided. This embodiment mainly applies this method to the above Figure 1 Taking the computer device (terminal 102 or server 104) in the example as an example, the following steps are included:
[0157] Step S1102, obtaining a signal to be identified.
[0158] The signal to be recognized is a voice signal to be recognized by the method provided in the embodiment of the present application. The signal to be recognized can be a voice signal received in a voice interaction scenario, such as a virtual robot interaction scenario, a smart device control scenario, a machine translation scenario, a voice message text conversion scenario, etc.
[0159] Step S1104, input the signal to be recognized into the trained speech recognition model to obtain the speech features output by the encoder in the speech recognition model, and the speech recognition results output by the classifier in the speech recognition model based on the speech features; wherein the speech recognition model and the decoder are jointly trained based on the speech recognition loss and the semantic recognition loss, the speech recognition loss is calculated based on the first predicted character sequence and the annotated character sequence corresponding to the sample signal, the semantic recognition loss is calculated based on the second predicted character sequence and the annotated character sequence, the first predicted character sequence is obtained by classification based on the speech features output by the encoder, the second predicted character sequence is obtained by predicting the speech and semantic joint features obtained by decoding the speech features using the semantic features corresponding to the forward character sequence corresponding to the annotated character sequence by the decoder, and the forward character sequence is generated based on the previous character of each character in the annotated character sequence.
[0160] In one embodiment, a computer device obtains a signal to be recognized, inputs the signal to be recognized into a trained speech recognition model, obtains speech features output by an encoder in the speech recognition model, and obtains a speech recognition result output by a classifier in the speech recognition model based on the speech features. The speech recognition result may be the phoneme or text corresponding to the signal to be recognized.
[0161] Regarding the training method of the speech recognition model, please refer to the above embodiment and will not be repeated here.
[0162] In the above-mentioned speech recognition method, the signal to be recognized is input into a trained speech recognition model to obtain speech features output by the encoder in the speech recognition model, and speech recognition results output by the classifier in the speech recognition model based on the speech features. Since the trained speech recognition model can utilize contextual information at the semantic level for speech recognition, the speech recognition accuracy can be improved.
[0163] It should be understood that although Figure 3 , 8 The steps in the flowchart of 11 are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 3 , 8, at least part of the steps in 11 may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.
[0164] In one embodiment, Fig.12 As shown, a processing device for a speech recognition model is provided. The device can adopt a software module or a hardware module, or a combination of the two to become a part of a computer device. The device specifically includes: an acquisition module 1202, an encoding module 1204, an input module 1206, a decoding module 1208 and a training module 1210, wherein:
[0165] An acquisition module 1202 is used to acquire a sample signal and a corresponding marked character sequence;
[0166] The encoding module 1204 is used to input the sample signal into the speech recognition model to obtain the speech features corresponding to the sample signal and the first predicted character sequence output based on the speech features;
[0167] An input module 1206, for inputting a forward character sequence corresponding to the marked character sequence into a decoder, wherein the forward character sequence is generated based on the previous character of each character in the marked character sequence;
[0168] A decoding module 1208 is used to decode the speech feature in the decoder according to the semantic feature corresponding to the forward character sequence, obtain the speech and semantic joint feature corresponding to the sample signal, and make a prediction based on the speech and semantic joint feature to obtain a second predicted character sequence corresponding to the sample signal;
[0169] The training module 1210 is used to jointly train the speech recognition model and the decoder based on the speech recognition loss calculated according to the marked character sequence and the first predicted character sequence, and the semantic recognition loss calculated according to the marked character sequence and the second predicted character sequence.
[0170] In one embodiment, the encoding module 1204 is also used to: input the sample signal into the speech recognition model; output the speech features corresponding to the sample signal through the encoder of the speech recognition model; and output the first predicted character sequence based on the speech features through a classifier connected to the encoder in the speech recognition model.
[0171] In one embodiment, the encoder includes a feature extraction network and a self-attention-based speech context network; the encoding module 1204 is also used to: input the sample signal into the encoder to obtain a speech vector sequence corresponding to the sample signal output by the feature extraction network in the encoder; perform random masking on the speech vectors in the speech vector sequence; input the masked speech vector sequence into the speech context network to obtain contextual speech features output by the speech context network as speech features corresponding to the sample signal.
[0172] In one embodiment, the decoder includes a vectorization layer, a semantic context network based on self-attention, and a speech semantic context network based on cross-attention; the decoding module 1208 is also used to: convert the forward character sequence into a corresponding forward character vector sequence through the vectorization layer of the decoder, and input the forward character vector sequence into the semantic context network; calculate the contextual semantic features corresponding to the forward character sequence based on the forward character vector sequence through the semantic context network as the semantic features corresponding to the forward character sequence; calculate the speech semantic joint features corresponding to the sample signal based on the semantic features and speech features corresponding to the forward character sequence through the speech semantic context network.
[0173] In one embodiment, the decoding module 1208 is further used to: input the speech-semantics joint feature into the classifier of the decoder; and output a second predicted character sequence corresponding to the sample signal based on the speech-semantics joint feature through the classifier.
[0174] In one embodiment, the speech recognition model includes an encoder and a classifier connected to the encoder; the encoder is a pre-trained encoder obtained by self-supervised training using unlabeled sample signals; the training module 1210 is also used to: perform supervised training on the decoder and the classifier of the speech recognition model according to the speech recognition loss and the semantic recognition loss; when the supervised training stopping condition is met, the decoder and the speech recognition model are supervised trained according to the speech recognition loss and the semantic recognition loss.
[0175] In one embodiment, the encoder is a pre-trained encoder obtained by self-supervised training using an unlabeled sample signal; the speech recognition model also includes a pre-training module 1210, which is used to: obtain an unlabeled sample signal; input the unlabeled sample signal into the initial encoder to obtain a speech vector sequence corresponding to the unlabeled sample signal output by the feature extraction network in the initial encoder; perform a quantization operation on the speech vector sequence to obtain a speech quantization vector sequence; after randomly masking the speech vectors in the speech vector sequence, determine the masked speech vector; input the masked speech vector sequence into the speech context network of the initial encoder to obtain a predicted speech vector corresponding to the masked speech vector output by the speech context network; construct a self-supervised training loss based on the difference between the speech quantization vector corresponding to the masked speech vector in the speech quantization vector sequence and the predicted speech vector; after updating the network parameters of the initial encoder according to the self-supervised training loss, return to the step of obtaining the unlabeled sample signal to continue training until the training is completed, and obtain the pre-trained encoder.
[0176] In one embodiment, the training module 1210 is also used to: construct a speech recognition loss based on the difference between the labeled character sequence and the first predicted character sequence; construct a semantic recognition loss based on the difference between the labeled character sequence and the second predicted character sequence; weightedly sum the speech recognition loss and the semantic recognition loss according to a preset loss weighting coefficient to obtain a target loss; and jointly train the speech recognition model and the decoder according to the target loss.
[0177] In one embodiment, the processing device of the speech recognition model also includes a speech recognition module, which is used to: obtain a signal to be recognized; input the signal to be recognized into a trained speech recognition model to obtain speech features output by an encoder in the speech recognition model, and a speech recognition result output by a classifier in the speech recognition model based on the speech features.
[0178] For the specific definition of the processing device of the speech recognition model, please refer to the definition of the processing method of the speech recognition model in the above text, which will not be repeated here. The various modules in the processing device of the above speech recognition model can be implemented in whole or in part by software, hardware and a combination thereof. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0179] In the processing device of the above-mentioned speech recognition model, a sample signal is input into the speech recognition model to obtain speech features corresponding to the sample signal and a first predicted character sequence output based on the speech features, and a forward character sequence corresponding to the marked character sequence is input into the decoder. In the decoder, the speech features are decoded according to the semantic features corresponding to the forward character sequence to obtain the speech and semantic joint features corresponding to the sample signal. Since the forward character sequence is generated based on the previous character of each character in the marked character sequence, the speech features output by the encoder are decoded and encoded according to the semantic features corresponding to the forward character sequence. The speech and semantic joint features obtained carry context information at the semantic level. A second predicted character sequence corresponding to the sample signal is predicted based on the speech and semantic joint features. The semantic recognition loss constructed based on the second predicted character sequence and the marked character sequence is used to assist the speech recognition model in training, which can distill the context information at the semantic level into the speech recognition model, thereby improving the recognition accuracy of the speech recognition model.
[0180] In one embodiment, Fig.13 As shown, a speech recognition device is provided. The device can adopt a software module or a hardware module, or a combination of the two to become a part of a computer device. The device specifically includes: an acquisition module 1302 and a speech recognition module 1304, wherein:
[0181] An acquisition module 1302 is used to acquire a signal to be identified;
[0182] The speech recognition module 1304 is used to input the signal to be recognized into the trained speech recognition model, obtain the speech features output by the encoder in the speech recognition model, and the speech recognition results output by the classifier in the speech recognition model based on the speech features;
[0183] Among them, the speech recognition model and the decoder are jointly trained based on the speech recognition loss and the semantic recognition loss. The speech recognition loss is calculated based on the first predicted character sequence and the labeled character sequence corresponding to the sample signal. The semantic recognition loss is calculated based on the second predicted character sequence and the labeled character sequence. The first predicted character sequence is obtained after classification based on the speech features output by the encoder. The second predicted character sequence is obtained by predicting the joint features of speech and semantics obtained by decoding the speech features using the semantic features corresponding to the forward character sequence corresponding to the labeled character sequence by the decoder. The forward character sequence is generated based on the previous character of each character in the labeled character sequence.
[0184] For the specific definition of the speech recognition device, please refer to the definition of the speech recognition method above, which will not be repeated here. Each module in the above-mentioned speech recognition device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0185] In the above-mentioned speech recognition device, the signal to be recognized is input into a trained speech recognition model to obtain speech features output by the encoder in the speech recognition model, and speech recognition results output by the classifier in the speech recognition model based on the speech features. Since the trained speech recognition model can utilize contextual information at the semantic level for speech recognition, the accuracy of speech recognition can be improved.
[0186] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Fig.14 As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store processing data and / or image generation data of a speech recognition model. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a processing method for a speech recognition model and / or a speech recognition method is implemented.
[0187] In one embodiment, a computer device is provided. The computer device may be a terminal or a face acquisition device. The internal structure diagram thereof may be as follows: Fig.15 As shown. The computer device includes a processor, a memory, a communication interface and a voice collection device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, an operator network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a method for processing a speech recognition model and / or a speech recognition method is implemented.
[0188] Those skilled in the art will understand that Fig.14 and Fig.15 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0189] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.
[0190] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0191] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps in the above-mentioned method embodiments.
[0192] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0193] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0194] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for processing a speech recognition model, characterized in that: The method comprises: Obtaining sample signals and corresponding labeled character sequences; Inputting the sample signal into a speech recognition model to obtain speech features corresponding to the sample signal and a first predicted character sequence output based on the speech features; Inputting a forward character sequence corresponding to the marked character sequence into a decoder, wherein the forward character sequence is generated based on the previous character of each character in the marked character sequence; In the decoder, the speech feature is decoded according to the semantic feature corresponding to the forward character sequence to obtain the speech and semantic joint feature corresponding to the sample signal, and prediction is performed based on the speech and semantic joint feature to obtain a second predicted character sequence corresponding to the sample signal; The speech recognition model and the decoder are jointly trained based on the speech recognition loss calculated based on the labeled character sequence and the first predicted character sequence, and the semantic recognition loss calculated based on the labeled character sequence and the second predicted character sequence.
2. The method according to claim 1, characterized in that The step of inputting the sample signal into a speech recognition model to obtain speech features corresponding to the sample signal and outputting a first predicted character sequence based on the speech features includes: Inputting the sample signal into the speech recognition model; Outputting speech features corresponding to the sample signal through an encoder of the speech recognition model; The first predicted character sequence is output based on the speech features through a classifier connected to the encoder in the speech recognition model.
3. The method according to claim 2, characterized in that The encoder includes a feature extraction network and a self-attention-based speech context network; The outputting of the speech features corresponding to the sample signal by the encoder of the speech recognition model includes: Inputting the sample signal into the encoder to obtain a speech vector sequence corresponding to the sample signal output by a feature extraction network in the encoder; Performing random masking processing on the speech vectors in the speech vector sequence; The masked speech vector sequence is input into the speech context network to obtain context speech features output by the speech context network as speech features corresponding to the sample signal.
4. The method according to claim 1, characterized in that: The decoder includes a vectorization layer, a self-attention-based semantic context network, and a cross-attention-based speech semantic context network; The decoding of the speech feature according to the semantic feature corresponding to the forward character sequence to obtain the speech and semantic joint feature corresponding to the sample signal includes: The forward character sequence is converted into a corresponding forward character vector sequence through the vectorization layer of the decoder, and the forward character vector sequence is input into the semantic context network; By means of the semantic context network, based on the forward character vector sequence, a contextual semantic feature corresponding to the forward character sequence is calculated as a semantic feature corresponding to the forward character sequence; The speech semantic context network is used to calculate the speech semantic joint feature corresponding to the sample signal based on the semantic feature corresponding to the forward character sequence and the speech feature.
5. The method according to claim 4, characterized in that The step of performing prediction based on the speech-semantic joint feature to obtain a second predicted character sequence corresponding to the sample signal includes: Inputting the speech-semantic joint feature into the classifier of the decoder; The classifier outputs a second predicted character sequence corresponding to the sample signal based on the speech-semantic joint feature.
6. The method according to claim 1, characterized in that The speech recognition model includes an encoder and a classifier connected to the encoder; the encoder is a pre-trained encoder obtained by self-supervised training using unlabeled sample signals; The method of jointly training the speech recognition model and the decoder based on the speech recognition loss calculated according to the labeled character sequence and the first predicted character sequence, and the semantic recognition loss calculated according to the labeled character sequence and the second predicted character sequence, comprises: Performing supervised training on the decoder and the classifier of the speech recognition model according to the speech recognition loss and the semantic recognition loss; When the supervised training stop condition is met, supervised training is performed on the decoder and the speech recognition model according to the speech recognition loss and the semantic recognition loss.
7. The method according to claim 2, 3 or 6, characterized in that: The encoder is a pre-trained encoder obtained by self-supervised training using unlabeled sample signals; The method further comprises: Acquire the unlabeled sample signal; Inputting the unlabeled sample signal into an initial encoder to obtain a speech vector sequence corresponding to the unlabeled sample signal output by a feature extraction network in the initial encoder; Performing a quantization operation on the speech vector sequence to obtain a speech quantization vector sequence; After performing random masking processing on the speech vectors in the speech vector sequence, a masked speech vector is determined; Inputting the masked speech vector sequence into the speech context network of the initial encoder to obtain a predicted speech vector corresponding to the masked speech vector output by the speech context network; constructing a self-supervised training loss based on a difference between a speech quantization vector in the speech quantization vector sequence corresponding to the masked speech vector and the predicted speech vector; After updating the network parameters of the initial encoder according to the self-supervised training loss, returning to the step of obtaining the unlabeled sample signal to continue training until the training is completed, thereby obtaining the pre-trained encoder.
8. The method according to claim 1, characterized in that: The method of jointly training the speech recognition model and the decoder based on the speech recognition loss calculated according to the labeled character sequence and the first predicted character sequence, and the semantic recognition loss calculated according to the labeled character sequence and the second predicted character sequence, comprises: constructing the speech recognition loss based on the difference between the labeled character sequence and the first predicted character sequence; constructing a semantic recognition loss based on the difference between the labeled character sequence and the second predicted character sequence; The speech recognition loss and the semantic recognition loss are weighted and summed according to a preset loss weighting coefficient to obtain a target loss; The speech recognition model and the decoder are jointly trained according to the target loss.
9. The method according to claim 1, characterized in that: The method further comprises: Acquire a signal to be identified; The signal to be recognized is input into a trained speech recognition model to obtain speech features output by an encoder in the speech recognition model, and speech recognition results output by a classifier in the speech recognition model based on the speech features.
10. A speech recognition method, characterized in that: The method comprises: Acquire a signal to be identified; Inputting the signal to be recognized into a trained speech recognition model to obtain speech features output by an encoder in the speech recognition model, and speech recognition results output by a classifier in the speech recognition model based on the speech features; Among them, the speech recognition model and decoder are obtained by joint training based on speech recognition loss and semantic recognition loss, the speech recognition loss is calculated based on a first predicted character sequence and a labeled character sequence corresponding to a sample signal, and the semantic recognition loss is calculated based on a second predicted character sequence and the labeled character sequence. The first predicted character sequence is obtained after classification based on the speech features output by the encoder, and the second predicted character sequence is obtained by predicting the speech and semantic joint features obtained by decoding the speech features using the semantic features corresponding to the forward character sequence corresponding to the labeled character sequence by the decoder, and the forward character sequence is generated based on the previous character of each character in the labeled character sequence.
11. A processing device for a speech recognition model, characterized in that: The device comprises: An acquisition module, used to acquire a sample signal and a corresponding labeled character sequence; An encoding module, used for inputting the sample signal into a speech recognition model to obtain speech features corresponding to the sample signal and a first predicted character sequence output based on the speech features; An input module, used for inputting a forward character sequence corresponding to the marked character sequence into a decoder, wherein the forward character sequence is generated based on the previous character of each character in the marked character sequence; A decoding module, configured to decode the speech feature in the decoder according to the semantic feature corresponding to the forward character sequence, obtain the speech and semantic joint feature corresponding to the sample signal, and make a prediction based on the speech and semantic joint feature to obtain a second predicted character sequence corresponding to the sample signal; A training module is used to jointly train the speech recognition model and the decoder based on the speech recognition loss calculated according to the labeled character sequence and the first predicted character sequence, and the semantic recognition loss calculated according to the labeled character sequence and the second predicted character sequence.
12. The device according to claim 11, characterized in that The encoding module is also used to: input the sample signal into the speech recognition model; output the speech features corresponding to the sample signal through the encoder of the speech recognition model; and output the first predicted character sequence based on the speech features through a classifier connected to the encoder in the speech recognition model.
13. The device according to claim 12, characterized in that The encoder includes a feature extraction network and a self-attention-based speech context network; the encoding module is also used to: input the sample signal into the encoder to obtain a speech vector sequence corresponding to the sample signal output by the feature extraction network in the encoder; perform random masking on the speech vectors in the speech vector sequence; input the masked speech vector sequence into the speech context network to obtain contextual speech features output by the speech context network as speech features corresponding to the sample signal.
14. The device according to claim 11, characterized in that The decoder includes a vectorization layer, a semantic context network based on self-attention, and a speech semantic context network based on cross-attention; the decoding module is also used to: convert the forward character sequence into a corresponding forward character vector sequence through the vectorization layer of the decoder, and input the forward character vector sequence into the semantic context network; calculate the contextual semantic features corresponding to the forward character sequence based on the forward character vector sequence through the semantic context network as the semantic features corresponding to the forward character sequence; calculate the speech semantic joint features corresponding to the sample signal based on the semantic features corresponding to the forward character sequence and the speech features through the speech semantic context network.
15. The device according to claim 14, characterized in that The decoding module is also used to: input the speech-semantics joint feature into the classifier of the decoder; and output a second predicted character sequence corresponding to the sample signal based on the speech-semantics joint feature through the classifier.
16. The device according to claim 11, characterized in that The speech recognition model includes an encoder and a classifier connected to the encoder; the encoder is a pre-trained encoder obtained by self-supervised training using unlabeled sample signals; the training module is also used to: perform supervised training on the decoder and the classifier of the speech recognition model according to the speech recognition loss and the semantic recognition loss; when the supervised training stopping condition is met, perform supervised training on the decoder and the speech recognition model according to the speech recognition loss and the semantic recognition loss.
17. The device according to claim 12, 13 or 16, characterized in that The encoder is a pre-trained encoder obtained by self-supervised training using an unlabeled sample signal; the speech recognition model also includes a pre-training module, which is used to: obtain the unlabeled sample signal; input the unlabeled sample signal into an initial encoder to obtain a speech vector sequence corresponding to the unlabeled sample signal output by a feature extraction network in the initial encoder; perform a quantization operation on the speech vector sequence to obtain a speech quantization vector sequence; perform a random masking process on the speech vectors in the speech vector sequence to determine the masked speech vector; input the masked speech vector sequence into a speech context network of the initial encoder to obtain a predicted speech vector corresponding to the masked speech vector output by the speech context network; construct a self-supervised training loss based on the difference between the speech quantization vector corresponding to the masked speech vector in the speech quantization vector sequence and the predicted speech vector; after updating the network parameters of the initial encoder according to the self-supervised training loss, return to the step of obtaining the unlabeled sample signal to continue training until the training is completed, and obtain the pre-trained encoder.
18. The device according to claim 11, characterized in that The training module is also used to: construct the speech recognition loss based on the difference between the labeled character sequence and the first predicted character sequence; construct the semantic recognition loss based on the difference between the labeled character sequence and the second predicted character sequence; weightedly sum the speech recognition loss and the semantic recognition loss according to a preset loss weighting coefficient to obtain a target loss; and jointly train the speech recognition model and the decoder according to the target loss.
19. The device according to claim 11, characterized in that The processing device of the speech recognition model also includes a speech recognition module, which is used to: obtain a signal to be recognized; input the signal to be recognized into a trained speech recognition model to obtain speech features output by an encoder in the speech recognition model, and a speech recognition result output by a classifier in the speech recognition model based on the speech features.
20. A speech recognition device, characterized in that: The device comprises: An acquisition module, used for acquiring a signal to be identified; A speech recognition module, used for inputting the signal to be recognized into a trained speech recognition model, obtaining speech features output by an encoder in the speech recognition model, and a speech recognition result output by a classifier in the speech recognition model based on the speech features; Among them, the speech recognition model and decoder are obtained by joint training based on speech recognition loss and semantic recognition loss, the speech recognition loss is calculated based on a first predicted character sequence and a labeled character sequence corresponding to a sample signal, and the semantic recognition loss is calculated based on a second predicted character sequence and the labeled character sequence. The first predicted character sequence is obtained after classification based on the speech features output by the encoder, and the second predicted character sequence is obtained by predicting the speech and semantic joint features obtained by decoding the speech features using the semantic features corresponding to the forward character sequence corresponding to the labeled character sequence by the decoder, and the forward character sequence is generated based on the previous character of each character in the labeled character sequence.
21. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 are implemented.
22. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
23. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Voice processing method and electronic device
CN106992012A
Voice interaction method, device and equipment
CN107464564A