Speech recognition and model training method, device, equipment and computer program product
By jointly training the speech recognition model and text reconstruction model, using the semantic information of the large language model, the problem of poor recognition effect of end-to-end speech recognition system in specific fields is solved, and better domain customization capabilities and recognition effects are achieved.
Patent Information
- Application Number
- CN202510625610.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-09-25
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-15
AI Technical Summary
When facing a specific field, the recognition effect of end-to-end speech recognition system has significantly decreased, mainly due to the data long-tail effect and the difficulty in obtaining speech and text data in specific fields.
The joint training strategy is adopted to jointly train the speech recognition model and the text reconstruction model. The two share the same decoder. The text reconstruction model includes a text encoder built on a large language model. By calculating the feature alignment loss value between audio semantic representation and text semantic representation, unify the semantic representation space, and use the rich semantic information of large language models to improve the feature extraction capability of audio encoder.
It improves the recognition effect of speech recognition models in specific fields, enhances the domain customization capabilities, simplifies model structure, reduces inference time, and eliminates the need for external language models or additional TTS systems.
Smart Images

Figure CN120126459A_ABST
Abstract
Description
[0001] This application claims the priority of a domestic application filed with the China National Patent Office on September 25, 2024, with the application number 202411338067.6 and the invention title "Speech Recognition and Model Training Method, Device, Equipment and Computer Program Product", the entire content of which is incorporated herein by reference. Technical Field
[0002] This application relates to the field of speech recognition technology, and more specifically, to a speech recognition and model training method, device, equipment and computer program product. Background Art
[0003] In recent years, with the development of intelligent speech technology, end-to-end speech recognition systems have become a research hotspot and have been widely used in usage scenarios such as voice assistants, meeting transcription, and call transcription.
[0004] Compared with traditional speech recognition systems, end-to-end speech recognition systems implicitly integrate modules such as acoustic models, speech models, and pronunciation dictionaries, directly learning the mapping relationship from speech signals to text, simplifying the training and inference processes, and avoiding error accumulation in cascaded systems. With the rise of deep learning technology, especially the application of deep neural networks (DNNs), the accuracy of end-to-end speech recognition systems has been significantly improved compared with traditional speech recognition systems. However, end-to-end speech recognition systems still face many challenges when dealing with specific (professional) fields, such as medical, legal, financial and other professional fields.
[0005] Since end-to-end speech recognition models rely on a large amount of training data to achieve good results, for some specific fields, on the one hand, due to the long-tail effect of data, the data in specific fields usually accounts for a very low proportion in the training data, and on the other hand, it is particularly difficult to obtain a sufficient amount of speech texts in specific fields. Therefore, when facing specific fields, the recognition effect of end-to-end speech recognition models will drop significantly. Summary of the Invention
[0006] In view of the above problems, this application is proposed to provide a speech recognition and model training method, device, equipment and computer program product to improve the recognition effect of speech recognition models in specific fields. The specific solutions are as follows:
[0007] In the first aspect of this application, a speech recognition method is provided, including:
[0008] Obtain the acoustic features of the speech signal to be recognized;
[0009] Input the acoustic features into the trained speech recognition model to obtain the speech recognition result output by the model; where:
[0010] The speech recognition model is jointly trained with the text reconstruction model during the training phase. The speech recognition model and the text reconstruction model share the same decoder. The text reconstruction model further includes a text encoder constructed based on a large language model. In the joint training process, an audio sample is used as the input to the audio encoder in the speech recognition model, and the recognition text label corresponding to the audio sample is used as the input to the text encoder. The total loss value of the joint training includes a text recognition loss value calculated based on the text output by the decoder and the recognition text label, and a feature alignment loss value calculated based on the audio semantic representation extracted by the audio encoder and the text semantic representation extracted by the text encoder.
[0011] In a second aspect of the present application, a method for training a speech recognition model is provided, including:
[0012] Obtain audio-text pair training data, where the audio-text pair training data includes audio samples and corresponding recognition text labels;
[0013] Use the audio sample as the input to the audio encoder in the speech recognition model, and use the recognition text label as the input to the text encoder constructed based on a large language model in the text reconstruction model. The text reconstruction model and the speech recognition model share the same decoder. The audio encoder is used to extract an audio semantic representation and send it to the decoder with a first probability. The text encoder is used to extract a text semantic representation and send it to the decoder with a second probability. The decoder is used to output the decoded text, and the sum of the first probability and the second probability is equal to 1;
[0014] Calculate a text recognition loss value based on the text output by the decoder and the recognition text label, calculate a feature alignment loss value based on the audio semantic representation and the text semantic representation, and calculate the total loss value of the text recognition loss value and the feature alignment loss value;
[0015] Update the model parameters according to the total loss value until the training end condition is reached, and obtain the trained speech recognition model.
[0016] In a possible design, in another implementation manner of the second aspect of the embodiments of the present application, the speech recognition model further includes a shared encoder located between the audio encoder and the decoder. The shared encoder encodes the input features and sends the encoded features to the decoder. The audio semantic representation extracted by the audio encoder is sent to the shared encoder with a first probability, and the text semantic representation extracted by the text encoder is sent to the shared encoder with a second probability.
[0017] In a possible design, in another implementation manner of the second aspect of the embodiments of the present application, the speech recognition model further includes an alignment module;
[0018] When the input of the shared encoder is the audio semantic representation, the alignment module is used to predict audio-text alignment information based on the audio semantic representation;
[0019] The total loss value further includes: a connectionist temporal classification (CTC) loss calculated based on the audio-text alignment information and the recognition text label.
[0020] In a possible design, in another implementation manner of the second aspect of the embodiments of the present application, the text encoder includes a large language model, a resampling module, and a duration prediction module;
[0021] The large language model is used to extract an initial text semantic representation from the input text;
[0022] When the input of the shared encoder is the audio semantic representation, the resampling module is used to resample the initial text semantic representation based on the audio-text alignment information output by the alignment module to obtain a resampled text semantic representation with the same number of frames as the audio semantic representation, and calculate the feature alignment loss value using the resampled text semantic representation and the audio semantic representation;
[0023] The duration prediction module is used to predict the duration of each frame in the initial text semantic representation to obtain duration prediction information;
[0024] The total loss value further includes: a duration prediction loss value calculated based on the duration prediction information and the audio-text alignment information.
[0025] In a possible design, in another implementation manner of the second aspect of the embodiments of the present application, after updating the model parameters according to the total loss value until the training end condition is reached, it further includes:
[0026] Obtain a text corpus in the target domain;
[0027] Send the text corpus into the large language model to obtain the target text semantic representation of the text corpus;
[0028] Predict the duration of each frame in the target text semantic representation through the duration prediction module to obtain target duration prediction information;
[0029] Through the resampling module, resample the target text semantic representation according to the target duration prediction information to obtain a resampled text semantic representation;
[0030] Encode the resampled text semantic representation through the shared encoder, and send the encoded features to the decoder to predict the decoded text through the decoder;
[0031] Calculate the text reconstruction loss value based on the decoded text predicted by the decoder and the text corpus, and update the model parameters according to the text reconstruction loss value until the training end condition is reached. The final speech recognition model is composed of the audio encoder, the shared encoder, and the decoder.
[0032] In a possible design, in another implementation manner of the second aspect of the embodiments of the present application, after obtaining the target duration prediction information, it further includes:
[0033] Adjust the target duration prediction information according to the configured duration perturbation parameter to obtain the perturbed target duration prediction information, so as to guide the resampling module to perform resampling.
[0034] In a possible design, in another implementation manner of the second aspect of the embodiments of the present application, the process of obtaining the text corpus of the target domain includes:
[0035] Send the first prompt instruction prompt to the large language model to obtain the text corpus of the target domain generated by the large language model, where the first prompt instruction is used to instruct the model to generate the text corpus of the target domain.
[0036] The third aspect of the present application provides a speech recognition device, including:
[0037] An acoustic feature acquisition unit, configured to acquire the acoustic features of the speech signal to be recognized;
[0038] A model recognition unit, configured to input the acoustic features into the trained speech recognition model to obtain the speech recognition result output by the model; where:
[0039] The speech recognition model is jointly trained with the text reconstruction model in the training stage. The speech recognition model and the text reconstruction model share the same decoder. The text reconstruction model further includes a text encoder constructed based on the large language model. The total loss value of the joint training uses the audio sample as the input of the audio encoder in the speech recognition model, and the recognition text label corresponding to the audio sample as the input of the text encoder. The total loss value of the joint training includes the text recognition loss value calculated based on the text output by the decoder and the recognition text label, and the feature alignment loss value calculated based on the audio semantic representation extracted by the audio encoder and the text semantic representation extracted by the text encoder.
[0040] The fourth aspect of the present application provides a speech recognition model training device, including:
[0041] A training data acquisition unit for acquiring audio-text pair training data, where the audio-text pair training data includes audio samples and corresponding recognition text labels;
[0042] A model calculation unit for using the audio sample as the input of an audio encoder in a speech recognition model, and using the recognition text label as the input of a text encoder constructed based on a large language model in a text reconstruction model. The text reconstruction model and the speech recognition model share the same decoder. The audio encoder is used to extract audio semantic representations and send them to the decoder with a first probability. The text encoder is used to extract text semantic representations and send them to the decoder with a second probability. The decoder is used to output the decoded text, and the sum of the first probability and the second probability is equal to 1;
[0043] A loss value calculation unit for calculating a text recognition loss value based on the text output by the decoder and the recognition text label, calculating a feature alignment loss value based on the audio semantic representation and the text semantic representation, and calculating the total loss value of the text recognition loss value and the feature alignment loss value;
[0044] A parameter update unit for updating model parameters according to the total loss value until the training end condition is reached, and obtaining a trained speech recognition model.
[0045] In a fifth aspect of the present application, an electronic device is provided, including: a memory and a processor;
[0046] The memory is used for storing programs;
[0047] The processor is used for executing the program to implement the speech recognition method described in the first aspect of the present application, or to implement each step of the speech recognition model training method described in any one of the second aspects of the present application.
[0048] In a sixth aspect of the present application, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the speech recognition method described in the first aspect of the present application, or implements each step of the speech recognition model training method described in any one of the second aspects of the present application.
[0049] In a seventh aspect of the present application, a computer program product is provided, including a computer program. When the computer program is executed by a processor, it implements the speech recognition method described in the first aspect of the present application, or implements each step of the speech recognition model training method described in any one of the second aspects of the present application.
[0050] With the above technical solution, the present application adopts a joint training strategy to jointly train a speech recognition model and a text reconstruction model. The two share the same decoder. The text reconstruction model further includes a text encoder constructed based on a large language model, which can extract text semantic representations from the recognition text labels corresponding to audio samples. During the training process, the feature alignment loss value between the audio semantic representation extracted by the audio encoder from the audio sample and the text semantic representation is calculated, and the text recognition loss value between the text output by the decoder and the recognition text label is calculated. Through the feature alignment loss value, the semantic representations of audio and text are unified into the same semantic representation space. At the same time, since the large language model can extract rich semantic information, the rich semantic information of the large language model can be transferred to the audio encoder through the feature alignment loss value, enabling the audio encoder to extract more rich audio semantic representations, thereby improving the domain customization ability of the speech recognition model, that is, improving the recognition effect of the trained speech recognition model in a specific domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0052] Figure 1 is a flowchart of a method for training a speech recognition model provided by an embodiment of the present application;
[0053] Figure 2 illustrates a schematic diagram of a speech recognition model training architecture;
[0054] Figure 3 illustrates another schematic diagram of a speech recognition model training architecture;
[0055] Figure 4 illustrates yet another schematic diagram of a speech recognition model training architecture;
[0056] Figure 5 illustrates yet another schematic diagram of a speech recognition model training architecture;
[0057] Figure 6 is a schematic diagram of the structure of a speech recognition device provided by an embodiment of the present application;
[0058] Figure 7 is a schematic diagram of the structure of a speech recognition model training device provided by an embodiment of the present application;
[0059] Figure 8 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] Before introducing the solution of this application, first, the relevant concepts involved in this article are explained:
[0061] prompt: An instruction. When interacting with an AI (such as an artificial intelligence model), it is the instruction that needs to be sent to the AI. It can be a text description, such as "Please recommend a pop song for me" when you interact with the AI, or a parameter description in a certain format. For example, when asking the AI to draw a picture in a certain format, relevant drawing parameters need to be described.
[0062] Large language model: (Large language model, LLM) Generally refers to a language model with a large number of parameters and capabilities. It learns the statistical laws and semantic relationships of language through pre-training on a large amount of text data. These models usually use unsupervised learning methods to predict the next word or fill in the missing word to capture the context and semantic information of the language. Large language models can generate coherent sentences, answer questions, complete translation tasks, etc. The characteristic of LLM is its huge scale, containing billions or even more parameters, which helps them learn complex patterns in language data. The emerging capabilities of large language models include in-context learning, instruction following, and step-by-step reasoning capabilities, etc.
[0063] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0064] Traditional end-to-end speech recognition systems face many challenges when dealing with specific fields, such as professional fields like medical, legal, and financial.
[0065] These challenges mainly stem from the following aspects: First, since the terms and concepts in professional fields often have significant differences from everyday language and their frequency of occurrence in training data is extremely low, end-to-end speech recognition systems usually cannot accurately recognize these professional terms and concepts (such as diseases and drugs in the medical field, professional terms in the financial field, etc.); Second, although it is possible to perform domain adaptation training on the end-to-end speech recognition system by constructing domain-specific annotation data, the annotation work of domain-specific data requires huge costs, and some fields have a high degree of openness, making it difficult for the annotation data to cover all domain scenarios, resulting in insufficient generalization ability of the model when facing domain scenarios.
[0066] To improve the recognition performance of an end-to-end speech recognition system in a specific domain, one solution is to train an external language model (LM) using the text corpus of the target domain and integrate it into the end-to-end speech recognition model during the inference (decoding) process. During the decoding stage, the probabilities of the external language model are combined with those of the end-to-end speech recognition model to increase the probability of the decoding results related to the target domain, thereby improving the recognition performance of the end-to-end speech recognition system in the target domain.
[0067] Although this method can improve the recognition performance in the target domain, it involves using an external language model during inference, introducing a more complex model architecture, significantly increasing the computational cost during the decoding process, resulting in a slower inference speed, and making it inapplicable to offline devices with limited computational resources. In addition, this method also requires adjusting additional hyperparameters, increasing the training and deployment complexity of the end-to-end speech recognition system.
[0068] To avoid introducing an external language model (LM) during the decoding process and thus increasing the computational cost, one solution is to use text-to-speech (TTS) technology to generate paired speech and text data based on the text corpus of the target domain, and then use this data to train or fine-tune the end-to-end speech recognition system to improve the recognition performance of the end-to-end speech recognition system in the target domain.
[0069] However, the above solution requires introducing an additional TTS system. The generated speech data is usually based on rules or deep learning algorithms and cannot fully simulate the speech variations and complexities in the real world. There is a significant gap in data authenticity and diversity compared to real audio data, resulting in a decline in the performance of the speech recognition system in real-world applications.
[0070] Therefore, this application further provides a speech recognition method and a speech recognition model training method to address at least some of the problems existing in the several solutions of the foregoing examples and improve the recognition performance of the speech recognition model in a specific domain.
[0071] The following embodiments of this application respectively provide a speech recognition method and a speech recognition model training method. Among them, the provided speech recognition model training method can be applied to an electronic device. Exemplarily, the electronic device can be, for example, a server, a robot, a smartphone, a personal computer (PC), a laptop, a wireless electronic device in industrial control, etc.
[0072] The method provided in this application can be divided into a training stage and an inference stage. The training stage is the stage of training the speech recognition model, and the inference stage is the stage of using the trained speech recognition model to perform speech recognition on the speech signal to be recognized. Among them, the training stage and the inference stage can be deployed on the same device or on different devices. For example, the training stage can be deployed in the cloud or on a server, and the inference stage can be deployed in an intelligent terminal, such as a mobile phone, a tablet, a voice recorder, a translator, a robot, a vehicle-mounted terminal, or a wearable device, etc.
[0073] For ease of understanding, this application introduces the processes of the training stage and the inference stage respectively.
[0074] I. Training Stage
[0075] It should be noted that the training process of the model can usually be divided into multiple iterations or multiple executions. In this application, by way of example, one iteration or execution process is taken as an example for exemplary introduction. The steps mentioned below can be repeatedly executed and will not be elaborated further.
[0076] Refer to Figure 1 , the schematic flow chart of a training method for a speech recognition model provided in this application is as follows. The method includes:
[0077] Step S100: Obtain audio-text pair training data, where the audio-text pair training data includes audio samples and corresponding recognition text labels.
[0078] Specifically, audio samples and paired recognition text labels can be collected to form audio-text pair training data.
[0079] Step S110: Jointly train the speech recognition model and the text reconstruction model using the audio-text pair training data. The two share the same decoder. The text reconstruction model also includes a text encoder constructed based on a large language model. In the joint training process, the audio sample is used as the input of the audio encoder in the speech recognition model, and the recognition text label is used as the input of the text encoder.
[0080] Specifically, as shown in Figure 2 , the speech recognition model can include an audio encoder and a decoder. The audio sample is used as the input of the audio encoder, and the audio semantic representation of the audio sample is extracted through the audio encoder.
[0081] The text reconstruction model includes a text encoder and a decoder. The speech recognition model and the text reconstruction model share the same decoder.
[0082] The text encoder is built based on a large language model, specifically the ability of the large language model to extract rich text semantic representations. During the training phase, the recognized text labels corresponding to the audio samples can be input into the text encoder to obtain the text semantic representations extracted by the text encoder. In this embodiment, in order to facilitate the calculation of the feature alignment loss between the text semantic representation and the audio semantic representation in subsequent steps, the text semantic representation extracted by the text encoder here can be a feature with the same dimension as the audio semantic representation.
[0083] For a piece of audio-text pair training data, only one of the audio semantic representation extracted by the audio encoder and the text semantic representation extracted by the text encoder can be sent into the decoder. In this embodiment, the first probability P1 and the second probability (1 - P1) can be preset. The audio semantic representation is sent into the decoder with the first probability P1, and the text semantic representation is sent into the decoder with the second probability (1 - P1). The size of P1 can be set by the user, and its value range is (0, 1).
[0084] The decoder decodes based on the input feature representation to obtain the output text.
[0085] It can be understood that when the input of the decoder is the audio semantic representation, the text output by the decoder can be regarded as the speech recognition result; when the input of the decoder is the text semantic representation, the text output by the decoder can be regarded as the reconstructed text.
[0086] Step S120: Calculate the text recognition loss value based on the text output by the decoder and the recognized text label, calculate the feature alignment loss value based on the audio semantic representation and the text semantic representation, and calculate the total loss value of the text recognition loss value and the feature alignment loss value.
[0087] Combined with Figure 2 As shown, when the input of the decoder is the audio semantic representation, the text recognition loss value calculated between the text output by the decoder and the recognized text label can specifically be the speech recognition loss value; when the input of the decoder is the text semantic representation, the text recognition loss value calculated between the text output by the decoder and the recognized text label can specifically be the text reconstruction loss value. In this embodiment, for simplicity of expression, the text recognition loss function is defined as ASR_Loss.
[0088] The parameters of the encoder and the decoder as a whole can be updated through the text recognition loss value.
[0089] For the audio semantic representation output by the audio encoder and the text semantic representation output by the text encoder, calculate the feature alignment loss value between the two. Combined with Figure 2 As shown, the feature alignment loss function is defined as MSE_Loss.
[0090] Under the constraint of the feature alignment loss function, the rich semantic information of the large language model (the text encoder is built based on the large language model and has the feature processing ability of the large language model) can be transferred to the audio encoder, enabling the audio encoder to extract richer audio semantic representations. On the premise of limited audio-text pair training data, the text data can be fully utilized to improve the feature extraction ability of the audio encoder and enhance the recognition effect of the speech recognition model in a specific domain.
[0091] Step S130: Update the model parameters according to the total loss value until the training end condition is reached, and obtain the trained speech recognition model.
[0092] Specifically, when updating the parameters of the speech recognition model and the text reconstruction model according to the above total loss value, for the large language model in the text encoder, a pre-trained large language model can be used and does not participate in the parameter update process. Of course, in some other optional implementation manners, the large language model can also jointly participate in the parameter update with other network modules.
[0093] After reaching the set training end condition, the trained speech recognition model can be obtained.
[0094] Obviously, in the speech recognition model training method provided in this embodiment, the text reconstruction model is used to assist in jointly training the speech recognition model, and the semantic representations of audio and text are unified into the same semantic representation space through the feature alignment loss value. At the same time, since the large language model can extract rich semantic information, the rich semantic information of the large language model can be transferred to the audio encoder through the feature alignment loss value, enabling the audio encoder to extract richer audio semantic representations, thereby improving the domain customization ability of the speech recognition model, that is, enhancing the recognition effect of the trained speech recognition model in a specific domain.
[0095] In the speech recognition model training method provided in this embodiment, the trained speech recognition model does not need to load a language model inside or outside the model, which simplifies the structure of the speech recognition model and reduces the inference time. In addition, this application does not need to train the speech recognition model through audio synthesis technology to synthesize audio-text pair training data. By fully utilizing the text data in the existing audio-text pair training data, it guides the audio encoder to extract richer audio semantic representations, thereby improving the domain customization ability of the speech recognition model and enhancing the recognition effect of the speech recognition model in a specific domain.
[0096] Refer to Figure 3 , another speech recognition model training architecture is provided. Compared with the previous embodiment, the speech recognition model in this embodiment can further include a shared encoder, which is shared by the speech recognition model and the text reconstruction model during the training phase.
[0097] The shared encoder is located between the audio encoder and decoder, and also between the text encoder and decoder. The audio semantic representation extracted by the audio encoder is sent to the shared encoder with a first probability P1, and the text semantic representation extracted by the text encoder is sent to the shared encoder with a second probability (1 - P1).
[0098] The shared encoder performs re-encoding processing on the input feature representation and sends the encoded features to the decoder.
[0099] In this embodiment, by adding a shared encoder between the audio encoder and decoder of the speech recognition model, during the joint training phase, when the text semantic representation is sent to the shared encoder, the loss function at this time is the text reconstruction loss function. When updating the parameters according to this text reconstruction loss function, the parameters of the decoder and the shared encoder can be updated simultaneously, that is, the network modules participating in parameter update further include the shared encoder. In an end-to-end speech recognition model, the role of the encoder is usually greater than that of the decoder. Therefore, by adding the shared encoder to participate in parameter update, the shared encoder can also be fine-tuned in the domain, further improving the recognition effect of the trained speech recognition model on a specific domain.
[0100] Referring to Figure 4 , another speech recognition model training architecture is provided. Compared with the foregoing embodiments, the speech recognition model in this embodiment may further include an alignment module.
[0101] When the input of the shared encoder is the audio semantic representation, the alignment module is used to predict audio-text alignment information based on the audio semantic representation, that is, the alignment relationship between each frame in the audio sample and each token unit in the text.
[0102] On this basis, the connectionist temporal classification (CTC) loss can be calculated based on the audio-text alignment information and the recognized text label, and the total loss value in the joint training process can further include the CTC loss value.
[0103] In this embodiment, by adding the CTC loss function, it can guide the speech recognition model to learn the alignment relationship between audio and text and improve the recognition effect of the speech recognition model.
[0104] It should be noted that after training is completed according to the Figure 4 architecture, only the audio encoder, the shared encoder, and the decoder can be taken to form the final speech recognition model for subsequent inference use.
[0105] Referring to Figure 4 as shown, for the text encoder introduced in the foregoing embodiments, an optional composition structure of the text encoder is provided in this embodiment, which may include: a large language model, a resampling module, and a duration prediction module.
[0106] The large language model is used to extract the initial text semantic representation from the input text.
[0107] When the input of the shared encoder is the audio semantic representation, the resampling module is used to resample the initial text semantic representation based on the audio-text alignment information output by the alignment module, so as to obtain the resampled text semantic representation with the same number of frames as the audio semantic representation.
[0108] Specifically, due to the difference in the granularity of the modeling unit, the text sequence is usually converted into unit sequences such as character sequences, pinyin sequences, and phoneme sequences and sent into the large language model. The number of frames of the text sequence semantic representation extracted by the large language model is usually less than that of the audio semantic representation extracted by the acoustic encoder. In this embodiment, the resampling module is used to expand the frames of the text semantic representation so that the number of frames is consistent with the audio semantic representation, and then the feature alignment loss MSE_Loss is calculated using the resampled text semantic representation and the audio semantic representation.
[0109] The duration prediction module is used to predict the duration of each frame in the initial text semantic representation to obtain duration prediction information. On this basis, the duration prediction loss value can be calculated based on the duration prediction information and the audio-text alignment information output by the alignment module, and the duration prediction loss function is defined as Duration_Loss.
[0110] Then the total loss value of the joint training process can also include the above duration prediction loss value.
[0111] In summary, in the training process of the speech recognition model provided in this embodiment, the text reconstruction model is jointly trained. The total loss function of the joint training process can be expressed as:
[0112] Loss total =ASR_Loss + αMSE_Loss + βCTC_Loss + γDuration_Loss;
[0113] Among them, α, β, and γ are the weights of MSE_Loss, CTC_Loss, and Duration_Loss respectively.
[0114] The speech recognition model training method provided in this embodiment designs the structure of the text encoder. While ensuring its ability to extract rich semantic features of the large language model, the resampling module is guided by the alignment information output by the alignment module to perform feature resampling, so that the dimension of the resampled text semantic representation is the same as that of the audio semantic representation, facilitating the calculation of MSE_Loss. In addition, the text encoder also includes a duration prediction module, which can predict the duration information of each frame of the initial text semantic representation extracted by the large language model, and calculate the duration prediction loss value by combining the duration prediction information and the audio-text alignment information output by the alignment module to guide the training process of the duration prediction module.
[0115] When there is a lack of audio-text pairs in the training data, only the text data in the target domain can be obtained at this time. After extracting the initial text semantic representation through the large language model, the duration prediction module predicts the duration information of each frame of the feature in the initial text semantic representation, and then guides the resampling module to perform resampling. Since the text semantic representation and the audio semantic representation are mapped to the same semantic representation space in the previous training stage, the resampled text semantic representation can simulate the audio semantic representation extracted by the corresponding audio encoder at this time, and is sent to the subsequent shared encoder and decoder for processing, and the decoded text is output, and the text recognition loss value between the decoded text and the input text is calculated, and the model parameter update process is guided according to the text recognition loss value. Under the condition of lacking the corresponding audio, the domain adaptation effect of the shared encoder can be greatly improved.
[0116] After the training according to the training strategy of the foregoing embodiment is completed, in this embodiment, a process of further performing domain adaptation training on the speech recognition model using the text in the target domain can be added.
[0117] Combined Figure 5 As shown, there is only domain text and no corresponding audio in this stage. Therefore, the network modules related to the audio branch in the corresponding training architecture can be removed, and the training architecture in this stage is as Figure 4 shown, including a large language model, a duration prediction module, a resampling module, a shared encoder, and a decoder. Figure 5 After the training strategy introduced in the foregoing embodiment, the duration prediction module already has the ability to predict duration information.
[0118] The domain adaptation training stage may include the following training steps:
[0119] S1. Obtain the text corpus in the target domain.
[0120] Specifically, the target domain can be the domain where the speech recognition model is to be applied. For example, if the speech recognition model is to be applied in the medical field, the medical field can be used as the target domain in this step.
[0121] Specifically, the target domain can be the domain where the speech recognition model is to be applied. For example, if the speech recognition model is to be applied in the medical field, the medical field can be used as the target domain in this step.
[0122] For scenarios where there is a scarcity of paired audio and text data in the target domain, in this embodiment, only the text corpus of the target domain can be obtained. Compared with audio data, the text corpus is easier to obtain.
[0123] In this step, the text corpus of the target domain can be obtained from a public dataset or automatically generated. This embodiment provides a method for automatically generating the text corpus of the target domain, which can utilize the text generation ability of a large language model. Specifically:
[0124] Send the first prompt instruction "prompt" into the large language model to obtain the text corpus of the target domain generated by the large language model. Herein, the first prompt instruction is used to instruct the model to generate the text corpus of the target domain.
[0125] Among them, the first prompt instruction "prompt" may include the target domain theme, reference keywords, reference examples, word count requirements for the generated text, text magnitude requirements, etc. The following provides an example of the first prompt instruction "prompt":
[0126] "The target domain is medical conferences, related themes include Theme 1, Theme 2, Theme 3, etc., possible keywords include Keyword 1, Keyword 2, Keyword 3, etc., and typical sample texts in this domain are Sample Text 1, Sample Text 2, Sample Text 3. Please generate 1000 lines of text related to this domain based on the above information, with each line having a word count between 10 and 50."
[0127] S2. Send the text corpus into the large language model to obtain the target text semantic representation of the text corpus.
[0128] As shown in combination with Figure 5 Take the text corpus generated in the previous step as the input text sequence and input it into the large language model, and use the large language model to extract the target text semantic representation of the input text corpus.
[0129] S3. Predict the duration of each frame in the target text semantic representation through the duration prediction module to obtain the target duration prediction information.
[0130] After the training process of the foregoing embodiment, the duration prediction module already has the duration prediction ability. In this step, the duration prediction module can be used to predict the duration of each frame in the target text semantic representation to obtain the target duration prediction information.
[0131] S4. Through the resampling module, resample the target text semantic representation according to the target duration prediction information to obtain the resampled text semantic representation.
[0132] Specifically, under the guidance of the target duration prediction information, the resampling module resamples the target text semantic representation to obtain the resampled text semantic representation. Since the text semantic representation and the audio semantic representation are mapped to the same semantic representation space in the aforementioned training stage, the resampled text semantic representation can simulate the audio semantic representation extracted by the corresponding audio encoder at this time.
[0133] S5. Encode the resampled text semantic representation through a shared encoder, and send the encoded features to the decoder to predict the decoded text through the decoder.
[0134] S6. Calculate the text reconstruction loss value based on the decoded text predicted by the decoder and the text corpus, and update the model parameters according to the text reconstruction loss value until the training end condition is reached. The final speech recognition model is composed of an audio encoder, a shared encoder, and a decoder.
[0135] Similar to the previous text, define the text reconstruction loss function as ASR_Loss. Then, using the text corpus obtained in step S1 as the text label, calculate the text reconstruction loss value between the decoded text predicted by the decoder and the text label. The text reconstruction loss function can adopt various loss function types. For example, the cross-entropy CE loss function can be used.
[0136] After the above domain adaptation training stage, the text corpus of the target domain can be used to perform domain adaptation fine-tuning on the shared encoder and the decoder, further improving the recognition effect of the speech recognition model in the target domain.
[0137] In the above domain adaptation training stage, the parameters of the large language model can be updated synchronously. The loss function in the domain adaptation training stage can be expressed as:
[0138] Loss total =ASR_Loss;
[0139] Furthermore, before steps S3 and S4 above, the following steps can be further added:
[0140] Adjust the target duration prediction information according to the configured duration perturbation parameter to obtain the perturbed target duration prediction information, so as to guide the resampling module in step S4 to perform resampling.
[0141] That is, in the domain adaptation training stage, by adding appropriate duration perturbation parameters to the duration prediction module (such as multiplying the predicted duration of each frame by a parameter greater than 1 or less than 1 with a set probability P2), the generated resampled text semantic representation can be made more abundant, thereby better training the shared encoder and the decoder.
[0142] In summary, for the speech recognition model training method provided in this application, on the one hand, there is no need for an external language model for decoding fusion, which will not additionally increase the inference calculation amount and will not increase the complexity of training and deployment; on the other hand, by leveraging the rich semantic representation ability of the large language model, the audio semantic representation and the text semantic representation are mapped to the same semantic representation space, and the rich semantic feature extraction ability of the large language model is transferred to the audio encoder, improving the semantic perception ability of the audio encoder.
[0143] In addition, this application further adds a domain adaptation training stage, and uses the text corpus of the target domain to adaptively train the model. This process can simultaneously perform domain fine-tuning on both the shared encoder and the decoder, improving the customization ability of the speech recognition model for the target domain. It can simulate the diversity of audio semantic representations under the condition of lacking audio-text pairs for training data, and has a better domain adaptation effect.
[0144] Based on the speech recognition model training method introduced in any of the above embodiments, the speech to be recognized can be recognized based on the trained speech recognition model.
[0145] In some embodiments of this application, a speech recognition method is further provided, which may include the following steps:
[0146] S1. Obtain the acoustic features of the speech signal to be recognized.
[0147] S2. Input the acoustic features into the trained speech recognition model to obtain the speech recognition result output by the model; where:
[0148] During the training stage, the speech recognition model is jointly trained with the text reconstruction model. The speech recognition model and the text reconstruction model share the same decoder. The text reconstruction model further includes a text encoder constructed based on the large language model. During the joint training process, the audio sample is used as the input of the audio encoder in the speech recognition model, and the recognition text label corresponding to the audio sample is used as the input of the text encoder. The total loss value of the joint training includes the text recognition loss value calculated based on the text output by the decoder and the recognition text label, and the feature alignment loss value calculated based on the audio semantic representation extracted by the audio encoder and the text semantic representation extracted by the text encoder.
[0149] For the training process of the speech recognition model, for details, reference can be made to the introduction of the training process of the speech recognition model in the foregoing embodiments, which will not be elaborated in this embodiment.
[0150] After being trained by the above training strategy, the speech recognition model has the customization ability for the target domain. Therefore, for the speech signal to be recognized in the target domain, a better recognition effect can be obtained by using the speech recognition model.
[0151] The voice recognition device provided by the embodiments of the present application will be described below. The voice recognition device described below can be correspondingly referred to the voice recognition method described above.
[0152] See Figure 6 , Figure 6 which is a schematic structural diagram of a voice recognition device disclosed in the embodiments of the present application.
[0153] As Figure 6 shown, the device may include:
[0154] An acoustic feature acquisition unit 11, configured to acquire acoustic features of a voice signal to be recognized;
[0155] A model recognition unit 12, configured to input the acoustic features into a trained voice recognition model to obtain a voice recognition result output by the model; wherein:
[0156] The voice recognition model is jointly trained with a text reconstruction model in the training stage. The voice recognition model and the text reconstruction model share the same decoder. The text reconstruction model further includes a text encoder constructed based on a large language model. In the joint training process, an audio sample is used as the input of the audio encoder in the voice recognition model, and the recognition text label corresponding to the audio sample is used as the input of the text encoder. The total loss value of the joint training includes a text recognition loss value calculated based on the text output by the decoder and the recognition text label, and a feature alignment loss value calculated based on the audio semantic representation extracted by the audio encoder and the text semantic representation extracted by the text encoder.
[0157] Furthermore, the voice recognition model training device provided by the embodiments of the present application will be described. The voice recognition model training device described below can be correspondingly referred to the voice recognition model training method described above.
[0158] See Figure 7 , Figure 7 which is a schematic structural diagram of a voice recognition model training device disclosed in the embodiments of the present application.
[0159] As Figure 7 shown, the device may include:
[0160] A training data acquisition unit 21, configured to acquire audio-text pair training data, where the audio-text pair training data includes an audio sample and a corresponding recognition text label;
[0161] The model calculation unit 22 is configured to use the audio sample as the input of the audio encoder in the speech recognition model, and use the recognition text label as the input of the text encoder based on the large language model in the text reconstruction model. The text reconstruction model and the speech recognition model share the same decoder. The audio encoder is configured to extract audio semantic representations and send them to the decoder with a first probability. The text encoder is configured to extract text semantic representations and send them to the decoder with a second probability. The decoder is configured to output the decoded text, and the sum of the first probability and the second probability is equal to 1;
[0162] The loss value calculation unit 23 is configured to calculate the text recognition loss value based on the text output by the decoder and the recognition text label, calculate the feature alignment loss value based on the audio semantic representation and the text semantic representation, and calculate the total loss value of the text recognition loss value and the feature alignment loss value;
[0163] The parameter update unit 24 is configured to update the model parameters according to the total loss value until the training end condition is reached, and obtain the trained speech recognition model.
[0164] In a possible implementation, the speech recognition model further includes a shared encoder located between the audio encoder and the decoder. The shared encoder encodes the input features and sends the encoded features to the decoder. The audio semantic representation extracted by the audio encoder is sent to the shared encoder with a first probability, and the text semantic representation extracted by the text encoder is sent to the shared encoder with a second probability.
[0165] In a possible implementation, the speech recognition model further includes an alignment module;
[0166] When the input of the shared encoder is the audio semantic representation, the alignment module is configured to predict audio-text alignment information based on the audio semantic representation;
[0167] The total loss value further includes: the connectionist temporal classification (CTC) loss calculated based on the audio-text alignment information and the recognition text label.
[0168] In a possible implementation, the text encoder includes a large language model, a resampling module, and a duration prediction module;
[0169] The large language model is configured to extract the initial text semantic representation from the input text;
[0170] When the input of the shared encoder is the audio semantic representation, the resampling module is used to resample the initial text semantic representation based on the audio-text alignment information output by the alignment module to obtain a resampled text semantic representation with the same number of frames as the audio semantic representation, and calculate the feature alignment loss value using the resampled text semantic representation and the audio semantic representation;
[0171] The duration prediction module is used to predict the duration of each frame in the initial text semantic representation to obtain duration prediction information;
[0172] The total loss value further includes: a duration prediction loss value calculated based on the duration prediction information and the audio-text alignment information.
[0173] In a possible implementation, the device of the present application may further include: a domain adaptation fine-tuning training unit, which is used to further perform model training using a domain adaptation fine-tuning training strategy after the parameter update unit updates the model parameters according to the total loss value until the training end condition is reached. This process may include:
[0174] Obtain the text corpus of the target domain;
[0175] Send the text corpus into the large language model to obtain the target text semantic representation of the text corpus;
[0176] Predict the duration of each frame in the target text semantic representation through the duration prediction module to obtain target duration prediction information;
[0177] Through the resampling module, resample the target text semantic representation according to the target duration prediction information to obtain a resampled text semantic representation;
[0178] Encode the resampled text semantic representation through the shared encoder, and send the encoded features into the decoder, and predict the decoded text through the decoder;
[0179] Calculate a text reconstruction loss value based on the decoded text predicted by the decoder and the text corpus, and update the model parameters according to the text reconstruction loss value until the training end condition is reached, and the final speech recognition model is composed of the audio encoder, the shared encoder, and the decoder.
[0180] In a possible implementation, after obtaining the target duration prediction information, the domain adaptation fine-tuning training unit may also be used to adjust the target duration prediction information according to the configured duration perturbation parameter to obtain a perturbed target duration prediction information to guide the resampling module to perform resampling.
[0181] In one possible implementation, the process by which the domain adaptation fine-tuning training unit obtains the text corpus of the target domain includes:
[0182] Sending the first prompt instruction "prompt" to the large language model to obtain the text corpus of the target domain generated by the large language model, where the first prompt instruction is used to instruct the model to generate the text corpus of the target domain.
[0183] An embodiment of the present application also provides an electronic device. Refer to Figure 8 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, fixed terminals such as servers, personal computers, mobile phones, and the like. Figure 8 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0184] As Figure 8 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603 to implement the speech recognition method or the speech recognition model training method in the foregoing embodiments of the present application. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0185] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 shows an electronic device having various devices, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.
[0186] An embodiment of the present application also provides a computer program product, including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement any speech recognition method or speech recognition model training method provided in the embodiments of the present application.
[0187] In an embodiment of the present application, a computer-readable storage medium is further provided. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any voice recognition method or voice recognition model training method provided by the embodiment of the present application.
[0188] In addition, it should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in the present application, the connection relationship between modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.
[0189] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for the present application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in various embodiments of the present application.
[0190] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0191] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0192] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.
Claims
1. A speech recognition method, characterized in that: include: Acquiring acoustic features of a speech signal to be recognized; The acoustic features are input into the trained speech recognition model to obtain the speech recognition results output by the model; wherein: The speech recognition model is jointly trained with the text reconstruction model during the training phase. The speech recognition model and the text reconstruction model share the same decoder. The text reconstruction model also includes a text encoder built based on a large language model. The joint training process uses audio samples as input to the audio encoder in the speech recognition model, and uses the recognized text labels corresponding to the audio samples as input to the text encoder. The total loss value of the joint training includes a text recognition loss value calculated based on the text output by the decoder and the recognized text label, and a feature alignment loss value calculated based on the audio semantic representation extracted by the audio encoder and the text semantic representation extracted by the text encoder.
2. A speech recognition model training method, characterized in that: include: Acquire audio-text pair training data, wherein the audio-text pair training data includes audio samples and corresponding recognition text labels; The audio sample is used as the input of an audio encoder in a speech recognition model, and the recognized text label is used as the input of a text encoder constructed based on a large language model in a text reconstruction model, wherein the text reconstruction model and the speech recognition model share the same decoder, wherein the audio encoder is used to extract an audio semantic representation and feed it into the decoder with a first probability, wherein the text encoder is used to extract a text semantic representation and feed it into the decoder with a second probability, wherein the decoder is used to output a decoded text, and the sum of the first probability and the second probability is equal to 1; Calculating a text recognition loss value based on the text output by the decoder and the recognized text label, calculating a feature alignment loss value based on the audio semantic representation and the text semantic representation, and calculating a total loss value of the text recognition loss value and the feature alignment loss value; The model parameters are updated according to the total loss value until the training end condition is reached, thereby obtaining a trained speech recognition model.
3. The method according to claim 2, characterized in that The speech recognition model also includes a shared encoder located between the audio encoder and the decoder, which encodes input features and sends the encoded features to the decoder; the audio semantic representation extracted by the audio encoder is sent to the shared encoder with a first probability, and the text semantic representation extracted by the text encoder is sent to the shared encoder with a second probability.
4. The method according to claim 3, characterized in that The speech recognition model also includes an alignment module; When the input of the shared encoder is the audio semantic representation, the alignment module is used to predict audio-text alignment information based on the audio semantic representation; The total loss value also includes: a temporal classification CTC loss calculated based on the audio-text alignment information and the recognition text label.
5. The method according to claim 4, characterized in that The text encoder includes a large language model, a resampling module and a time length prediction module; The large language model is used to extract initial text semantic representation from the input text; When the input of the shared encoder is the audio semantic representation, the resampling module is used to resample the initial text semantic representation based on the audio-text alignment information output by the alignment module to obtain a resampled text semantic representation with the same frame number as the audio semantic representation, and calculate the feature alignment loss value using the resampled text semantic representation and the audio semantic representation; The duration prediction module is used to predict the duration of each frame in the initial text semantic representation to obtain duration prediction information; The total loss value also includes: a duration prediction loss value calculated based on the duration prediction information and the audio-text alignment information.
6. The method according to claim 5, characterized in that After updating the model parameters according to the total loss value until the training end condition is reached, the method further includes: Obtain text corpus in the target domain; Sending the text corpus into the large language model to obtain a target text semantic representation of the text corpus; Predicting the duration of each frame in the semantic representation of the target text by the duration prediction module to obtain target duration prediction information; Resampling the target text semantic representation according to the target duration prediction information through the resampling module to obtain a resampled text semantic representation; Encoding the resampled text semantic representation through the shared encoder, sending the encoded features to the decoder, and predicting the decoded text through the decoder; A text reconstruction loss value is calculated based on the decoded text predicted by the decoder and the text corpus, and the model parameters are updated according to the text reconstruction loss value until the training end condition is reached. The audio encoder, the shared encoder and the decoder form a final speech recognition model.
7. The method according to claim 6, characterized in that After obtaining the target duration prediction information, it also includes: The target duration prediction information is adjusted according to the configured duration disturbance parameter to obtain the disturbed target duration prediction information to guide the resampling module to perform resampling.
8. The method according to claim 6 or 7, characterized in that: The process of obtaining text corpus in the target domain includes: The first prompt instruction prompt is sent to the large language model to obtain the text corpus of the target domain generated by the large language model, wherein the first prompt instruction is used to instruct the model to generate the text corpus of the target domain.
9. A speech recognition device, characterized in that: include: An acoustic feature acquisition unit, used to acquire acoustic features of a speech signal to be recognized; A model recognition unit is used to input the acoustic features into the trained speech recognition model to obtain the speech recognition results output by the model; wherein: The speech recognition model is jointly trained with the text reconstruction model during the training phase. The speech recognition model and the text reconstruction model share the same decoder. The text reconstruction model also includes a text encoder built based on a large language model. The joint training process uses audio samples as input to the audio encoder in the speech recognition model, and uses the recognized text labels corresponding to the audio samples as input to the text encoder. The total loss value of the joint training includes a text recognition loss value calculated based on the text output by the decoder and the recognized text label, and a feature alignment loss value calculated based on the audio semantic representation extracted by the audio encoder and the text semantic representation extracted by the text encoder.
10. A speech recognition model training device, characterized in that: include: A training data acquisition unit, used to acquire audio-text pair training data, wherein the audio-text pair training data includes an audio sample and a corresponding recognition text label; A model calculation unit, used to use the audio sample as an input of an audio encoder in a speech recognition model, and use the recognized text label as an input of a text encoder constructed based on a large language model in a text reconstruction model, wherein the text reconstruction model and the speech recognition model share the same decoder, wherein the audio encoder is used to extract an audio semantic representation and feed it into the decoder with a first probability, wherein the text encoder is used to extract a text semantic representation and feed it into the decoder with a second probability, wherein the decoder is used to output a decoded text, and wherein the sum of the first probability and the second probability is equal to 1; a loss value calculation unit, configured to calculate a text recognition loss value based on the text output by the decoder and the recognized text label, calculate a feature alignment loss value based on the audio semantic representation and the text semantic representation, and calculate a total loss value of the text recognition loss value and the feature alignment loss value; A parameter updating unit is used to update the model parameters according to the total loss value until the training end condition is reached to obtain a trained speech recognition model.
11. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement the speech recognition method as described in claim 1, or to implement each step of the speech recognition model training method as described in any one of claims 2 to 8.
12. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the speech recognition method as claimed in claim 1, or implements the various steps of the speech recognition model training method as claimed in any one of claims 2 to 8.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the speech recognition method as claimed in claim 1, or implements the various steps of the speech recognition model training method as claimed in any one of claims 2 to 8.
Citation Information
Patent Citations
Voice recognition method and device, electronic equipment and storage medium
CN113643694A
Vietnamese speech recognition corpus construction method
CN115223549A
Speech recognition method and device, electronic equipment and storage medium
CN117711378A
Speech recognition method based on large language model
CN118447827A
Efficient self-adaptive speech recognition engine-oriented hot word error correction method and system
CN118471201A
Cited By
Audio understanding model training method and device, audio understanding method and device, storage medium and program product
CN120356465A