Audio-to-text model training method, audio-to-text method and device, electronic equipment, storage medium and computer program product
By using a training method combining real audio samples and virtual audio samples, the problem of poor training effect of streaming speech recognition model caused by a single data set is solved, and higher training effect and recognition accuracy are achieved.
Patent Information
- Application Number
- CN202510258567.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, streaming speech recognition models usually use a single type of data set during training, resulting in poor training effects and affecting the recognition accuracy.
Two different types of data sets are trained, real audio samples and virtual audio samples. By obtaining audio samples, inputting audio to text model, calculating losses, and adjusting model parameters according to the losses to improve training effect.
By using multiple types of data sets for training, the amount of training data and content richness is increased, and the training effect and recognition accuracy of the audio-to-text model are improved.
Smart Images

Figure CN120048264A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and more particularly, to a method for training an audio-to-text model, an audio-to-text method, an apparatus, an electronic device, a storage medium, and a computer program product. Background Art
[0002] Automatic Speech Recognition (ASR), also known as Stream Speech Recognition, is used to convert a user's speech into text content for output. Specifically, after receiving the speech input of a speaker, an ASR system can convert the speech signal into corresponding text content for output by performing steps such as audio processing, speech feature extraction, and model matching.
[0003] In related technologies, when training a stream speech recognition model, a single type of dataset is often used for training, and the amount of training data is small and the content of the training data is not rich enough. This will result in a poor training effect of the stream speech recognition model, thereby affecting the recognition accuracy of the stream speech recognition model. Summary of the Invention
[0004] The present disclosure provides a method for training an audio-to-text model, an audio-to-text method, an apparatus, an electronic device, a storage medium, and a computer program product to at least solve the problem in the above-mentioned related technologies that using a single type of dataset for training results in a poor training effect of the stream speech recognition model, thereby affecting the recognition accuracy of the stream speech recognition model.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a method for training an audio-to-text model, including: obtaining an audio sample, where the audio sample includes a real audio sample and a virtual audio sample, the real audio sample is a sample obtained by collecting a user's real speech, the virtual audio sample is a sample obtained by converting known text into audio using a speech synthesis engine, and the audio sample corresponds to a text content label; inputting the audio sample into the audio-to-text model to obtain a predicted text content; calculating a loss based on the predicted text content and the text content label; and training the audio-to-text model by adjusting parameters of the audio-to-text model according to the loss.
[0006] Optionally, the step of inputting the audio sample into the audio-to-text model to obtain the predicted text content includes: inputting the virtual audio sample into the audio-to-text model to obtain the predicted virtual text content; inputting the real audio sample into the audio-to-text model to obtain the predicted real text content; wherein, calculating the loss based on the predicted text content and the text content label includes: calculating a first loss based on the predicted virtual text content and the virtual text content label corresponding to the virtual audio sample; calculating a second loss based on the predicted real text content and the real text content label corresponding to the real audio sample; wherein, training the audio-to-text model by adjusting the parameters of the audio-to-text model according to the loss includes: training the audio-to-text model by adjusting the parameters of the audio-to-text model according to the first loss to obtain a first audio-to-text model; fine-tuning the parameters of the first audio-to-text model according to the second loss to obtain the finally trained audio-to-text model.
[0007] Optionally, before inputting the audio sample into the audio-to-text model to obtain the predicted text content, it further includes: performing sample enhancement on the audio sample to obtain an enhanced audio sample; the step of inputting the audio sample into the audio-to-text model to obtain the predicted text content includes: inputting the enhanced audio sample into the audio-to-text model to obtain the predicted text content.
[0008] Optionally, the sample enhancement includes at least one of the following: changing the speech rate of the audio sample, changing the pitch of the audio sample, changing the gain of the audio sample, adding background noise to the audio sample, performing truncation processing on the audio sample, and randomly adjusting the spectrum corresponding to the audio sample.
[0009] According to a second aspect of the embodiments of the present disclosure, there is provided an audio-to-text method, including: obtaining audio data; inputting the obtained audio data into an audio-to-text model to obtain the text content corresponding to the output audio data, wherein the audio-to-text model is trained according to the training method of the present disclosure.
[0010] Optionally, before inputting the obtained audio data into an audio-to-text model to obtain the text content corresponding to the output audio data, it further includes: determining at least one blank audio region included in the audio data, where the blank audio region is an audio region with a volume lower than a preset volume threshold; determining a target blank audio region in the at least one blank audio region with a duration greater than or equal to a preset duration threshold; based on the target blank audio region, splitting the audio data into multiple audio segments, where the audio data between two adjacent target blank audio regions is split into one audio segment; the step of inputting the obtained audio data into the audio-to-text model to obtain the text content corresponding to the output audio data includes: sequentially inputting the multiple audio segments into the audio-to-text model in chronological order to obtain the text content corresponding to each sequentially output audio segment.
[0011] Optionally, the step of inputting the obtained audio data into the audio-to-text model includes: obtaining a first input at the current moment, where the first input at least includes a first audio segment located in an audio buffer, and the audio buffer is used to limit the processing length of the real-time obtained audio stream; inputting the first input into the audio-to-text model to obtain a first text content; obtaining a second input at the next moment of the current moment, where the second input at least includes a second audio segment located in the audio buffer; inputting the second input into the audio-to-text model to obtain a second text content; comparing the text content of the same audio segment included in the first text content with the text content of the same audio segment included in the second text content to obtain a comparison result, where the text content of the same audio segment is the text content corresponding to the overlapping audio segment of the first audio segment and the second audio segment; and outputting the text content of the same audio segment when the comparison result indicates consistency.
[0012] Optionally, the step of comparing the text content of the same audio segment included in the first text content with the text content of the same audio segment included in the second text content to obtain a comparison result includes: comparing a partial text content located at a preset position included in the text content of the same audio segment included in the first text content with a partial text content located at the preset position included in the text content of the same audio segment included in the second text content to obtain the comparison result.
[0013] Optionally, the obtaining of the second input at the next moment of the current moment includes: determining whether the text content of the same audio segment between the current moment and the previous moment contains a complete sentence; in the case where it is determined that the text content of the same audio segment contains the complete sentence, deleting the complete sentence audio segment corresponding to the complete sentence in the first audio segment; using the remaining audio segment in the first audio segment except the complete sentence audio segment, the audio segment that is continuously rolled into the audio buffer in real time during the process from the current moment to the next moment, and the complete sentence content as the second input.
[0014] Optionally, the audio-to-text method further includes: adding the complete sentence content to a complete sentence content storage space, where the complete sentence content storage space stores historical complete sentence contents obtained at each moment before the current moment and arranged in chronological order; the step of using the remaining audio segment in the first audio segment except the complete sentence audio segment, the audio segment that is continuously rolled into the audio buffer in real time during the process from the current moment to the next moment, and the complete sentence content as the second input includes: using the remaining audio segment, the audio segment that is continuously rolled into the audio buffer in real time during the process from the current moment to the next moment, the complete sentence content, and the historical complete sentence content as the second input.
[0015] Optionally, before inputting the obtained audio data into the audio-to-text model, it further includes: determining the speech data and non-speech data in the audio data; the step of inputting the obtained audio data into the audio-to-text model to obtain the text content corresponding to the output audio data includes: inputting the speech data into the audio-to-text model to obtain the text content corresponding to the output speech data.
[0016] According to a third aspect of the embodiments of the present disclosure, there is provided a training device for an audio-to-text model, including: an audio sample acquisition module configured to acquire audio samples, where the audio samples include real audio samples and virtual audio samples, the real audio samples are samples obtained by collecting the real speech of a user, the virtual audio samples are samples obtained by converting known text into audio using a speech synthesis engine, and the audio samples correspond to text content labels; a sample input module configured to input the audio samples into the audio-to-text model to obtain predicted text content; a loss calculation module configured to calculate a loss based on the predicted text content and the text content labels; a parameter adjustment module configured to train the audio-to-text model by adjusting the parameters of the audio-to-text model according to the loss.
[0017] Optionally, the sample input module is configured to: input the virtual audio sample into the audio-to-text model to obtain the predicted virtual text content; input the real audio sample into the audio-to-text model to obtain the predicted real text content; the loss calculation module is configured to: calculate a first loss based on the predicted virtual text content and the virtual text content label corresponding to the virtual audio sample; calculate a second loss based on the predicted real text content and the real text content label corresponding to the real audio sample; the parameter adjustment module is configured to: train the audio-to-text model by adjusting the parameters of the audio-to-text model according to the first loss to obtain a first audio-to-text model; fine-tune the parameters of the first audio-to-text model according to the second loss to obtain a finally trained audio-to-text model.
[0018] Optionally, the training device further includes: a sample enhancement module configured to perform sample enhancement on the audio sample to obtain an enhanced audio sample; the sample input module is configured to: input the enhanced audio sample into the audio-to-text model to obtain the predicted text content.
[0019] Optionally, the sample enhancement includes at least one of the following: changing the speech rate of the audio sample, changing the pitch of the audio sample, changing the gain of the audio sample, adding background noise to the audio sample, performing truncation processing on the audio sample, and randomly adjusting the spectrum corresponding to the audio sample.
[0020] According to a fourth aspect of the embodiments of the present disclosure, there is provided an audio-to-text device, including: an audio data acquisition module configured to acquire audio data; an audio data input module configured to input the acquired audio data into an audio-to-text model to obtain the text content corresponding to the output audio data, where the audio-to-text model is trained according to the training method of the present disclosure.
[0021] Optionally, the audio-to-text device further includes: a blank audio region determination module configured to determine at least one blank audio region included in the audio data, where the blank audio region is an audio region with a corresponding volume lower than a preset volume threshold; a target blank audio region determination module configured to determine a target blank audio region among the at least one blank audio region with a corresponding duration greater than or equal to a preset duration threshold; a segmentation module configured to segment the audio data into multiple audio segments based on the target blank audio region, where the audio data between two adjacent target blank audio regions is segmented into one audio segment; the audio data input module is configured to: sequentially input the multiple audio segments into the audio-to-text model in chronological order to obtain the text content corresponding to each sequentially output audio segment.
[0022] Optionally, the audio data input module is configured to: obtain a first input at the current moment, where the first input at least includes a first audio segment located in an audio buffer, and the audio buffer is used to limit the processing length of the audio stream obtained in real time; input the first input into the audio-to-text model to obtain first text content; obtain a second input at the next moment of the current moment, where the second input at least includes a second audio segment located in the audio buffer; input the second input into the audio-to-text model to obtain second text content; compare the text content of the same audio segment included in the first text content with the text content of the same audio segment included in the second text content to obtain a comparison result, where the text content of the same audio segment is the text content corresponding to the overlapping audio segment of the first audio segment and the second audio segment; and output the text content of the same audio segment when the comparison result indicates consistency.
[0023] Optionally, the audio data input module is configured to: compare a partial text content located at a preset position included in the text content of the same audio segment included in the first text content with a partial text content located at the preset position included in the text content of the same audio segment included in the second text content to obtain the comparison result.
[0024] Optionally, the audio data input module is configured to: determine whether the text content of the same audio segment between the current moment and the previous moment includes a complete sentence; delete the complete sentence audio segment corresponding to the complete sentence included in the first audio segment when it is determined that the text content of the same audio segment includes the complete sentence; and use the remaining audio segment except the complete sentence audio segment in the first audio segment, the audio segment that is continuously rolled into the audio buffer during the process from the current moment to the next moment, and the complete sentence content as the second input.
[0025] Optionally, the audio-to-text device further includes: a full-sentence content adding module configured to add the full-sentence content to a full-sentence content storage space, where the full-sentence content storage space stores historical full-sentence contents obtained at various times before the current time and arranged in chronological order; the audio data input module is configured to: use the remaining audio segment, the audio segments that are real-time rolled into the audio buffer during the process from the current time to the next time, the full-sentence content, and the historical full-sentence contents as the second input.
[0026] Optionally, the audio-to-text device further includes: a voice data determination module configured to determine voice data and non-voice data in the audio data; the audio data input module is configured to: input the voice data into the audio-to-text model to obtain the text content corresponding to the output voice data.
[0027] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the training method of the audio-to-text model according to the present disclosure, or to implement the audio-to-text method according to the present disclosure.
[0028] According to a sixth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the training method of the audio-to-text model according to the present disclosure, or to execute the audio-to-text method according to the present disclosure.
[0029] According to a seventh aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the training method of the audio-to-text model according to the present disclosure, or implements the audio-to-text method according to the present disclosure.
[0030] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0031] In the present disclosure, when training an audio-to-text model, two different types of data sets, namely real audio samples and virtual audio samples, can be used for training. The training data volume is larger, the training data content is richer, the training coverage is more comprehensive and extensive, which can improve the training effect of the audio-to-text model, and further ensure the conversion accuracy of the audio-to-text model.
[0032] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an undue limitation of the present disclosure.
[0034] Figure 1 is a schematic diagram showing the voice recognition process in the related art;
[0035] Figure 2 is a flowchart showing a method for training an audio-to-text model according to an exemplary embodiment of the present disclosure;
[0036] Figure 3 is a flowchart showing an audio-to-text method according to an exemplary embodiment of the present disclosure;
[0037] Figure 4 is a schematic diagram showing at least one blank audio region included in audio data according to an exemplary embodiment of the present disclosure;
[0038] Figure 5 is a schematic diagram showing a first audio segment and a second audio segment located in an audio buffer according to an exemplary embodiment of the present disclosure;
[0039] Figure 6 is a schematic diagram showing the speech transcription process at the previous moment, the current moment, and the next moment according to an exemplary embodiment of the present disclosure;
[0040] Figure 7 is a block diagram showing a training device for an audio-to-text model according to an exemplary embodiment of the present disclosure;
[0041] Figure 8 is a block diagram showing an audio-to-text device according to an exemplary embodiment of the present disclosure;
[0042] Figure 9 is a block diagram showing an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0044] It should be noted that the terms "first", "second", etc. in the specification, claims and the above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0045] It should be noted here that "at least one of several items" in the present disclosure all represents three parallel situations, including "any one of the several items", "a combination of any multiple of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example, "performing at least one of step one and step two" means the following three parallel situations: (1) performing step one; (2) performing step two; (3) performing step one and step two.
[0046] Automatic speech recognition (ASR), also known as streaming speech recognition, is used to convert the user's speech into text content for output. Specifically, after receiving the speech input from the speaker, the ASR system can convert the speech signal into the corresponding text content for output by performing steps such as audio processing, speech feature extraction, and model matching. This technology can be applied to applications such as voice assistants, voice searches, and speech transcription, making the interaction between users and computers more natural and convenient. Streaming speech recognition specifically refers to the ability to process continuous speech input in real time and perform instant speech recognition and text output without interrupting the speech stream. Figure 1 is a schematic diagram showing the speech recognition process in the related art.
[0047] In the related art, when training a streaming speech recognition model, a single type of dataset is often used for training, and the training data volume is small and the training data content is not rich enough. This will result in a poor training effect of the streaming speech recognition model, thereby affecting the recognition accuracy of the streaming speech recognition model. In addition, the streaming speech recognition model in the related art has poor recognition performance on specific languages, that is, the accuracy of streaming speech recognition for different languages is poor.
[0048] To solve the above problems existing in the related art, the training method, audio-to-text method, device, electronic device, storage medium, and computer program product of the audio-to-text model provided by the present disclosure can use two different types of data sets, namely real audio samples and virtual audio samples, for training when training the audio-to-text model. The training data volume is larger, the training data content is richer, the training coverage is more comprehensive and extensive, which can improve the training effect of the audio-to-text model, and thus can ensure the conversion accuracy of the audio-to-text model.
[0049] Figure 2 is a flowchart showing a training method of an audio-to-text model according to an exemplary embodiment of the present disclosure.
[0050] Referring to Figure 2 , in step 201, an audio sample can be obtained, where the audio sample can include real audio samples and virtual audio samples. The "real audio sample" can be a sample obtained by collecting the real speech of a user. Specifically, real samples can be collected according to the needs of the sample user (for example, specific language, specific tone, specific scenario, etc.). The "virtual audio sample" can be a sample obtained by converting known text into audio using a speech synthesis engine. And the audio sample can correspond to a text content label.
[0051] It should be noted that in the present disclosure, the audio-to-text model can be the Whisper model, which is based on deep learning and can achieve high-accuracy speech recognition. Specifically, the working principle of the Whisper model can be: first convert the audio into a log-Mel spectrogram, and then pass it to the encoder. The trained encoder will try to predict the corresponding audio as a text caption and then display it to the user.
[0052] According to an exemplary embodiment of the present disclosure, the above "real audio sample" can use the Common Voice data set, which is a multi-language open-source speech academic data set and a large public corpus collected in a crowdsourcing manner. It mainly records voices and classifies the voices recorded by others using a microphone. The Common Voice data set contains speech data of all ages, different genders, and various accents, and can cover approximately 104 languages. Among them, each language can include a training set, a development set, and a test set required to establish a speech recognition model for that language.
[0053] In addition, the advantages of the open-source dataset, namely the CommonVoice dataset, are as follows: The speech data is sourced from news, film and television programs, actual recordings, TV shows, etc., and is manually annotated, with good speech diversity and conforming to the actual usage situation. Its disadvantages are: low data volume, and even the corpus of some languages is missing. Moreover, there is still a certain gap between the actual data distribution and the usage scenario.
[0054] The above-mentioned "virtual audio samples" can be audio samples generated by using text-to-speech (TTS) technology, that is, TTS can convert written text into natural speech for output. Specifically, after receiving text input, the TTS system can convert it into corresponding speech signals through a speech synthesis engine and output them to users. This technology can be applied to applications such as voice prompts, voice navigation, audiobooks, voice assistants, etc., enabling the computer system to interact with users in the form of natural speech.
[0055] In addition, the advantages of TTS are: large data volume and easy to obtain; its disadvantages are: the speech quality is limited by the limitations of the TTS model itself, resulting in insufficient diversity in scenarios such as timbre, scene, and pitch. This disclosure mainly combines the open-source dataset CommonVoice and the virtual audio samples generated by TTS for training the audio-to-text model, achieving the complementarity between different datasets. It should be noted that generating corpus through TTS and recognizing corpus through ASR are essentially a reciprocal process, so adversarial training can be carried out.
[0056] In step 202, the above audio samples can be input into the audio-to-text model to obtain the predicted text content, that is, the above speech sample data (numpy array) and the corresponding text content label (string) can be input into the audio-to-text model to obtain the predicted text content, and then the audio-to-text model can be trained based on the predicted text content and the corresponding text content label.
[0057] Specifically, first, the speech sample data can be uniformly adjusted to a sampling rate of 16 kHz. And for audio with a length less than 30s, it can be padded. Exemplarily, "0" can be filled into the audio; for audio with a length greater than 30s, it can be truncated. In this way, the length of all audio can be 30s. Next, all audio can be transformed into log mel spectrograms; for the text content label corresponding to the audio sample, the text can be tokenized through the whisper tokenizer and transformed into corresponding tokens. And for fine-tuning of various languages, the token id corresponding to that language needs to be added as a sequence prefix. In addition, the text can also be filled to the maximum batch length like the audio.
[0058] Then, the training problem for the audio-to-text model can be regarded as a multi-classification prediction problem, that is, the log Mel spectrogram of the speech can be input into the audio-to-text model, and then the audio-to-text model can output the predicted token id. Next, based on the predicted token id and the tokens obtained by tokenizing the text content label corresponding to the audio sample, the audio-to-text model can be trained.
[0059] According to an exemplary embodiment of the present disclosure, before inputting the audio sample into the audio-to-text model, sample augmentation can also be performed on the audio sample to obtain an augmented audio sample. Then, the augmented audio sample can be input into the audio-to-text model to obtain the predicted text content.
[0060] In this way, by performing sample augmentation on the audio sample, the number of audio samples can be increased, the sample content can be made more abundant, and the sample types can be made more comprehensive. Furthermore, the training samples can cover various speech scenarios as much as possible, thereby improving the training effect of the audio-to-text model and improving the prediction accuracy and accuracy of the audio-to-text model.
[0061] According to an exemplary embodiment of the present disclosure, the above "sample augmentation" may include at least one of the following items:
[0062] Changing the speech rate of the audio sample (e.g., 0.9x - 1.1x), changing the pitch of the audio sample (e.g., ±4 semitones), changing the gain of the audio sample (e.g., ±6 dB), adding background noise to the audio sample (e.g., musan noise dataset, Gaussian noise), performing truncation processing on the audio sample (e.g., horizontally shifting ±5 s left and right), randomly adjusting the spectrum corresponding to the audio sample (mask).
[0063] In step 203, the loss can be calculated based on the predicted text content and the text content label. Specifically, the cross-entropy loss can be calculated based on the predicted text content and the text content label. Exemplarily, the problem of the audio-to-text model predicting the text content corresponding to the audio sample can be regarded as a multi-classification problem. At this time, the output of the audio-to-text model can be a series of token IDs. The cross-entropy loss function in the present disclosure can be as follows:
[0064]
[0065] where N is the length of the token sequence output by the audio-to-text model, M is the total number of types of token IDs, I is the true ID of the i-th token, and p ij represents the probability that the audio-to-text model predicts the ID of the i-th token as j.
[0066] In step 204, the audio-to-text model can be trained by adjusting the parameters of the audio-to-text model according to the loss.
[0067] According to an exemplary embodiment of the present disclosure, a virtual audio sample can be input into the audio-to-text model to obtain predicted virtual text content; a real audio sample can also be input into the audio-to-text model to obtain predicted real text content.
[0068] Next, a first loss can be calculated based on the predicted virtual text content and the virtual text content label corresponding to the virtual audio sample; a second loss can also be calculated based on the predicted real text content and the real text content label corresponding to the real audio sample.
[0069] Then, the audio-to-text model can be trained by adjusting the parameters of the audio-to-text model according to the first loss to obtain a first audio-to-text model; next, the parameters of the first audio-to-text model can be fine-tuned according to the second loss to obtain a finally trained audio-to-text model. That is, coarse-grained training can be first performed on the TTS dataset, and then fine-grained fine-tuning can be performed on the open-source dataset.
[0070] In this way, in the present disclosure, corpus collection can include two aspects. Exemplarily, the open-source dataset: CommonVoice and the virtual audio samples generated by TTS can be combined and used to train the audio-to-text model. On the one hand, these two datasets are relatively easy to obtain, and on the other hand, complementarity between the datasets can be achieved. In addition, through this two-stage training method of first performing coarse-grained training on the TTS dataset and then performing fine-grained fine-tuning on the open-source dataset, the model training process is more hierarchical and targeted, thereby improving the training accuracy of the audio-to-text model and ensuring the prediction accuracy of the audio-to-text model.
[0071] It should be noted that in the present disclosure, the audio-to-text model can also be fine-tuned using corresponding language data for different languages to improve the performance of the model in the corresponding language. That is, in the present disclosure, corpus data from a specific language can be used to fine-tune the model to improve the speech recognition accuracy of the model in that specific language. That is, in the present disclosure, the parameters of the model can be further adjusted by using a dataset in a specific domain or specific task to improve its speech recognition performance in the corresponding domain or task.
[0072] Figure 3 is a flowchart showing an audio-to-text method according to an exemplary embodiment of the present disclosure.
[0073] Refer toFigure 3 In step 301, audio data can be obtained.
[0074] In step 302, the obtained audio data can be input into an audio-to-text model to obtain the text content corresponding to the output audio data, where the audio-to-text model can be trained according to the training method of the present disclosure.
[0075] According to an exemplary embodiment of the present disclosure, before inputting the obtained audio data into the audio-to-text model, at least one blank audio region included in the audio data can be determined, where the blank audio region can be an audio region corresponding to a volume lower than a preset volume threshold. Figure 4 is a schematic diagram showing at least one blank audio region included in the audio data according to an exemplary embodiment of the present disclosure. Refer to Figure 4 , the audio data can altogether include 9 blank audio regions, namely (1), (2), (3), (4), (5), (6), (7), (8), and (9).
[0076] Next, a target blank audio region corresponding to a duration greater than or equal to a preset duration threshold in at least one blank audio region can be determined. Exemplarily, refer to Figure 4 , a target blank audio region corresponding to a duration greater than or equal to the preset duration threshold can be determined among the above 9 blank audio regions. Suppose a total of 4 target blank audio regions are determined, namely: (3), (5), (6), and (9). In addition, in the present disclosure, the above "preset duration threshold" can be determined according to the median of the sentence intervals in the training corpus.
[0077] Then, based on the target blank audio region, the audio data can be segmented into multiple audio segments, where the audio data between two adjacent target blank audio regions can be segmented into one audio segment. Exemplarily, refer to Figure 4 , the above 4 target blank audio regions can altogether segment the audio data into 5 audio segments, namely: a, b, c, d, and e.
[0078] Next, in chronological order, the multiple audio segments can be sequentially input into the audio-to-text model to obtain the text content corresponding to each sequentially output audio segment. Exemplarily, refer to Figure 4 , in chronological order, the above 5 audio segments a, b, c, d, and e can be sequentially input into the audio-to-text model to obtain the text content corresponding to each sequentially output audio segment.
[0079] In the present disclosure, by determining a target blank audio region in the audio data where the corresponding volume is lower than a preset volume threshold and the duration exceeds a preset duration threshold, the pauses in the audio data can be roughly determined, that is, the audio data can be divided into multiple independent audio segments, and then the independent audio segments can be sequentially input into an audio-to-text model to achieve speech recognition for each audio segment. In this way, it is possible to prevent the audio-to-text model from processing overly long audio and avoid adverse interference of subsequent audio recognition results on previous audio recognition content.
[0080] According to an exemplary embodiment of the present disclosure, a first input at the current moment can be obtained, where the first input can at least include a first audio segment located in an audio buffer. The "audio buffer" is used to limit the processing length of the real-time acquired audio stream. Then, the above first input can be input into an audio-to-text model to obtain a first text content. Exemplarily, an audio buffer that limits the processing of at most 30 seconds of audio stream can be set. In this way, whenever a new audio stream is acquired, the new audio stream can be appended to the audio buffer, and then the Whisper model can be used to transcribe the audio in the audio buffer. Assume that the above "first audio segment" is the audio segment from the 0th second to the 30th second included in the audio data.
[0081] Next, a second input at the next moment of the current moment can be obtained, where the second input can at least include a second audio segment located in the audio buffer. Then, the second input can be input into the audio-to-text model to obtain a second text content. It should be noted that as the user continuously outputs speech and the transcription continues, the audio stream included in the above "audio buffer" changes in real time. Exemplarily, as described above, the first input at the current moment can be the audio segment from the 0th second to the 30th second located in the audio buffer; then at the next moment of the current moment, the audio segment located in the audio buffer may change to the audio segment from the 1st second to the 31st second in the audio data.
[0082] Figure 5 is a schematic diagram showing the first audio segment and the second audio segment located in the audio buffer according to an exemplary embodiment of the present disclosure. Refer to Figure 5 , at the current moment, the audio segment located in the audio buffer is the audio segment from the 0th second to the 30th second in the audio data; at the next moment of the current moment, the audio segment located in the audio buffer changes to the audio segment from the 1st second to the 31st second in the audio data.
[0083] Then, the text content of the same audio segment included in the first text content can be compared with the text content of the same audio segment included in the second text content to obtain a comparison result, where the "text content of the same audio segment" can be the text content corresponding to the overlapping audio segment of the first audio segment and the second audio segment. Specifically, the localagreement-2 strategy can be used to compare the text content corresponding to the overlapping audio segments included in different audio segments.
[0084] Exemplarily, referring to Figure 5 , for the audio segment from the 0th second to the 30th second in the audio buffer at the current moment, and the audio segment from the 1st second to the 31st second in the audio buffer at the next moment of the current moment, the "overlapping audio segment" in these two audio segments is the audio segment in the time interval from the 1st second to the 30th second in the audio data. At this time, the text content obtained by performing speech transcription on the audio segment in the time interval from the 1st second to the 30th second included in the first audio segment can be compared with the text content obtained by performing speech transcription on the audio segment in the time interval from the 1st second to the 30th second included in the second audio segment to obtain a comparison result.
[0085] Next, when the comparison result indicates consistency, the above text content of the same audio segment will be appended to "CONTEXT" and can be output to the user, that is, the above "overlapping audio segment": the text content corresponding to the audio segment in the time interval from the 1st second to the 30th second can be output. In addition, the above "overlapping audio segment" will be temporarily retained in the audio buffer. When the comparison result indicates inconsistency, it means that there is newly recognized text content. In this case, the newly recognized text content can be temporarily not output this time, but left to be displayed to the user after being verified and error-free later.
[0086] In this way, in the present disclosure, the local agreement-2 strategy can be used to compare the text content corresponding to the overlapping audio segments included in different audio segments. When the comparison result indicates consistency, it means that the text content corresponding to the overlapping audio segment has withstood the two-way verification. At this time, it can be determined that the text content corresponding to the overlapping audio segment belongs to the correct and error-free transcription result, and then the transcription result can be displayed to the user; when the comparison result indicates inconsistency, it means that the text content corresponding to the overlapping audio segment has not withstood the two-way verification, that is, the second speech recognition has probably recognized new text content. At this time, the newly recognized text content can be temporarily not output, but left to be displayed to the user after being verified and error-free later. In this way, by adopting the two-way verification method, the previous speech recognition result can be corrected, and then it can be ensured that the speech transcription result finally displayed to the user is a basically error-free transcription result, that is, the accuracy of audio-to-text conversion can be ensured.
[0087] According to an exemplary embodiment of the present disclosure, partial text content located at a preset position in the same audio segment text content included in the first text content can be compared with partial text content located at the preset position in the same audio segment text content included in the second text content to obtain a comparison result. Exemplarily, the above "partial text content located at the preset position" can be partial text content of a preset length located at the end of the entire text content.
[0088] Specifically, in the present disclosure, "n-gram" can be used to compare partial transcription results located at a preset position in two consecutive transcription results, that is, "n-gram" can be used to compare word sequences located at a preset position in two consecutive transcription results to determine whether these word sequences correspond to each other. It should be noted that "n" in n-gram can represent the number of words included in a sequence. Exemplarily, the value of "n" can be any integer between 1 and 5, inclusive.
[0089] In this way, in the present disclosure, when performing two-way verification of two consecutive transcription results, only partial transcription results located at the preset position of the transcription results need to be compared one by one. Compared with the method of comparing all transcription results, only comparing partial transcription results located at the specified position can not only play a role in two-way verification, but also save computing resources to a relatively high degree, that is, a better balance is achieved between verification accuracy and resource consumption.
[0090] According to an exemplary embodiment of the present disclosure, it can also be determined whether the same audio segment text content between the current moment and the previous moment contains a complete sentence content. The "complete sentence content" can refer to text content ending with a specific sentence-ending symbol, that is, it can be determined whether the above "CONTEXT" contains the sentence-ending symbol of the corresponding language. Exemplarily, for Chinese, the "specific sentence-ending symbol" can be, but is not limited to, "period 。", "question mark ?", "exclamation mark !", etc.; for English, the "specific sentence-ending symbol" can be, but is not limited to, "period.", "question mark?", "exclamation mark!", etc.
[0091] In the case where it is determined that the same audio segment text content contains complete sentence content, the complete sentence audio segment corresponding to the complete sentence content included in the first audio segment can be deleted. At the same time, "CONTEXT" can also be cleared. Moreover, the remaining audio segment in the first audio segment except the complete sentence audio segment, the audio segment that is continuously rolled into the audio buffer during the process from the current moment to the next moment, and the above complete sentence content can be used as the second input at the next moment of the current moment.
[0092] In this way, by using the entire sentence content recognized at the current moment as the input to the model for the next moment of the current moment again, the text content recognized subsequently can be made to align as much as possible with the text content recognized previously, that is, the consistency in semantics and style of the speech transcription results at each moment can be ensured.
[0093] According to an exemplary embodiment of the present disclosure, the above entire sentence content can also be added to the entire sentence content storage space, where the entire sentence content storage space can store historical entire sentence contents obtained at each moment before the current moment and arranged in chronological order.
[0094] It should be noted that the "entire sentence content storage space" can be "PROMPT", and this "PROMPT" can accommodate up to 200 characters at most. When the length of the characters included in "PROMPT" is about to exceed 200 characters, truncation processing can be performed on the characters included in "PROMPT". Exemplarily, truncation processing can be performed on the characters at the front end of "PROMPT", that is, the characters with a relatively early transcription order in "PROMPT" can be directly discarded, and then only the characters at the back end of "PROMPT" are retained, that is, only the characters with a relatively late transcription order in "PROMPT" are retained.
[0095] Next, the remaining audio fragments in the first audio fragment except for the entire sentence audio fragment, the audio fragments that are continuously rolled into the audio buffer in real time during the process from the current moment to the next moment, the entire sentence content obtained at the current moment, and the historical entire sentence contents included in "PROMPT" can be used as the input to the model for the next moment of the current moment.
[0096] In this way, in the present disclosure, when performing the transcription task each time, the entire sentence content transcribed at the current moment and the historical entire sentence contents included in "PROMPT" can be used as the input to the model for the next moment of the current moment again, and thus the text content recognized subsequently can be made to align as much as possible with the text content recognized previously, that is, the consistency in semantics and style of the speech transcription results at each moment can be further ensured.
[0097] Figure 6 is a schematic diagram showing the speech transcription process at the previous moment, the current moment, and the next moment according to an exemplary embodiment of the present disclosure. Refer to Figure 6, the black long strip area is the audio buffer, and the area in front of the audio buffer is the storage space "PROMPT" for the whole sentence content. The "PROMPT" input to the model at the current moment can be: "Thank you, Mr. President.". The text transcription result corresponding to the audio segment contained in the audio buffer at the current moment can be: "Today, I want to thank Mr. Brake for his great report. And".
[0098] In addition, Figure 6 The text content that has withstood the two-way verification is shown: "Today, I want to thank Mr. Brake for his great", and the newly transcribed text content: "report. And". At this time, the text content that has withstood the two-way verification: "Today, I want to thank Mr. Brake for his great" can be shown to the user; while the newly transcribed text content: "report. And" needs to be used together with the transcription result at the next moment of the current moment to perform two-way verification using the local agreement strategy.
[0099] According to an exemplary embodiment of the present disclosure, before inputting the obtained audio data into the audio-to-text model, the voice data and non-voice data in the audio data can also be determined. Then, only the voice data can be input into the audio-to-text model to obtain the text content corresponding to the output voice data.
[0100] It should be noted that Voice Activity Detection (VAD) technology is usually used in speech processing systems to distinguish speech signals from non-speech signals (such as silence or noise), so as to perform speech recognition, speech enhancement or other speech processing tasks more effectively. Therefore, in the present disclosure, VAD can be used to detect whether there is effective voice activity in the audio data, that is, to detect when the speaker starts speaking and when the speaker stops speaking, and then only the detected effective voice signal can be input into the audio-to-text model for speech transcription.
[0101] In this way, in the present disclosure, by using VAD to pre-distinguish voice data and non-voice data in advance, it is possible to achieve speech recognition using the audio-to-text model only when there is voice activity, avoid recognizing noise, reduce unnecessary computational effort, and thus reduce resource consumption and improve the speech recognition accuracy and efficiency of the model.
[0102] Figure 7 is a block diagram showing a training device of an audio-to-text model according to an exemplary embodiment of the present disclosure.
[0103] Referring to Figure 7 , the training device 700 of the audio-to-text model may include an audio sample acquisition module 701, a sample input module 702, a loss calculation module 703, and a parameter adjustment module 704.
[0104] The audio sample acquisition module 701 can acquire audio samples, where the audio samples can include real audio samples and virtual audio samples. The "real audio samples" can be samples obtained by collecting the real voices of users, and the "virtual audio samples" can be samples obtained by converting known text into audio using a speech synthesis engine. And the audio samples can correspond to text content labels.
[0105] The sample input module 702 can input the above audio samples into the audio-to-text model to obtain the predicted text content, that is, the above voice sample data (numpy array) and the corresponding text content labels (string) can be input into the audio-to-text model to obtain the predicted text content, and then the audio-to-text model can be trained based on the predicted text content and the corresponding text content labels.
[0106] According to an exemplary embodiment of the present disclosure, the above training device 700 may further include a sample enhancement module.
[0107] Before inputting the audio samples into the audio-to-text model, the sample enhancement module can also perform sample enhancement on the audio samples to obtain enhanced audio samples. Then, the sample input module 702 can input the enhanced audio samples into the audio-to-text model to obtain the predicted text content.
[0108] In this way, by performing sample enhancement on the audio samples, the number of audio samples can be increased, the sample content can be made more abundant, and the sample types can be made more comprehensive. Furthermore, the training samples can cover various speech scenarios as much as possible, thereby improving the training effect of the audio-to-text model and improving the prediction accuracy and accuracy of the audio-to-text model.
[0109] According to an exemplary embodiment of the present disclosure, the above "sample enhancement" may include at least one of the following items:
[0110] Changing the speech rate of the audio samples (e.g., 0.9x - 1.1x), changing the pitch of the audio samples (e.g., ±4 semitones), changing the gain of the audio samples (e.g., ±6 dB), adding background noise to the audio samples (e.g., musan noise dataset, Gaussian noise), performing truncation processing on the audio samples (e.g., horizontal shift of ±5 s left and right), randomly adjusting the spectrum corresponding to the audio samples (mask).
[0111] The loss calculation module 703 can calculate the loss based on the predicted text content and the text content tags. Specifically, the cross-entropy loss can be calculated based on the predicted text content and the text content tags.
[0112] The parameter adjustment module 704 can train the audio-to-text model by adjusting the parameters of the audio-to-text model according to the loss.
[0113] According to an exemplary embodiment of the present disclosure, the sample input module 702 can input virtual audio samples into the audio-to-text model to obtain predicted virtual text content; the sample input module 702 can also input real audio samples into the audio-to-text model to obtain predicted real text content.
[0114] Next, the loss calculation module 703 can calculate a first loss based on the predicted virtual text content and the virtual text content tags corresponding to the virtual audio samples; the loss calculation module 703 can also calculate a second loss based on the predicted real text content and the real text content tags corresponding to the real audio samples.
[0115] Then, the parameter adjustment module 704 can train the audio-to-text model by adjusting the parameters of the audio-to-text model according to the first loss to obtain a first audio-to-text model; next, the parameter adjustment module 704 can fine-tune the parameters of the first audio-to-text model according to the second loss to obtain a finally trained audio-to-text model. That is, coarse-grained training can be first performed on the TTS dataset, and then fine-grained fine-tuning can be performed on the open-source dataset.
[0116] In this way, in the present disclosure, corpus collection can include two aspects. Exemplarily, the open-source dataset: CommonVoice and the virtual audio samples generated by TTS can be combined and used to train the audio-to-text model. On the one hand, these two datasets are relatively easy to obtain, and on the other hand, the complementarity between the datasets can be achieved. In addition, through this two-stage training method of first performing coarse-grained training on the TTS dataset and then performing fine-grained fine-tuning on the open-source dataset, the model training process is more hierarchical and targeted, thereby improving the training accuracy of the audio-to-text model and ensuring the prediction accuracy of the audio-to-text model.
[0117] Figure 8 It is a block diagram showing an audio-to-text device according to an exemplary embodiment of the present disclosure.
[0118] Referring to Figure 8 This audio-to-text device 800 may include an audio data acquisition module 801 and an audio data input module 802.
[0119] The audio data acquisition module 801 can acquire audio data.
[0120] The audio data input module 802 can input the acquired audio data into the audio-to-text model to obtain the text content corresponding to the output audio data, where the audio-to-text model can be trained according to the training method of the present disclosure.
[0121] According to an exemplary embodiment of the present disclosure, the above audio-to-text device 800 may further include a blank audio region determination module, a target blank audio region determination module, and a segmentation module.
[0122] Before inputting the acquired audio data into the audio-to-text model, the blank audio region determination module may further determine at least one blank audio region included in the audio data, where the blank audio region may be an audio region corresponding to a volume lower than a preset volume threshold.
[0123] Next, the target blank audio region determination module may determine a target blank audio region in at least one blank audio region corresponding to a duration greater than or equal to a preset duration threshold. Additionally, in the present disclosure, the above "preset duration threshold" may be determined according to the median of the sentence intervals in the training corpus.
[0124] Then, the segmentation module may segment the audio data into multiple audio segments based on the target blank audio region, where the audio data between two adjacent target blank audio regions may be segmented into one audio segment.
[0125] Next, the audio data input module 802 may sequentially input the multiple audio segments into the audio-to-text model in chronological order to obtain the text content corresponding to each sequentially output audio segment.
[0126] In the present disclosure, by determining the target blank audio region in the audio data corresponding to a volume lower than the preset volume threshold and a duration exceeding the preset duration threshold, the pause in the audio data can be roughly determined, that is, the audio data can be divided into multiple independent audio paragraphs, and then the independent audio paragraphs can be sequentially input into the audio-to-text model to implement speech recognition for each audio paragraph. In this way, it is possible to prevent the audio-to-text model from processing overly long audio and the subsequent audio recognition results from having an adverse interference on the previous audio recognition content.
[0127] According to an exemplary embodiment of the present disclosure, the audio data input module 802 may obtain a first input at the current moment, where the first input may at least include a first audio segment located in the audio buffer. The "audio buffer" is used to limit the processing length of the real-time acquired audio stream. Then, the above first input may be input into the audio-to-text model to obtain the first text content. Exemplarily, an audio buffer that limits the processing of at most 30 seconds of audio stream may be set. In this way, whenever a new audio stream is acquired, the new audio stream may be appended to the audio buffer, and then the Whisper model may be used to transcribe the audio in the audio buffer.
[0128] Next, the audio data input module 802 may obtain a second input at the next moment of the current moment, where the second input may at least include a second audio segment located in the audio buffer. Then, the second input may be input into the audio-to-text model to obtain the second text content. It should be noted that as the user continuously outputs speech and the transcription continues, the audio stream contained in the above "audio buffer" changes in real time.
[0129] Then, the audio data input module 802 may compare the text content of the same audio segment included in the first text content with the text content of the same audio segment included in the second text content to obtain a comparison result, where the "text content of the same audio segment" may be the text content corresponding to the overlapping audio segment of the first audio segment and the second audio segment. Specifically, the local agreement-2 strategy may be adopted to compare the text content corresponding to the overlapping audio segments included in different audio segments.
[0130] Next, when the comparison result indicates consistency, the above text content of the same audio segment will be appended to "CONTEXT" and output to the user. In addition, the above "overlapping audio segment" will be temporarily retained in the audio buffer. When the comparison result indicates inconsistency, it means that there is newly recognized text content. In this case, the newly recognized text content may not be output this time, but will be displayed to the user after being verified correctly later.
[0131] In this way, in the present disclosure, the local agreement-2 strategy can be adopted to compare the text contents corresponding to the overlapping audio segments included in different audio segments. When the comparison result indicates consistency, it means that the text content corresponding to the overlapping audio segment has withstood the two-way verification. At this time, it can be determined that the text content corresponding to the overlapping audio segment belongs to the correct transcription result, and then the transcription result can be presented to the user. When the comparison result indicates inconsistency, it means that the text content corresponding to the overlapping audio segment has not withstood the two-way verification, that is, the second speech recognition has probably recognized new text content. At this time, the newly recognized text content can be temporarily not output, but can be left to be presented to the user after being verified correctly later. In this way, by adopting the two-way verification method, the previous speech recognition result can be corrected, and then the speech transcription result finally presented to the user can be ensured to be a basically error-free transcription result, that is, the accuracy of audio-to-text conversion can be ensured.
[0132] According to an exemplary embodiment of the present disclosure, the audio data input module 802 can compare a partial text content located at a preset position included in the same audio segment text content included in the first text content with a partial text content located at a preset position included in the same audio segment text content included in the second text content to obtain a comparison result. Exemplarily, the above-mentioned "partial text content located at a preset position" can be a partial text content of a preset length located at the end of the entire text content.
[0133] Specifically, in the present disclosure, "n-gram" can be used to compare the partial transcription results located at the preset position in the consecutive two transcription results, that is, "n-gram" can be used to compare the word sequences located at the preset position in the consecutive two transcription results to determine whether these word sequences correspond to the same. It should be noted that "n" in n-gram can represent the number of words included in a sequence. Exemplarily, the value of "n" can be, but is not limited to, any integer between 1 and 5.
[0134] In this way, in the present disclosure, when performing two-way verification of consecutive two transcription results, only the partial transcription results located at the preset position of the transcription results can be compared one by one. Compared with the method of comparing all the transcription results, only comparing the partial transcription results located at the specified position can not only play the role of two-way verification, but also save computing resources to a relatively high degree, that is, a better balance is achieved between verification accuracy and resource consumption.
[0135] According to an exemplary embodiment of the present disclosure, the audio data input module 802 may further determine whether the text content of the same audio segment between the current moment and the previous moment contains a complete sentence content. The "complete sentence content" may refer to the text content ending with a specific sentence-ending symbol, that is, it may be determined whether the above "CONTEXT" contains the sentence-ending symbol of the corresponding language. Exemplarily, for Chinese, the "specific sentence-ending symbol" may be, but is not limited to: "period 。", "question mark ?", "exclamation mark !", etc.; for English, the "specific sentence-ending symbol" may be, but is not limited to: "period.", "question mark?", "exclamation mark!", etc.
[0136] In the case where it is determined that the text content of the same audio segment contains a complete sentence content, the audio data input module 802 may delete the complete sentence audio segment corresponding to the complete sentence content included in the first audio segment. At the same time, it may also clear the "CONTEXT". In addition, the audio data input module 802 may also use the remaining audio segments in the first audio segment except the complete sentence audio segment, the audio segments that are continuously rolled into the audio buffer during the process from the current moment to the next moment, and the above complete sentence content as the second input for the next moment of the current moment.
[0137] In this way, by using the complete sentence content recognized at the current moment as the model input for the next moment of the current moment again, the text content recognized subsequently can be made to be as close as possible to the text content recognized previously, that is, the consistency of the speech-to-text results at each moment in terms of semantics and style can be ensured.
[0138] According to an exemplary embodiment of the present disclosure, the above audio-to-text device 800 may further include a complete sentence content adding module.
[0139] The complete sentence content adding module may further add the above complete sentence content to the complete sentence content storage space, where the complete sentence content storage space may store historical complete sentence contents obtained at each moment before the current moment and arranged in chronological order.
[0140] It should be noted that the "complete sentence content storage space" may be "PROMPT", and this "PROMPT" can accommodate up to 200 characters at most. In the case where the length of the characters included in "PROMPT" is about to exceed 200 characters, truncation processing may be performed on the characters included in "PROMPT". Exemplarily, truncation processing may be performed on the characters at the front end of "PROMPT", that is, the characters with a relatively early transcription order in "PROMPT" may be directly discarded, and then only the characters at the back end of "PROMPT" are retained, that is, only the characters with a relatively late transcription order in "PROMPT" are retained.
[0141] Next, the audio data input module 802 may use the remaining audio segments in the first audio segment except for the full-sentence audio segments, the audio segments that are rolled into the audio buffer in real time from the current moment to the next moment, the full-sentence content obtained at the current moment, and the historical full-sentence content included in "PROMPT" as the model input for the next moment of the current moment.
[0142] In this way, in the present disclosure, when performing a transcription task each time, the full-sentence content transcribed at the current moment and the historical full-sentence content included in "PROMPT" can be used again as the model input for the next moment of the current moment, so as to make the text content recognized subsequently as close as possible to the text content recognized previously, that is, the consistency of the speech transcription results at each moment in semantics and style can be further ensured.
[0143] According to an exemplary embodiment of the present disclosure, the above audio-to-text device 800 may further include a voice data determination module.
[0144] Before inputting the obtained audio data into the audio-to-text model, the voice data determination module may further determine the voice data and non-voice data in the audio data. Then, the audio data input module 802 may input only the voice data into the audio-to-text model to obtain the text content corresponding to the output voice data.
[0145] It should be noted that voice activity detection (VAD) technology is usually used in speech processing systems to distinguish between speech signals and non-speech signals (such as silence or noise), so as to perform speech recognition, speech enhancement, or other speech processing tasks more effectively. Therefore, in the present disclosure, VAD can be used to detect whether there is effective voice activity in the audio data, that is, to detect when the speaker starts speaking and when the speaker stops speaking, and then only the detected effective voice signal can be input into the audio-to-text model for speech transcription.
[0146] In this way, in the present disclosure, by using VAD to pre-distinguish voice data and non-voice data in advance, it is possible to realize that the audio-to-text model is used for speech recognition only when there is voice activity, avoid recognizing noise, reduce unnecessary computational complexity, and thus reduce resource consumption and improve the speech recognition accuracy and efficiency of the model.
[0147] Figure 9 is a block diagram showing an electronic device according to an exemplary embodiment of the present disclosure.
[0148] Refer to Figure 9, the electronic device 900 includes at least one memory 901 and at least one processor 902. Instructions are stored in the at least one memory 901. When the instructions are executed by the at least one processor 902, a method for training an audio-to-text model or an audio-to-text method according to an exemplary embodiment of the present disclosure is executed.
[0149] As an example, the electronic device 900 may be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instructions. Here, the electronic device 900 does not have to be a single electronic device, and may also be a collection of devices or circuits that can execute the above instructions (or instruction sets) individually or jointly. The electronic device 900 may also be a part of an integrated control system or system manager, or may be configured as a portable electronic device that can be interfaced with a local or remote (e.g., via wireless transmission).
[0150] In the electronic device 900, the processor 902 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0151] The processor 902 may run instructions or code stored in the memory 901. The memory 901 may also store data. The instructions and data may also be sent and received via a network interface device over a network, where the network interface device may employ any known transmission protocol.
[0152] The memory 901 may be integrated with the processor 902. For example, RAM or flash memory may be arranged within an integrated circuit microprocessor, etc. In addition, the memory 901 may include a separate device, such as an external disk drive, a storage array, or other storage devices that can be used by any database system. The memory 901 and the processor 902 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, etc., such that the processor 902 can read files stored in the memory.
[0153] In addition, the electronic device 900 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 900 may be connected to each other via a bus and / or a network.
[0154] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium may also be provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the above-mentioned training method or audio-to-text method of the audio-to-text model. Examples of such computer-readable storage media include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above-mentioned computer-readable storage medium may run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0155] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, including a computer program, which when executed by a processor implements the training method or audio-to-text method of the audio-to-text model according to the present disclosure.
[0156] According to the training method, audio-to-text method, device, electronic device, storage medium, and computer program product of the audio-to-text model according to the present disclosure, when training the audio-to-text model, two different types of data sets, namely real audio samples and virtual audio samples, can be used for training. The training data volume is larger, the training data content is richer, the training coverage is more comprehensive and extensive, which can improve the training effect of the audio-to-text model, and thus can ensure the conversion accuracy of the audio-to-text model.
[0157] According to an exemplary embodiment of the present disclosure, by performing sample enhancement on audio samples, the number of audio samples can be increased, the sample content can be made richer, and the sample types can be made more comprehensive. Furthermore, the training samples can cover various speech scenarios as much as possible, thereby improving the training effect of the audio-to-text model and enhancing the prediction accuracy and precision of the audio-to-text model.
[0158] According to an exemplary embodiment of the present disclosure, through a two-stage training method of first performing coarse-grained training on a TTS dataset and then performing fine-grained fine-tuning on an open-source dataset, the model training process becomes more hierarchical and targeted. Furthermore, the training accuracy of the audio-to-text model can be improved, thus ensuring the prediction accuracy of the audio-to-text model.
[0159] According to an exemplary embodiment of the present disclosure, by determining a target blank audio region in the audio data where the corresponding volume is lower than a preset volume threshold and the continuous duration exceeds a preset duration threshold, the pauses in the audio data can be roughly determined. That is, the audio data can be divided into multiple independent audio segments, and then the independent audio segments can be sequentially input into the audio-to-text model to achieve speech recognition for each audio segment. In this way, it is possible to prevent the audio-to-text model from processing overly long audio and avoid adverse interference of subsequent audio recognition results on the previous audio recognition content.
[0160] According to an exemplary embodiment of the present disclosure, by adopting a two-way verification method, it is possible to correct the previous speech recognition results. Furthermore, it can be ensured that the speech transcription result finally presented to the user is a basically error-free transcription result, that is, the accuracy of audio-to-text can be guaranteed.
[0161] According to an exemplary embodiment of the present disclosure, when performing two-way verification of consecutive transcription results, only the partial transcription results at the preset positions in the transcription results need to be compared one by one. Compared with the method of comparing all transcription results, only comparing the partial transcription results at the specified positions can not only achieve the effect of two-way verification but also save computational resources to a high degree, that is, a better balance is achieved between verification accuracy and resource consumption.
[0162] According to an exemplary embodiment of the present disclosure, by using the entire sentence content recognized at the current moment as the model input for the next moment of the current moment again, the text content recognized subsequently can be made to align as much as possible with the text content recognized previously. That is, the semantic and stylistic consistency of the speech transcription results at each moment can be ensured.
[0163] According to an exemplary embodiment of the present disclosure, when performing a transcription task each time, the entire sentence content transcribed at the current moment and the historical entire sentence content included in "PROMPT" can be used again as the model input for the next moment at the current moment. Furthermore, the text content recognized subsequently can be made to align as much as possible with the text content recognized previously, that is, the consistency in semantics and style of the speech transcription results at each moment can be further ensured.
[0164] According to an exemplary embodiment of the present disclosure, by using VAD to pre-distinguish speech data and non-speech data, it is possible to achieve speech recognition using an audio-to-text model only when there is speech activity, avoid recognizing noise, reduce unnecessary computational effort, and thus reduce resource consumption and improve the speech recognition accuracy and efficiency of the model.
[0165] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0166] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A training method for an audio-to-text model, characterized in that: include: Acquire an audio sample, wherein the audio sample includes a real audio sample and a virtual audio sample, the real audio sample is a sample obtained by collecting the user's real voice, the virtual audio sample is a sample obtained by converting a known text into audio using a speech synthesis engine, and the audio sample corresponds to a text content label; Inputting the audio sample into the audio-to-text model to obtain predicted text content; Calculating a loss based on the predicted text content and the text content label; The audio-to-text model is trained by adjusting parameters of the audio-to-text model according to the loss.
2. The training method according to claim 1, characterized in that: The step of inputting the audio sample into the audio-to-text model to obtain predicted text content includes: Inputting the virtual audio sample into the audio-to-text model to obtain predicted virtual text content; Inputting the real audio sample into the audio-to-text model to obtain predicted real text content; The step of calculating the loss based on the predicted text content and the text content label includes: Calculating a first loss based on the predicted virtual text content and the virtual text content label corresponding to the virtual audio sample; Calculating a second loss based on the predicted real text content and the real text content label corresponding to the real audio sample; The step of adjusting the parameters of the audio-to-text model according to the loss to train the audio-to-text model comprises: By adjusting the parameters of the audio-to-text model according to the first loss, the audio-to-text model is trained to obtain a first audio-to-text model; The parameters of the first audio-to-text model are fine-tuned according to the second loss to obtain a finally trained audio-to-text model.
3. An audio-to-text method, characterized in that: include: Get audio data; The acquired audio data is input into an audio-to-text model to obtain text content corresponding to the output audio data, wherein the audio-to-text model is trained according to the training method according to any one of claims 1 to 2.
4. The audio-to-text method according to claim 3, wherein: Before inputting the acquired audio data into the audio-to-text model to obtain the output text content corresponding to the audio data, the method further includes: Determine at least one blank audio region included in the audio data, wherein the blank audio region is an audio region whose corresponding volume is lower than a preset volume threshold; Determine a target blank audio region in the at least one blank audio region whose corresponding duration is greater than or equal to a preset duration threshold; Based on the target blank audio area, the audio data is divided into a plurality of audio segments, wherein the audio data between two adjacent target blank audio areas is divided into one audio segment; The step of inputting the acquired audio data into an audio-to-text model to obtain text content corresponding to the outputted audio data includes: The multiple audio clips are sequentially input into the audio-to-text model in chronological order to obtain text contents corresponding to the audio clips output sequentially.
5. The audio-to-text method according to claim 3, wherein: The step of inputting the acquired audio data into the audio-to-text model includes: Obtaining a first input at a current moment, wherein the first input at least includes a first audio segment in an audio buffer, and the audio buffer is used to limit the processing length of the audio stream obtained in real time; inputting the first input into the audio-to-text model to obtain a first text content; Obtaining a second input at a time point next to the current time point, wherein the second input at least includes a second audio segment located in the audio buffer; inputting the second input into the audio-to-text model to obtain second text content; Comparing the text content of the same audio segment contained in the first text content with the text content of the same audio segment contained in the second text content to obtain a comparison result, wherein the text content of the same audio segment is the text content corresponding to the overlapping audio segment of the first audio segment and the second audio segment; When the comparison result indicates consistency, the text content of the same audio segment is output.
6. A training device for an audio-to-text model, characterized in that: include: An audio sample acquisition module is configured to acquire audio samples, wherein the audio samples include real audio samples and virtual audio samples, the real audio samples are samples acquired by collecting the real voice of the user, the virtual audio samples are samples acquired by converting known text into audio using a speech synthesis engine, and the audio samples correspond to text content labels; A sample input module, configured to input the audio sample into the audio-to-text model to obtain predicted text content; A loss calculation module, configured to calculate a loss based on the predicted text content and the text content label; A parameter adjustment module is configured to train the audio-to-text model by adjusting the parameters of the audio-to-text model according to the loss.
7. An audio-to-text device, characterized in that: include: An audio data acquisition module, configured to acquire audio data; The audio data input module is configured to input the acquired audio data into an audio-to-text model to obtain text content corresponding to the output audio data, wherein the audio-to-text model is trained according to the training method according to any one of claims 1 to 2.
8. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the training method of the audio-to-text model as described in any one of claims 1 to 2, or to implement the audio-to-text method as described in any one of claims 3 to 5.
9. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the training method of the audio-to-text model as described in any one of claims 1 to 2, or to execute the audio-to-text method as described in any one of claims 3 to 5.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the training method of the audio-to-text model according to any one of claims 1 to 2, or implements the audio-to-text method according to any one of claims 3 to 5.