Method, device and storage medium for punctuation restoration in speech recognition
Through the UniPunc framework, the multimodal punctuation model trained with audio and without audio texts is solved, and the accuracy and multilingual adaptability of punctuation marks are improved.
Patent Information
- Application Number
- CN202111335102.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-11
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-11-11
AI Technical Summary
The text output by existing speech recognition systems lacks punctuation, resulting in dyslexia and degradation of downstream tasks. The conventional multimodal method is poor in audio-free text scenarios and scarce training data.
The UniPunc framework is adopted to build a multimodal punctuation model using audio and audio-free text, and optimize the punctuation recovery model through collaborative training of virtual audio samples and multilingual languages.
It improves the accuracy of punctuation mark addition and model performance in multilingual scenarios, solves the problems of modal missing and data scarcity, and enhances the punctuation recovery effect.
Smart Images

Figure CN114120975B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to speech recognition, including punctuation restoration in speech recognition. Background Art
[0002] Automatic speech recognition (ASR) is a technology that converts human speech into text, has a wide range of applications, and can serve as an upstream component for multiple tasks, such as voice assistants and speech translation, etc. The existing commercial speech recognition systems often output text without punctuation, which may lead to misunderstandings and affect the performance of downstream tasks such as machine translation and information extraction. Specifically, on the one hand, text without punctuation is difficult to read, has poor readability, unclear sentence breaks and ambiguity, and on the other hand, downstream tasks such as machine translation and information extraction assume that the input is punctuated, and text without punctuation will lead to degradation of the performance of downstream tasks. Summary of the Invention
[0003] This Summary of the Invention section is provided to introduce concepts in a brief form, which will be described in detail in the Detailed Implementation section later. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.
[0004] According to some embodiments of the present disclosure, there is provided a method for training a model for speech recognition punctuation restoration, including the following steps: obtaining a text sample for model training and a corresponding audio sample, wherein for a text sample obtained from text without audio, its corresponding audio sample is a virtual sample; and training a model for speech recognition punctuation restoration based on the obtained text sample for model training and the corresponding audio sample.
[0005] According to some other embodiments of the present disclosure, there is provided a training device for a model for speech recognition punctuation restoration, including: a sample acquisition unit configured to obtain a text sample for model training and a corresponding audio sample, wherein for a text sample obtained from text without audio, its corresponding audio sample is a virtual sample; and a training unit configured to train a model for speech recognition punctuation restoration based on the obtained text sample for model training and the corresponding audio sample.
[0006] According to some embodiments of the present disclosure, there is provided an electronic device, including: a memory; and a processor coupled to the memory, the processor being configured to execute the method of any one of the embodiments described in the present disclosure based on instructions stored in the memory.
[0007] According to some embodiments of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program, which when executed by a processor causes the method of any one of the embodiments described in the present disclosure to be implemented.
[0008] According to some embodiments of the present disclosure, there is provided a computer program product including instructions, which when executed by a processor causes the method of any one of the embodiments described in the present disclosure to be implemented.
[0009] Other features, aspects, and advantages of the present disclosure will become clear through the following detailed description of the exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Preferred embodiments of the present disclosure will be described below with reference to the accompanying drawings. The accompanying drawings described herein are used to provide a further understanding of the present disclosure, and together with the following specific description, they are included in this specification and form a part of this specification for explaining the present disclosure. It should be understood that the accompanying drawings in the following description only relate to some embodiments of the present disclosure and do not constitute a limitation to the present disclosure. In the drawings:
[0011] Figures 1A to 1C A flowchart showing a training method of a model for speech recognition punctuation restoration according to some embodiments of the present disclosure.
[0012] Figure 2 A flowchart showing a speech recognition punctuation restoration method according to some embodiments of the present disclosure.
[0013] Figure 3 An exemplary implementation of model training for speech recognition punctuation restoration according to some embodiments of the present disclosure.
[0014] Figure 4A A block diagram showing a training apparatus of a model for speech recognition punctuation restoration according to some embodiments of the present disclosure, and Figure 4B A block diagram showing a speech recognition punctuation restoration apparatus according to some embodiments of the present disclosure.
[0015] Figure 5 A block diagram showing some embodiments of an electronic device of the present disclosure.
[0016] Figure 6 A block diagram showing other embodiments of an electronic device of the present disclosure.
[0017] It should be understood that for the sake of description, the sizes of the various parts shown in the accompanying drawings are not necessarily drawn according to actual proportional relationships. The same or similar reference numerals are used in the various drawings to represent the same or similar components. Therefore, once an item is defined in one drawing, it may not be further discussed in subsequent drawings. Detailed implementation manners
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present disclosure in conjunction with the accompanying drawings in the embodiments of the present disclosure. However, it is obvious that the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The following description of the embodiments is actually only illustrative and in no way limits the present disclosure or its application or use. It should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.
[0019] It should be understood that the steps recorded in the method implementation manners of the present disclosure can be executed in different orders and / or executed in parallel. In addition, the method implementation manners may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard. Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments should be construed as merely exemplary and do not limit the scope of the present disclosure.
[0020] The term "comprising" and its variants used in the present disclosure mean open terms that include at least the subsequent elements / features, but do not exclude other elements / features, that is, "including but not limited to". In addition, the term "containing" and its variants used in the present disclosure mean open terms that include at least the subsequent elements / features, but do not exclude other elements / features, that is, "containing but not limited to". Therefore, "comprising" and "containing" are synonymous. The term "based on" means "at least partially based on".
[0021] Throughout the specification, the terms "one embodiment", "some embodiments", or "embodiments" mean that specific features, structures, or characteristics described in connection with the embodiments are included in at least one embodiment of the present invention. For example, the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Moreover, the occurrences of the phrases "in one embodiment", "in some embodiments", or "in embodiments" throughout the specification do not necessarily all refer to the same embodiment, but may also refer to the same embodiment.
[0022] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order of the functions performed by these devices, modules, or units or their interdependent relationships. Unless otherwise specified, the concepts such as "first" and "second" are not intended to imply that the objects so described must be in a given order in terms of time, space, ranking, or any other way.
[0023] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless clearly specified otherwise in the context, it should be understood as "one or more".
[0024] The names of the data, messages or information exchanged in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these data, messages or information.
[0025] To solve the problem of lack of punctuation in the output text of speech recognition, a punctuation restoration task is proposed and a corresponding punctuation model is designed. The punctuation restoration task aims to correctly add punctuation marks to the output results of an automatic speech recognition system, such as the output text sentence, and in implementation, relevant punctuation models are adopted for punctuation restoration. Conventional punctuation models either only use text or require the audio corresponding to the text, but are often restricted by real scenarios.
[0026] On the one hand, conventional punctuation models only perform punctuation restoration based on text information. However, the single text information is ambiguous, which easily causes some problems. A sentence may have different meanings when different punctuation marks are added. For example, the meaning of "I don’t want anymore kids" is significantly different from "I don’t wantanymore,kids", which shows the importance of the comma, and inappropriate addition will lead to obvious ambiguity or errors. In addition, due to the lack of knowledge of the speaker's tone, it may be difficult for the model to determine whether a sentence should end with a full stop or a question mark.
[0027] On the other hand, considering the rich information in speech audio, such as pauses and intonations, the sound signal can help reduce the ambiguity of the punctuation model. Therefore, some multi-modal punctuation models are proposed, which extract acoustic features from speech audio and fuse acoustic and lexical features by addition / concatenation to facilitate the addition of punctuation marks. Here, the modalities can respectively involve text and audio. However, conventional multi-modal methods face the problem of modality missing in practical applications. First, due to storage limitations or privacy policies, sometimes the corresponding audio cannot be accessed, and previous multi-modal methods cannot perform punctuation restoration on text sentences without audio; second, the cost of manually annotated text audio is very high and difficult to obtain, which results in a scarcer training set for these multi-modal models and makes it difficult to obtain a good multi-modal model.
[0028] In view of this, the present disclosure proposes an improved solution that can use both audio-containing text and audio-free text for punctuation model training. In particular, noting that audio-free text is easily obtainable, the solution of the present disclosure can make full use of audio-free text to construct a large training set for punctuation model training, and use limited audio to provide further optimized training, thereby improving the accuracy of the punctuation model and providing a punctuation model that can use optional audio to enhance the punctuation restoration effect of text.
[0029] In particular, the present disclosure proposes a so-called unified multi-modal punctuation framework, which can be referred to as UniPunc, and can use both audio-containing text and audio-free text for punctuation model training. Specifically, for audio-containing text, the corresponding text and audio pairs can be obtained, and for audio-free text, its corresponding audio can be fabricated by constructing virtual content, and the corresponding text and audio pairs can also be constructed. In this way, the punctuation model training can be carried out based on the text and audio pairs obtained from both audio-containing text and audio-free text. Thus, a large amount of easily obtainable audio-free text / corpus can be used as the training set, and voice input can be used to reduce the ambiguity in punctuation symbols, so that the obtained punctuation model has improved accuracy.
[0030] In addition, the solution of the present disclosure can further improve the training and application of punctuation models in multi-language scenarios. Conventional multi-modal methods can only process single-language inputs, and data for some languages is difficult to obtain and very scarce. For example, it is difficult to obtain a large amount of high-quality text and audio data for many minority languages, which cannot meet the requirements of model training, and direct training will result in very poor model performance. In view of this, the present disclosure uses a multi-language collaborative training method to solve the problem of insufficient data for minority languages. In particular, the present disclosure can simultaneously use the audio and text of multiple languages for model training, so that the data of commonly used languages that are easily obtainable can be used to enhance the performance of minority languages, thereby obtaining an improved multi-language punctuation model.
[0031] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. In addition, in one or more embodiments, specific features, structures or characteristics can be combined in any suitable manner that will be clear to those of ordinary skill in the art from the present disclosure.
[0032] Figure 1AA training method for a model for speech recognition punctuation restoration according to some embodiments of the present disclosure is shown. In method 100, in step S101, text samples and corresponding audio samples for model training are obtained, where for text samples obtained from text without audio, the corresponding audio samples are virtual samples; and in step S102, a model for speech recognition punctuation restoration is trained based on the obtained text samples and corresponding audio samples for model training.
[0033] According to some embodiments of the present disclosure, the text samples and corresponding audio samples for model training substantially correspond to a set of paired text and audio samples, where for each text sample, its corresponding audio sample is obtained, which is either the corresponding audio content of the text sample content or a virtual audio sample for the audio of a fictional text sample.
[0034] According to some embodiments of the present disclosure, the text samples and corresponding audio samples for model training are obtained from both text with audio and text without audio, for example, derived from an initial set containing both text with audio and text without audio. Among them, paired text samples and audio samples are obtained from text with audio. For example, various appropriate methods can be used for conversion / extraction / splitting, etc. to obtain them, which will not be described in detail here. In addition, from text without audio, text samples (such text samples can be called single text samples) are obtained and virtual samples are obtained as the audio samples corresponding to the text samples. In this way, on the one hand, a large amount of pure text can be utilized without audio; on the other hand, acoustic features can be effectively utilized when audio exists. Thus, a training set can be constructed by using easily obtainable large amounts of text without audio / single text samples. In particular, for a large number of single text samples, text and audio sample pairs can still be constructed by using virtual samples, which helps to optimize the model training effect.
[0035] In some embodiments, the virtual sample can be a preset sample for the audio of a fictional text sample. The virtual sample can take various appropriate forms. In particular, it can be a vector with preset content, such as a vector with a fixed length and fixed content, preferably a vector of all zeros. Of course, it can also be other appropriate forms. In some embodiments, for all single text samples used for model training, virtual samples can be constructed for all of them, and their virtual samples can be the same or different from each other.
[0036] In some embodiments, an initial set containing both audio texts and non-audio texts can be obtained in any suitable manner. For example, they can be obtained from a conventional database, such as an existing training database in a storage device; or can be acquired using a suitable type of acquisition device. For example, a microphone can be used to obtain audio data, and a video device or a suitable text input device can be used to input text data; or it can be text input by a user and / or audio manually annotated. In particular, audio texts and non-audio texts can be stored in any suitable manner for obtaining text samples and audio samples. For example, in some embodiments, these texts are provided with indication information as to whether the associated text has audio, such as an indicator in binary form, etc., so that the model training device can determine whether to generate virtual samples for a text sample based on this indication information. For example, virtual samples are generated only when the indication information indicates that the text sample has no audio.
[0037] According to some embodiments of the present disclosure, training a punctuation model for speech recognition punctuation restoration may include generating a multimodal mixed representation based on the obtained text samples and audio samples for model training, as in Figure 1B step S1021 in, and performing model training based on the generated multimodal mixed representation, as in Figure 1B step S1022 in. In this way, the obtained multimodal mixed representation can contain information on text and audio obtained from a training sample set effectively created using audio texts and non-audio texts, and such a multimodal mixed representation can well support model training. In some embodiments, for each pair of text sample and audio sample for model training, a corresponding multimodal mixed representation is obtained, whereby multimodal mixed representations of all text samples and audio samples can be obtained.
[0038] According to some embodiments of the present disclosure, a multimodal mixed representation can be obtained by respectively performing processing based on text samples and audio samples, and combining the processing results based on text samples and the processing results based on audio samples. In some embodiments, a mixed representation is generated based on the sum of the processing results of text samples and audio samples. It should be noted that a mixed representation can also be generated by performing other suitable processing on the processing results of text samples and audio samples, such as weighted sum or other suitable mathematical operations, etc., which will not be described in detail here.
[0039] According to some embodiments of the present disclosure, the processing performed based on text samples and audio samples may be attention-based processing. Thus, the multimodal hybrid representation may be a multimodal hybrid representation obtained by performing attention-based processing based on the acquired text samples and audio samples for model training. In particular, attention-based processing is performed based on the text samples and audio samples respectively, and the processing results of the text samples and the processing results of the audio samples are combined to obtain the multimodal hybrid representation.
[0040] According to some embodiments of the present disclosure, the attention-based processing applied to the text samples and audio samples may be any suitable processing, preferably different from each other. In some embodiments, a self-attention mechanism is used to process the text samples. In some embodiments, a cross-attention mechanism is used to process the audio samples. The self-attention mechanism and the cross-attention mechanism may adopt various suitable architectures / algorithms / models, etc., which will not be described in detail here. Thus, the multimodal hybrid representation can be generated by combining the self-attention processing results of the text samples and the cross-attention processing results of the audio samples.
[0041] According to some embodiments of the present disclosure, the text information related to the text samples is also considered when performing attention-based processing on the audio samples for model training. That is, in the processing of the audio samples, the text information related to the text samples corresponding to the audio samples will also be used as input parameters to perform the processing. In some embodiments, the text information related to the text samples may be features converted / extracted from the text samples or the processing results obtained by performing attention-based processing, such as self-attention processing, on the text samples.
[0042] According to some embodiments of the present disclosure, training a model for speech recognition punctuation restoration based on the text samples for model training and the corresponding audio samples further includes performing punctuation model training based on the intermediate feature values / sequences converted from the text samples and audio samples. In some embodiments, the text samples and audio samples may be respectively converted to obtain intermediate feature values / sequences, and then a multimodal hybrid representation is generated based on the intermediate feature values / sequences converted from the text samples and audio samples for punctuation model training. Preferably, attention-based processing is performed on the intermediate feature values / sequences converted from the text samples and audio samples to generate a multimodal hybrid representation, and punctuation model training is performed based on the multimodal hybrid, as Figure 1C shown. The attention-based processing performed on the intermediate feature values / sequences may be performed as described above and will not be described in detail here.
[0043] In some embodiments, the intermediate eigenvalue / sequence can be in any suitable form and can be obtained through corresponding operations.
[0044] According to some embodiments of the present disclosure, a lexical embedding sequence can be converted from a text sample as the intermediate eigenvalue / sequence. In some embodiments, lexical encoding processing can be performed on the text sample to obtain the lexical embedding sequence. It should be noted that the encoding of the text sample can be processed in various suitable ways, such as a pre-trained lexical encoder or various encoding methods known in the art, which will not be described in detail here. In some embodiments, self-attention processing can be performed on the lexical embedding sequence. In some embodiments, the lexical embedding sequence or the processing result of the lexical embedding sequence after self-attention processing can also be used as the text information related to the text sample in the attention-based processing of the audio sample, such as being used as an input parameter in the cross-attention processing of the audio sample.
[0045] According to some embodiments of the present disclosure, the audio sample can be processed to obtain acoustic embedding content as the intermediate eigenvalue / sequence. In some embodiments, the processing of the audio sample can be implemented in various suitable ways. In some embodiments, the processing of the audio sample can include encoding and / or downsampling of the audio sample. In some embodiments, corresponding processing can be performed on the corresponding audio sample according to whether the text sample has audio. For example, for the audio sample obtained from the text with audio, encoding and / or downsampling can be performed on it to generate the acoustic embedding content, while for the virtual sample corresponding to the non-audio text, it can be directly used as the acoustic embedding content. In implementation, the indication information about whether the text sample has audio, such as a binary indicator, etc., can be input into the model training device together with the text sample and the audio sample, so that the model training device can perform corresponding processing on the audio sample corresponding to the text sample according to this indication information for model training. In some embodiments, attention-based processing, such as cross-attention processing, can be performed on the acoustic embedding content.
[0046] According to some embodiments of the present disclosure, training a model for speech recognition punctuation restoration can include performing punctuation learning / prediction based on the obtained text sample and audio sample for model training, so as to perform punctuation model training. In some embodiments, punctuation learning / prediction can be performed based on the multi-modal hybrid representation described above, so as to perform punctuation model training. In some embodiments, a classifier can be used to perform punctuation learning / prediction. The classifier can be a classifier in various suitable forms, such as a linear classifier (Linear+Softmax), and of course it can also be any other suitable form, which will not be described in detail here.
[0047] In this way, the solution of the present disclosure can use single text data and paired text and audio data for training simultaneously to solve the problem of insufficient paired text and audio data. In particular, for single text data without corresponding audio content, virtual samples are used to fabricate audio as its corresponding audio data for model training, so that a large amount of pure text data can be used for training, and audio data can also be effectively used to enhance the model training effect.
[0048] According to some embodiments of the present disclosure, the text and audio samples for training the punctuation model may include multilingual text and audio samples. In some embodiments, the multilingual text and audio samples for model training can be obtained as described above. In particular, for the text samples in each language, the corresponding audio samples can be obtained, which are either samples in the audio form of the text sample content or virtual samples. Thus, an improved punctuation model can be trained, which can more accurately perform punctuation restoration in multilingual scenarios.
[0049] As an example, in the multilingual case, there is often a problem of insufficient data for minority languages. It is difficult to obtain a large amount of high-quality text data for many minority languages. Directly using a small amount of minority language data for training will result in very poor model performance. The method of simultaneous multilingual training can use the data of commonly used languages that are easily obtained to enhance the performance of minority languages. In this way, a sample library for multilingual model training can be effectively constructed, and a more accurate punctuation model suitable for multiple languages can be obtained by using multilingual samples for simultaneous training.
[0050] According to some embodiments of the present disclosure, multilingual text and audio samples for model training can be equalized to further optimize the samples for multilingual model training. In some embodiments, by expanding text samples and / or audio samples of languages with low proportions in text and audio samples, such as language samples that are not easily obtained, minority language samples, and low-resource / low-data languages, the proportions of text and / or audio of each language used for model training can be balanced. In particular, the proportions of text samples and / or audio samples of languages with low proportions are increased. In some embodiments, various appropriate methods can be used for data expansion. For example, operations such as duplication and interpolation can be performed on text samples and / or audio samples of languages with low proportions. In some embodiments, temperature sampling or a similar sampling algorithm can also be used to process multilingual text samples and / or audio samples to achieve equalization. As an example, due to the scarcity of minority language data, the temperature sampling method is used to increase the proportion of minority language data and reduce the proportion of high-resource language data according to the data volume of different language data. This makes the proportions of various language data in the total multilingual data as balanced as possible, which can further optimize the model training effect. For example, the punctuation restoration of the model trained for various language texts is improved.
[0051] The following will refer to the attached Figure 3 to describe in detail an exemplary implementation of punctuation model training according to some embodiments of the present disclosure. The solution of the present disclosure receives text input and optional audio input to perform a punctuation restoration task to add punctuation to the text. The text input and optional audio input here can be the text samples and audio samples for training described above, or the multilingual text samples and audio samples described above.
[0052] The solution of the present disclosure mainly includes three stages of operations: the first stage processes the text input and audio input respectively, the second stage performs attention-based processing on the processed text input and audio input respectively to integrate the information of the audio into the representation of the text to obtain a multimodal mixed representation, and the third stage performs punctuation prediction based on the multimodal mixed representation to train the model. The training process of the present invention uses the classical SGD algorithm and the Adam optimizer to optimize the Cross-Entropy loss function.
[0053] Punctuation restoration is usually modeled as a sequence tagging task. Generally, a multimodal punctuation corpus is a set of sentence-audio-punctuation triples, denoted as S = {x, a, y}, where x is an unpunctuated sentence of length T, and a is the corresponding speech audio. The output of the model should be the predicted sequence of added punctuation y given x and a. Due to the nature of the sequence tagging task, the length of the sequence of punctuation y is the same as that of the unpunctuated sentence x, i.e., |x| = |y|. The punctuation can be any appropriate punctuation available in the text, for example, it can be of four types, namely comma (,), period (.), question mark (?), and no punctuation.
[0054] In the first stage, the text input can be encoded. Here, the encoding can be performed in various appropriate ways. In particular, the unpunctuated text sentence can be split into a subword sequence and converted into a sequence of lexical embeddings. As an example, a pre-trained natural language (NLP) processing model can be used as the backbone model to construct a lexical encoder, and the model of the lexical encoder can be fine-tuned according to the data of a specific task. In some embodiments, the encoding of the text can be performed by a text encoder, which can be included in the solution according to the embodiments of the present disclosure.
[0055] In the first stage, the audio input can be processed. In particular, based on the type of the audio input, the audio input can be converted into an acoustic embedding H , a , i , a , i , ,
[0057] ,
[0056] , ,
[0055] , or a virtual embedding H i . In some embodiments, the audio processing can be performed by various appropriate acoustic processing components, which can be included in the solution according to the embodiments of the present disclosure.
[0056] In particular, for an audio input that is an audio representation of the content of the text input, such as a transcription with audio annotations, it can be converted into an acoustic embedding H a . Specifically, the audio input can be converted into acoustic features by a pre-trained acoustic model / acoustic feature extractor. Generally, the acoustic feature extractor can first be preprocessed on an unlabeled audio dataset through self-supervised training and fine-tuned according to the downstream punctuation restoration task. Then, a downsampling network can be further applied to shorten the length of the extracted acoustic features to obtain the acoustic embedding content. The purpose is to make the length of the acoustic embedding close to or equal to the sentence embedding so that the model can better align cross-modal information. As an example, a multi-layer convolutional network can be selected as the core component of the downsampling network.
[0057] For text without audio, a virtual embedding H i is used to fabricate the possibly missing acoustic features, i.e., if Then H a =H i As an example, the virtual embedding H i can be set as an array of learnable parameters of a fixed length, and it is expected to learn the representation of the missing audio. That is to say, the virtual audio sample corresponding to the audio-free text as described above can be a pre-defined vector sequence, such as an all-zero sequence, and the virtual embedding can be obtained from it, for example, directly used as the virtual embedding, or its length can be shortened so that the length of the virtual embedding is close to or equal to the sentence embedding, and the shortened sequence is used as the virtual embedding.
[0058] Then, based on the sequence of lexical embeddings and the acoustic embedding or virtual embedding, the acoustic and lexical features can be combined to generate a multimodal hybrid representation. This operation can be performed by a coordinate bootstrapper, which can jointly train audio-free and audio texts to overcome the missing problem. In particular, the coordinate bootstrapper jointly utilizes the acoustic and lexical features and applies an attention-based operation to learn the hybrid representation of the two modalities.
[0059] Specifically, first perform a self-attention operation on the sequence of lexical embeddings H l to capture the long-range dependencies S l in the sentence without punctuation, and apply a cross-attention operation between the sequence of lexical embeddings H l and the sequence of acoustic embeddings H a to form a cross-modal representation S a :
[0060] S l =Att(H l , H l , H l ) (1)
[0061] S a =Att(H a , H a , H l ) (2)
[0062] Here is an attention operation, where d k is the dimension size of the model. Note that for modality missing samples, we replace the acoustic embedding H i with the virtual embedding H a , and in this case, if then S a =Att(H i , H i , H l )
[0063] Then, a hybrid representation H is obtained by adding the attention-processed representation and the residual connection. h :
[0064] H h = S l + S a + H l (3)
[0065] In implementation, the coordination guide can be stacked into multiple layers to further increase the model capacity.
[0066] Finally, predictions are made from the hybrid representation through an output classifier layer. The output classifier layer consists of a linear projection and a softmax activation function, and H is input to the classifier layer h to predict the punctuation symbol sequence
[0067] In this way, we enable the representations of audio samples and the representations without audio samples to share the same embedding space. Therefore, the model can receive mixed data in the same training batch, and the trained model can punctuate audio texts and texts without audio.
[0068] It should be noted that, according to some embodiments of the present disclosure, the foregoing text encoder, acoustic processing component, and coordination guide can all be included as sub-modules / frameworks in a model training device according to an embodiment of the present disclosure, namely, the UniPunc framework. In this way, the UniPunc framework according to the present disclosure provides a general framework for solving the modality missing in the multi-modal punctuation restoration task. In addition, the solution according to the present disclosure can also be effectively applied to some current punctuation models and serve as a beneficial supplement thereto. As an example, the acoustic processing component and coordination guide according to the present disclosure can be applied to / added to the current punctuation models for correction so that they can be improved to handle modality missing samples.
[0069] In addition, the above examples of model training are equally applicable to multi-language scenarios. In particular, data and audio in different languages can be used simultaneously to train the model. In particular, data in multiple different languages, including text data and / or audio data, are used as inputs to perform the processing of each stage in the above solution according to the present disclosure. For example, the audio of each language is input into the audio component, and a vocabulary encoder can be uniformly used for texts in different languages, so as to obtain a model with enhanced performance suitable for multi-language scenarios.
[0070] Moreover, although not shown, before performing the above processing on texts and audio, equalization processing can be performed on the multi-language inputs, and the processed texts and audio of each language after equalization are further processed as described above, such as the processing of the above three stages. Details will not be described here.
[0071] The punctuation model according to some embodiments of the present disclosure can be used in various suitable punctuation restoration applications. The punctuation model according to some embodiments of the present disclosure can have good universality and can be applied to any speech recognition system. For example, it can further process the text output by the speech recognition system to optimize the punctuation restoration of the text output by the speech recognition system. For example, punctuation can be added to the text without punctuation, or the text with punctuation can be verified for further correction.
[0072] Figure 2 A flowchart of a punctuation restoration method according to some embodiments of the present disclosure is shown. In method 200, at step S201, the text to be added with punctuation is obtained, such as the output of speech recognition text, and at step S202, the punctuation model trained by the model training method according to the present disclosure is applied to the obtained text output to restore the punctuation in the text output. Thus, punctuation can be appropriately added to the text, thereby achieving accurate punctuation restoration or punctuation verification / correction.
[0073] In some embodiments, the input text can be input into the model together with the relevant audio, so that punctuation can be appropriately added to the input text based on both of them. If only the text is input into the model, it means that there is no audio for this text, then a virtual sample can be generated and input into the model together with the text, and punctuation can be added to the input text based on this. In some other embodiments, the audio information of the text input into the punctuation model can be obtained. For example, according to an indicator indicating whether the text input has audio, when the indicator indicates that the text has audio, the punctuation model can obtain the text audio. For example, the text audio and the text are input into the model together, or the text audio is obtained from a predetermined storage location. When the indicator indicates that the text has no audio, such as in the case of a single text, a virtual audio sample can be generated and input into the punctuation model. Thus, the punctuation model will perform reasonably to predict the punctuation symbols, and the text with punctuation restoration can be obtained. This process can be executed in a similar manner as the model training process described above, such as the processing of the above three stages, which will not be described in detail here.
[0074] Figure 4A A punctuation model training device 400 according to some embodiments of the present disclosure is shown. The device 400 includes a sample acquisition unit 401 configured to acquire text samples and corresponding audio samples for model training, where for the text samples obtained from text without audio, the corresponding audio samples are virtual samples; and a training unit 402 configured to train a model for speech recognition punctuation restoration based on the acquired text samples and corresponding audio samples for model training.
[0075] In some embodiments, the training unit 402 is further configured to perform attention-based processing on the acquired text samples and audio samples for model training to generate a multimodal hybrid representation, whereby the training unit 402 can perform model training based on the generated multimodal hybrid representation. The above multimodal hybrid representation can be generated by the hybrid representation generation unit 4021. Although not shown, in an exemplary implementation, model training can be performed by a component for punctuation prediction, such as a classifier or a similar model training component. In this case, such a component can be included in the training unit.
[0076] In some embodiments, the hybrid representation generation unit 4021 may further include a text conversion unit 4022 configured to convert the text samples into a sequence of word embeddings, an audio conversion unit 4023 configured to convert the audio samples into a sequence of acoustic embeddings, and a joint processing unit 4024 configured to perform attention-based operations on the sequence of word embeddings and the sequence of acoustic embeddings respectively; and combine the processing results of the sequence of word embeddings and the sequence of acoustic embeddings to generate a multimodal hybrid representation. It should be noted that although not shown, the text conversion unit and the audio conversion unit may not be included in the hybrid representation generation unit, or even outside the training unit. That is to say, text conversion and audio conversion can be processing of samples before model training, so the samples used as input for model training are actually the converted samples.
[0077] The above units can be implemented in various appropriate ways. In an exemplary implementation, the text conversion unit 4022 and the audio conversion unit 4023 can at least respectively correspond to the text encoder and the acoustic processing component described above, the joint processing unit 4024 can correspond to the coordination director described above, the hybrid representation generation unit 4021 can at least include or correspond to the coordination director described above, and can also include the word encoder, the acoustic processing component, and the coordination director described above; the training unit 402 can at least include or correspond to Figure 3 the corresponding modules / devices of all the processing stages shown, such as the word encoder, the acoustic processing component, the coordination director, the classifier, etc. described above.
[0078] Figure 4BFig. 0 shows a punctuation restoration device 410 according to some embodiments of the present disclosure. The device 410 includes an acquisition unit 411 configured to acquire the text output of speech recognition, and a punctuation restoration unit 412 configured to apply a punctuation model trained by a training method according to the embodiments of the present disclosure to the acquired text output to restore the punctuation in the text output. The punctuation restoration unit here can perform operations / processes similar to those of the training unit described above. For example, it can include the text encoder and acoustic processing components, coordination guide, classifier, etc. described above.
[0079] It should be noted that the above-mentioned respective units are only logical modules divided according to their specific functions implemented, rather than for limiting the specific implementation manners. For example, they can be implemented in software, hardware, or a combination of software and hardware. In actual implementation, the above-mentioned respective units can be implemented as independent physical entities, or can also be implemented by a single entity (such as a processor (CPU or DSP, etc.), integrated circuit, etc.). In addition, the above-mentioned respective units are shown by dashed lines in the drawings indicating that these units may not actually exist, and the operations / functions they implement can be implemented by the processing circuit itself.
[0080] In addition, although not shown, the device may further include a memory that can store various information generated during the operation of the device and each unit included in the device, programs and data for operation, data to be sent by the communication unit, etc. The memory can be a volatile memory and / or a non-volatile memory. For example, the memory can include, but is not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), flash memory. Of course, the memory can also be located outside the device. Optionally, although not shown, the device may also include a communication unit that can be used to communicate with other devices. In one example, the communication unit can be implemented in a suitable manner known in the art, such as including communication components such as antenna arrays and / or radio frequency links, various types of interfaces, communication units, etc. Details will not be described here. In addition, the device may further include other components not shown, such as radio frequency links, baseband processing units, network interfaces, processors, controllers, etc. Details will not be described here.
[0081] The present disclosure proposes the training and application of an improved punctuation model for punctuation restoration, which can be trained using both single text data and paired text and audio data simultaneously, thereby solving the problem of insufficient paired text and audio data. In particular, by using predefined vectors to replace audio as the input when there is no audio, a large amount of pure text data can be used to construct a training set, and audio data can be used for assistance, thus further improving the model training effect and enhancing the accuracy of the trained model. In addition, the present disclosure can also be trained collaboratively using multilingual data. In particular, by using language data with a large amount of data to enhance the model performance for languages with a small amount of data, a further optimized punctuation model is obtained. Thus, pure text data, text and audio-aligned data, and data in different languages can be trained together, improving the model's performance on pure text data, text and audio-aligned data, and data in minority languages. Moreover, such a model is suitable for various application tasks, especially suitable for speech recognition tasks, and achieves better punctuation restoration effects.
[0082] The effectiveness of the solution of the present disclosure will be further demonstrated below with examples.
[0083] For the training and test datasets, experiments will be mainly conducted on two real-world corpora: MuST-C and Multilingual TEDx (mTEDx), whose audio is sourced from TED talks. Datasets are constructed based on the following two corpora: 1) English Audio (English-Audio): This set contains the English audio and sentences in MuST-C, and each sample has audio. 2) English-Mixed: This set contains all English audio sentences and audio-free sentences from the two corpora. Note that English Audio is a subset of English-Mixed.
[0084] On the above datasets, the UniPunc according to the solution of the present disclosure and multimodal models in the related art are tested to compare performance. The tests prove that the UniPunc according to the embodiments of the present disclosure is superior to the multimodal models in the related art on the English Audio set, and the performance of the UniPunc according to the embodiments of the present disclosure on the English-Mixed set is better than its performance on the English Audio set. This shows that even without using audio-free sentences, the UniPunc according to the embodiments of the present disclosure can be superior to existing multimodal models, and the performance of the punctuation model according to the present disclosure can be further improved by using audio-free sentences.
[0085] On the above-mentioned English mixed dataset, UniPunc according to the present disclosure was tested with a single-modal model in the related art to compare performance. The test proved that UniPunc according to an embodiment of the present disclosure can effectively obtain multimodal mixed representation and effectively represent the acoustic features in speech, which can significantly improve the punctuation recovery of text.
[0086] Furthermore, testing on mTEDx's multilingual data revealed that UniPunc, implemented according to embodiments of the present disclosure, achieves punctuation recovery closer to that of humans, better distinguishing pauses between commas and periods, and the tone of questions. Furthermore, UniPunc, according to the present disclosure, achieves better punctuation performance than other baselines in multilingual punctuation, demonstrating its robustness and generalization capabilities.
[0087] Moreover, the UniPunc disclosed in the present invention can be well applied to any existing punctuation recovery method. In particular, by introducing the effective modules in the UniPunc framework disclosed in the present invention, especially the acoustic auxiliary module and / or the coordination guide, into the existing punctuation recovery scheme, the performance of the existing punctuation recovery scheme can be further optimized. Experiments show that the UniPunc scheme disclosed in the present invention, especially the modules such as the acoustic auxiliary module and / or the coordination guide, are universal for solving the modality loss in punctuation recovery, and enable the previous single-modal model to process multimodal corpus, further improving the overall performance.
[0088] Some embodiments of the present disclosure also provide an electronic device that can be operated to implement the operations / functions of the aforementioned model pre-training device and / or model training device. Figure 5 Block diagrams of some embodiments of the electronic device of the present disclosure are shown. For example, in some embodiments, the electronic device 5 can be various types of devices, for example, including but not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. For example, the electronic device 5 can include a display panel for displaying data and / or execution results utilized in the scheme of the present disclosure. For example, the display panel can be of various shapes, such as a rectangular panel, an elliptical panel, or a polygonal panel. In addition, the display panel can be not only a flat panel, but also a curved panel or even a spherical panel.
[0089] like Figure 5 As shown, the electronic device 5 of this embodiment includes: a memory 51 and a processor 52 coupled to the memory 51. It should be noted that Figure 5The components of the electronic device 50 shown are merely exemplary and not restrictive. According to actual application requirements, the electronic device 50 may also have other components. The processor 52 may control other components in the electronic device 5 to perform desired functions.
[0090] In some embodiments, the memory 51 is used to store one or more computer-readable instructions. When the processor 52 is used to run the computer-readable instructions, the computer-readable instructions are implemented to perform the method according to any of the above embodiments when run by the processor 52. For the specific implementation of each step of the method and related explanatory content, reference may be made to the above embodiments, and repeated parts will not be elaborated here.
[0091] For example, the processor 52 and the memory 51 may communicate with each other directly or indirectly. For example, the processor 52 and the memory 51 may communicate through a network. The network may include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 52 and the memory 51 may also communicate with each other through a system bus, and the present disclosure does not limit this.
[0092] For example, the processor 52 may be embodied as various appropriate processors, processing devices, etc., such as a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The central processing unit (CPU) may be of the X86 or ARM architecture, etc. For example, the memory 51 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The memory 51 may, for example, include a system memory, and the system memory may store, for example, an operating system, application programs, a boot loader, a database, and other programs. Various application programs and various data may also be stored in the storage medium.
[0093] In addition, according to some embodiments of the present disclosure, when various operations / processes according to the present disclosure are implemented by software and / or firmware, they can be transferred from a storage medium or a network to a computer system with a dedicated hardware structure, such as Figure 6 The computer system 600 shown installs a program that constitutes the software. When the computer system installs various programs, it can perform various functions, including the functions described above, etc. Figure 6 is a block diagram showing an example structure of a computer system that can be adopted in some embodiments of the present disclosure.
[0094] In Figure 6In this case, the central processing unit (CPU) 601 executes various processes in accordance with the programs stored in the read-only memory (ROM) 602 or the programs loaded from the storage section 608 into the random access memory (RAM) 603. In the RAM 603, data required when the CPU 601 executes various processes and the like is also stored as needed. The central processing unit is merely exemplary, and it may also be other types of processors, such as the various processors described above. The ROM 602, the RAM 603, and the storage section 608 may be various forms of computer-readable storage media, as described below. It should be noted that although Figure 6 the ROM 602, the RAM 603, and the storage device 608 are shown separately, one or more of them may be combined or located in the same or different memories or storage modules.
[0095] The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The input / output interface 605 is also connected to the bus 604.
[0096] The following components are connected to the input / output interface 605: an input section 606, such as a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output section 607, including a display, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage section 608, including a hard disk, a magnetic tape, etc.; and a communication section 609, including a network interface card such as a LAN card, a modem, etc. The communication section 609 allows communication processing to be performed via a network such as the Internet. It is easy to understand that although Figure 6 the various devices or modules in the electronic device 600 are shown to communicate via the bus 604, they may also communicate via a network or other means, where the network may include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.
[0097] As needed, a drive 610 is also connected to the input / output interface 605. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 610 as needed, so that the computer program read therefrom is installed into the storage section 608 as needed.
[0098] In the case where the above series of processes are implemented by software, the programs constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 611.
[0099] According to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods according to some embodiments of the present disclosure. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by a CPU 601, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.
[0100] It should be noted that, in the context of the present disclosure, a computer-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable medium can be either a computer-readable signal medium or a computer-readable storage medium or any combination of the two. A computer-readable storage medium, for example, can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. And in the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program codes are carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program codes contained on a computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0101] The above computer-readable medium can be included in the above electronic device; or it can exist separately and not be assembled into the electronic device.
[0102] In some embodiments, a computer program is also provided, including: instructions that, when executed by a processor, cause the processor to execute the method of any of the above embodiments. For example, the instructions may be embodied as computer program code.
[0103] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or may be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).
[0104] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0105] The modules, components, or units described in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the names of the modules, components, or units do not, in some cases, constitute a limitation on the modules, components, or units themselves.
[0106] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0107] According to some embodiments of the present disclosure, a method for training a model for speech recognition punctuation restoration is proposed, including the following steps: obtaining text samples and corresponding audio samples for model training, where for text samples obtained from text without audio, the corresponding audio samples are virtual samples; and training a model for speech recognition punctuation restoration based on the obtained text samples and audio samples for model training.
[0108] In some embodiments, the text samples and corresponding audio samples for model training can be obtained from both text with audio and text without audio, where paired text samples and audio samples are obtained from text with audio, and from text without audio, text samples are obtained and virtual samples are obtained as the audio samples corresponding to the text samples.
[0109] In some embodiments, the virtual sample is a preset sample for fabricating the audio of the text sample.
[0110] In some embodiments, training a model for speech recognition punctuation restoration may include: generating a hybrid representation based on the obtained text samples and audio samples for model training, and performing model training based on the hybrid representation.
[0111] In some embodiments, the hybrid representation can be a multi-modal hybrid representation obtained by performing attention-based processing based on the obtained text samples and audio samples for model training.
[0112] In some embodiments, attention-based processing can be performed on the audio samples for model training based on the text information related to the text samples.
[0113] In some embodiments, the text information related to the text samples can be the lexical features obtained by converting the text samples or the processing results obtained by performing attention-based processing on the text samples.
[0114] In some embodiments, the text samples are converted into a lexical embedding sequence, the audio samples are converted into an acoustic embedding sequence, and attention-based operations are respectively performed on the lexical embedding sequence and the acoustic embedding sequence.
[0115] In some embodiments, when the audio sample is in the audio form of the text content corresponding to the text sample, an acoustic embedding sequence is obtained based on the acoustic features extracted from the audio sample; and when the audio sample is a virtual sample, the virtual sample is used as the acoustic embedding sequence.
[0116] In some embodiments, the attention-based processing performed on the text sample for model training may be self-attention-based processing, and the attention-based processing performed on the audio sample for model training may be cross-attention-based processing.
[0117] In some embodiments, the text samples and the corresponding audio samples obtained for model training may include multilingual text samples and audio samples.
[0118] In some embodiments, the multilingual text samples and audio samples may be equalized to increase the proportion of low-resource language samples.
[0119] According to some embodiments of the present disclosure, a method for speech recognition punctuation restoration is provided, including the following steps: obtaining the text output of speech recognition, and applying a punctuation model trained by the model training method described in any one of the embodiments of the present disclosure to the obtained text output to restore the punctuation in the text output.
[0120] According to some other embodiments of the present disclosure, a training device for a model for speech recognition punctuation restoration is provided, including: a sample acquisition unit configured to acquire text samples and corresponding audio samples for model training, where for the text samples obtained from text without audio, the corresponding audio samples are virtual samples; and a training unit configured to train a model for speech recognition punctuation restoration based on the acquired text samples and audio samples for model training.
[0121] In some embodiments, the training unit may be further configured to: generate a multimodal mixed representation by performing attention-based processing on the acquired text samples and audio samples for model training.
[0122] In some embodiments, the device may further include: a text conversion unit configured to convert the text sample into a lexical embedding sequence, an audio conversion unit configured to convert the audio sample into an acoustic embedding sequence, and the training unit is configured to perform attention-based operations on the lexical embedding sequence and the acoustic embedding sequence respectively.
[0123] According to some other embodiments of the present disclosure, a device for speech recognition punctuation restoration is provided, including: an acquisition unit configured to acquire the text output of speech recognition, and a punctuation restoration unit configured to apply a punctuation model trained by the model training method described in any of the embodiments of the present disclosure to the acquired text output to restore the punctuation in the text output.
[0124] According to some other embodiments of the present disclosure, an electronic device is provided, including: a memory; and a processor coupled to the memory, wherein instructions are stored in the memory, and when the instructions are executed by the processor, the electronic device is caused to execute the method of any of the embodiments of the present disclosure.
[0125] According to some other embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method of any of the embodiments of the present disclosure is implemented.
[0126] According to some other embodiments of the present disclosure, a computer program is provided, including: instructions, and when the instructions are executed by a processor, the processor is caused to execute the method of any of the embodiments of the present disclosure.
[0127] According to some embodiments of the present disclosure, a computer program product is provided, including instructions, and when the instructions are executed by a processor, the method of any of the embodiments of the present disclosure is implemented.
[0128] The above description is only some embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0129] In the description provided herein, many specific details are set forth. However, it is understood that embodiments of the present invention may be practiced without these specific details. In other instances, well-known methods, structures, and technologies have not been shown in detail in order not to obscure the understanding of the description.
[0130] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0131] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and not for limiting the scope of the present disclosure. Those skilled in the art should understand that the above embodiments may be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A method for training a model for speech recognition punctuation recovery, comprising the following steps: Obtaining text samples and corresponding audio samples for model training, wherein the text samples and corresponding audio samples for model training are obtained from both text with audio and text without audio, obtaining paired text samples and audio samples from the text with audio, and obtaining a text sample from the text without audio and obtaining a virtual sample as the audio sample corresponding to the text sample, wherein the virtual sample is a pre-set sample of audio for the fictitious text sample; as well as A model for speech recognition punctuation recovery is trained based on the acquired text samples and corresponding audio samples for model training.
2. The method according to claim 1, wherein The model trained for speech recognition punctuation recovery includes: Generate a multimodal mixed representation based on the text samples and audio samples obtained for model training, and Model training is performed based on this multimodal mixed representation.
3. The method according to claim 2, wherein: The multimodal mixed representation is obtained by performing attention-based processing based on text samples and audio samples obtained for model training.
4. The method according to claim 3, wherein: Perform attention-based processing on audio samples used for model training based on textual information associated with the text samples.
5. The method according to claim 4, wherein The text information related to the text sample is a lexical feature obtained by converting the text sample or a processing result obtained by performing attention-based processing on the text sample.
6. The method according to claim 1, wherein Convert text samples into sequences of word embeddings, Convert audio samples into sequences of acoustic embeddings, Perform attention-based operations on the vocabulary embedding sequence and the acoustic embedding sequence respectively; and The manipulated lexical embedding sequence and acoustic embedding sequence are combined to generate a multimodal hybrid representation.
7. The method according to claim 6, wherein: In a case where the audio sample is an audio form of text content corresponding to the text sample, obtaining an acoustic embedding sequence based on acoustic features extracted from the audio sample; In the case where the audio sample is a virtual sample, the virtual sample is used as the acoustic embedding sequence.
8. The method according to any one of claims 3 to 7, wherein The attention-based processing performed on the text samples for model training is a self-attention-based processing, and the attention-based processing performed on the audio samples for model training is a cross-attention-based processing.
9. The method according to any one of claims 1 to 8, wherein The text samples and corresponding audio samples obtained for model training include text samples and corresponding audio samples in multiple languages.
10. The method according to claim 9, wherein: Multilingual text samples and corresponding audio samples are balanced to increase the proportion of low-resource language samples.
11. A method for recovering punctuation marks in speech recognition, comprising the following steps: Get the text output of speech recognition, and A model trained by the method according to any one of claims 1 to 10 is applied to the acquired text output to restore punctuation marks in the text output.
12. A training device for a model for speech recognition punctuation recovery, comprising: A sample acquisition unit is configured to acquire text samples and corresponding audio samples for model training, wherein the text samples and corresponding audio samples for model training are acquired from both texts with audio and texts without audio, and pairs of text samples and audio samples are acquired from texts with audio, and from texts without audio, a text sample is acquired and a virtual sample is acquired as the audio sample corresponding to the text sample, where the virtual sample is a pre-set sample of audio for a fictitious text sample; as well as The training unit is configured to train a model for speech recognition and punctuation recovery based on the acquired text samples and corresponding audio samples for model training.
13. The device according to claim 12, wherein The training unit further includes a mixed representation generating unit configured to: Attention-based processing is performed based on the text samples and audio samples obtained for model training to generate multimodal mixed representations.
14. The device according to claim 13, wherein The training unit further comprises: A text conversion unit, configured to convert a text sample into a sequence of word embeddings, an audio conversion unit configured to convert the audio samples into an acoustic embedding sequence, and A joint processing unit is configured to perform attention-based operations on the vocabulary embedding sequence and the acoustic embedding sequence respectively; and combine the operated vocabulary embedding sequence and the acoustic embedding sequence to generate a multimodal hybrid representation.
15. A device for recovering punctuation marks in speech recognition, comprising: an acquisition unit configured to acquire a text output of speech recognition, and A punctuation recovery unit is configured to output a punctuation mark for the acquired text. Apply the model trained by the method according to any one of claims 1-10 to restore punctuation in text output.
16. An electronic device comprising: Memory; and A processor coupled to the memory, wherein the memory stores instructions, and when the instructions are executed by the processor, the electronic device performs the method according to any one of claims 1-10.
17. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
18. A computer program product comprising instructions which, when executed by a processor, result in the implementation of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Rhythm phrase recognition method and device and electronic equipment
CN111640418A