Text-to-speech method, speech recognition method, training methods, apparatus, electronic device and storage medium
By mixing input plain text, pure audio and text audio to train the autoregressive model for discrete encoding of data, the problem of insufficient utilization of label-free data is solved, efficient speech synthesis and recognition model training is achieved, and the comprehensive performance of the model is improved.
Patent Information
- Application Number
- PCT/CN2024/131601
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-31
- Filing Date
- 2024-11-12
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, in the training scheme of the TTS model based on the text audio band pair data, label-free plain text and pure audio data cannot be effectively utilized, and the unsupervised training method is complex and has high cost.
The mixed input of plain text, pure audio and text audio is used to train the data, and the autoregressive model is input after discrete encoding processing to generate the target speech synthesis or recognition model.
It reduces the difficulty and cost of model training, improves training speed, improves the model's language understanding and acoustic comprehension ability, enhances the ability to capture emotions and tone, and provides more accurate output results.
Smart Images

Figure CN2024131601_03072025_PF_FP_ABST
Abstract
Description
Speech synthesis, speech recognition method, training method, device, electronic device, storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This disclosure claims priority to Chinese patent application number 202311873032.8 filed with the Chinese Patent Office on December 31, 2023, entitled “Speech synthesis, speech recognition method, training method, device, electronic device, storage medium,” the entire contents of which are incorporated by reference into this disclosure. Technical Field
[0003] The present disclosure relates to the technical field of multimedia content processing, and specifically to a training method for a speech synthesis model, a training method for a speech recognition model, a speech synthesis method, a speech recognition method, and related devices, models, electronic devices, and storage media. Background Art
[0004] In the existing technology, TTS (speech synthesis) models are trained based on data from text and audio segments. In the current training scheme, unlabeled plain text and pure audio cannot be effectively utilized because they have no labels or annotations. In addition, if an unsupervised method is used to train the text encoder and audio encoder, fine-tuning (adaptation to specific tasks or fields) is required in downstream tasks. The training method is relatively complex, and the training cost and difficulty are high.
[0005] The description of the background technology is intended to help understand the relevant technology in the relevant field, and does not mean that the background technology content is admitted to be prior art.
[0006] Application Contents
[0007] Therefore, the embodiments of the present disclosure provide a speech synthesis model training method, a speech recognition model training method, a speech synthesis method, a speech recognition method, and related models, electronic devices, and storage media. Through the solutions of the embodiments of the present disclosure, text, audio, or audio-text pairs can be discretely encoded and mixed into an autoregressive model, enabling mixed input of training data from multiple modalities in unsupervised training, thereby reducing the cost and difficulty of model training.
[0008] The present disclosure provides a method for training a speech synthesis model, comprising the following steps:
[0009] Acquire a training data set, the training data set including a plurality of data entries, wherein the types of the data entries include plain text data entries, plain audio data entries, and text-audio pair data entries;
[0010] Selecting a plurality of data entries from the training data set to generate a plurality of batch data entry sets, wherein each batch data entry set includes a plain text data entry, a plain audio data entry, and a text-audio pair data entry, and a ratio between the plain text data entry, the plain audio data entry, and the text-audio pair data entry in the batch data entry set satisfies a set ratio condition;
[0011] Performing discrete coding processing on the data entries in the batch data entry set to generate a plurality of data entry discrete codes, wherein each data entry discrete code includes a text discrete code and a voice discrete code, and the text discrete code in the data entry discrete code is located before the voice discrete code;
[0012] The autoregressive model is trained according to the discrete encoding of the multiple data entries in the multiple batch data entry sets to generate a target speech synthesis model.
[0013] In some embodiments of the present disclosure, the text discrete code includes text discrete code content, the speech discrete code includes speech discrete code content; and the discrete coding of the data entries in the batch data entry set to generate multiple data entry discrete codes includes:
[0014] In response to the data entry being a plain text data entry, discretely encoding the text content in the plain text data entry, obtaining the text discrete encoding content, and setting the voice discrete encoding content to be empty;
[0015] In response to the data entry being a pure audio data entry, discretely encoding the audio content in the pure audio data entry, obtaining the discretely encoded speech content, and setting the discretely encoded text content to be empty;
[0016] In response to the data entry being a text-audio pair data entry, the text content in the text-audio pair data entry is discretely encoded to obtain the text discretely encoded content, and the audio content in the text-audio pair data entry is discretely encoded to obtain the voice discretely encoded content.
[0017] In some embodiments of the present disclosure, the text discrete coding also includes a text discrete coding start mark and a text discrete coding end mark, and the text discrete coding content is located between the text discrete coding start mark and the text discrete coding end mark; the speech discrete coding also includes a speech discrete coding start mark and a speech discrete coding end mark, and the speech discrete coding content is located between the speech discrete coding start mark and the speech discrete coding end mark.
[0018] In some embodiments of the present disclosure, discrete encoding processing of text content to obtain the discretely encoded text content includes:
[0019] The text content is converted into phonemes to generate the text discrete coding content.
[0020] In some embodiments of the present disclosure, discrete encoding processing of audio content to obtain the discrete encoded speech content includes:
[0021] The audio content is input into an audio encoder to obtain the discrete speech coding content.
[0022] The present disclosure also provides a speech synthesis method, comprising the following steps:
[0023] Acquiring input data, wherein the input data includes a target text;
[0024] Performing discrete coding processing on the input data to obtain a discrete code of the input data, wherein the discrete code of the input data includes a text discrete code and a voice discrete code, and the text discrete code is located before the voice discrete code;
[0025] Inputting the discrete code of the input data into the target speech synthesis model to obtain the discrete code of the output speech;
[0026] The output voice discrete code is input into the voice decoder for decoding to obtain the target output audio.
[0027] In some embodiments of the present disclosure, the input data further includes a target audio segment having a set timbre and ambient sound.
[0028] In some embodiments of the present disclosure, the text discrete code includes text discrete code content, and the speech discrete code includes speech discrete code content;
[0029] When the input data is a target text, performing discrete coding processing on the input data to obtain a discrete code of the input data includes:
[0030] Performing discrete coding processing on the target text, obtaining the discrete coding content of the text, and leaving the discrete coding content of the speech blank;
[0031] When the input data is a target text and a target audio segment, performing discrete coding on the input data to obtain discrete coding of the input data includes:
[0032] The target text is discretely coded to obtain the discrete coded content of the text; the target audio segment is discretely coded to obtain the discrete coded content of the speech.
[0033] In some embodiments of the present disclosure, when the input data is target text, inputting the discrete encoding of the input data into the target speech synthesis model to obtain the output speech discrete encoding includes:
[0034] Input the discrete code of the input data into the target speech synthesis model, and obtain the output speech discrete code after decoding calculation;
[0035] When the input data is a target text and a target audio segment, the discrete coding of the input data is input into a target speech synthesis model to obtain the output speech discrete coding, including:
[0036] The text discrete coding content input data is discretely coded and input into the target speech synthesis model, and the speech discrete coding content is input into the target speech synthesis model in the form of prompt words. After decoding calculation, the output speech discrete coding is obtained.
[0037] In some embodiments of the present disclosure, the target speech synthesis model is trained by the speech synthesis model training method in any embodiment of the present disclosure.
[0038] The present disclosure also provides a method for training a speech recognition model, comprising the following steps:
[0039] Acquire a training data set, the training data set including a plurality of data entries, wherein the types of the data entries include plain text data entries, plain audio data entries, and text-audio pair data entries;
[0040] Selecting a plurality of data entries from the training data set to generate a plurality of batch data entry sets, wherein each batch data entry set includes a plain text data entry, a plain audio data entry, and a text-audio pair data entry, and a ratio between the plain text data entry, the plain audio data entry, and the text-audio pair data entry in the batch data entry set satisfies a set ratio condition;
[0041] Performing discrete coding processing on the data entries in the batch data entry set to generate a plurality of data entry discrete codes, wherein each data entry discrete code includes a text discrete code and a voice discrete code, and the voice discrete code in the data entry discrete code is located before the text discrete code;
[0042] The autoregressive model is trained according to the discrete encoding of the plurality of data entries in the plurality of batch data entry sets to generate a speech recognition model.
[0043] In some embodiments of the present disclosure, the text discrete code includes text discrete code content, the speech discrete code includes speech discrete code content; and the discrete coding of the data entries in the batch data entry set to generate multiple data entry discrete codes includes:
[0044] In response to the data entry being a plain text data entry, discretely encoding the text content in the plain text data entry, obtaining the text discrete encoding content, and setting the voice discrete encoding content to be empty;
[0045] In response to the data entry being a pure audio data entry, discretely encoding the audio content in the pure audio data entry, obtaining the discretely encoded speech content, and setting the discretely encoded text content to be empty;
[0046] In response to the data entry being a text-audio pair data entry, the text content in the text-audio pair data entry is discretely encoded to obtain the text discretely encoded content, and the audio content in the text-audio pair data entry is discretely encoded to obtain the voice discretely encoded content.
[0047] In some embodiments of the present disclosure, the text discrete coding also includes a text discrete coding start mark and a text discrete coding end mark, and the text discrete coding content is located between the text discrete coding start mark and the text discrete coding end mark; the speech discrete coding also includes a speech discrete coding start mark and a speech discrete coding end mark, and the speech discrete coding content is located between the speech discrete coding start mark and the speech discrete coding end mark.
[0048] In some embodiments of the present disclosure, discrete encoding processing of text content to obtain the discretely encoded text content includes:
[0049] The text content is converted into phonemes to generate the text discrete coding content.
[0050] In some embodiments of the present disclosure, discrete encoding processing of audio content to obtain the discrete encoded speech content includes:
[0051] The audio content is input into an audio encoder to obtain the discrete speech coding content.
[0052] The present disclosure also provides a speech recognition method, comprising the following steps:
[0053] Acquire input data, wherein the input data includes target audio;
[0054] Performing discrete coding processing on the input data to obtain a discrete code of the input data, wherein the discrete code of the input data includes a text discrete code and a voice discrete code, and the voice discrete code is located before the text discrete code;
[0055] Inputting the discrete code of the input data into the target speech recognition model to obtain the discrete code of the output text;
[0056] The output text discrete code is subjected to inverse discrete coding processing to obtain a target output text.
[0057] In some embodiments of the present disclosure, the text discrete code includes text discrete code content, and the speech discrete code includes speech discrete code content;
[0058] The step of performing discrete coding on the input data to obtain discrete codes of the input data includes:
[0059] The target audio is discretely coded to obtain the discrete coding content of the speech, and the discrete coding content of the text is left blank.
[0060] In some embodiments of the present disclosure, the step of inputting the discrete encoding of the input data into a target speech recognition model to obtain the discrete encoding of the output text includes:
[0061] The discrete encoding of the input data is input into the target speech recognition model, and after decoding calculation, the discrete encoding of the output text is obtained.
[0062] In some embodiments of the present disclosure, the target speech recognition model is trained by the speech recognition model training method in any embodiment of the present disclosure.
[0063] The embodiment of the present disclosure also provides a speech synthesis model training device, including a training data set acquisition module, a batch data integration module, a discrete coding processing module and a training module, wherein:
[0064] The training data set acquisition module is configured to acquire a training data set, wherein the training data set includes a plurality of data entries, and the types of the data entries include plain text data entries, plain audio data entries, and text-audio pair data entries;
[0065] The batch data integration module is configured to select multiple data entries from the training data set to generate multiple batch data entry sets, wherein each batch data entry set includes plain text data entries, pure audio data entries, and text-audio pair data entries, and the ratio between the plain text data entries, pure audio data entries, and text-audio pair data entries in the batch data entry set meets a set ratio condition;
[0066] The discrete code processing module is configured to perform discrete code processing on the data entries in the batch data entry set to generate a plurality of data entry discrete codes, wherein each data entry discrete code includes a text discrete code and a voice discrete code, and the text discrete code in the data entry discrete code is located before the voice discrete code;
[0067] The training module trains the autoregressive model according to the discrete encoding of multiple data entries in the multiple batch data entry sets to generate a target speech synthesis model.
[0068] The embodiments of the present disclosure also provide a speech synthesis model, which is trained by the speech synthesis model training method in any embodiment of the present disclosure.
[0069] The embodiment of the present disclosure further provides a speech synthesis device, comprising an input data acquisition module, a discrete coding processing module, a model calculation module and a speech decoder module, wherein:
[0070] The input data acquisition module is configured to acquire input data, wherein the input data includes a target text;
[0071] The discrete coding processing module is configured to perform discrete coding processing on the input data to obtain a discrete code of the input data, wherein the discrete code of the input data includes a text discrete code and a voice discrete code, and the text discrete code is located before the voice discrete code;
[0072] The model calculation module is configured to input the discrete code of the input data into the target speech synthesis model to obtain the discrete code of the output speech;
[0073] The speech decoder module is configured to perform speech decoding on the output speech discrete code to obtain target output audio.
[0074] The embodiment of the present disclosure also provides a speech recognition model training device, which includes a training data set acquisition module, a batch data integration module, a discrete coding processing module and a training module, wherein:
[0075] The training data set acquisition module is configured to acquire a training data set, wherein the training data set includes a plurality of data entries, and the types of the data entries include plain text data entries, plain audio data entries, and text-audio pair data entries;
[0076] The batch data integration module is configured to select multiple data entries from the training data set to generate multiple batch data entry sets, wherein each batch data entry set includes plain text data entries, pure audio data entries, and text-audio pair data entries, and the ratio between the plain text data entries, pure audio data entries, and text-audio pair data entries in the batch data entry set meets a set ratio condition;
[0077] The discrete code processing module is configured to perform discrete code processing on the data entries in the batch data entry set to generate a plurality of data entry discrete codes, wherein each data entry discrete code includes a text discrete code and a voice discrete code, and the voice discrete code in the data entry discrete code is located before the text discrete code;
[0078] The training module is configured to train an autoregressive model according to discrete encodings of multiple data entries in the multiple batch data entry sets to generate a speech recognition model.
[0079] The embodiments of the present disclosure also provide a speech recognition model, which is trained by the speech recognition model training method in any embodiment of the present disclosure.
[0080] The embodiment of the present disclosure further provides a speech recognition device, comprising an input data acquisition module, a discrete coding processing module, a model calculation module and an inverse discrete coding processing module, wherein:
[0081] The input data acquisition module is configured to acquire input data, wherein the input data includes target audio;
[0082] The discrete coding processing module is configured to perform discrete coding processing on the input data to obtain a discrete code of the input data, wherein the discrete code of the input data includes a text discrete code and a voice discrete code, and the voice discrete code is located before the text discrete code;
[0083] The model calculation module is configured to input the discrete code of the input data into the target speech recognition model to obtain the discrete code of the output text;
[0084] The inverse discrete encoding module is configured to perform inverse discrete encoding processing on the discrete encoding of the output text to obtain a target output text.
[0085] An embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the method of any one of the embodiments of the present disclosure is implemented.
[0086] An embodiment of the present disclosure further provides an electronic device, comprising: a processor and a memory storing a computer program, wherein the processor is configured to execute any method of the embodiment of the present disclosure when running the computer program.
[0087] The disclosed embodiments provide a method for training a speech synthesis model and a speech recognition model, wherein pure text, pure audio, and text-audio pair data are mixed in a set ratio to form batch training data, and the data items in the batch training data are discretely coded to form data item discrete codes, wherein each data item discrete code contains a text discrete code and a speech discrete code, and the data item discrete codes are input into an autoregressive model for training, ultimately obtaining a target speech synthesis model or speech recognition model. By mixing and inputting the discrete codes corresponding to pure text, pure audio, and text-audio pairs into the autoregressive model, the model training speed is increased and the training difficulty is reduced. The language understanding and language generation capabilities of the autoregressive model are trained using the discrete codes corresponding to the pure text, and the acoustic understanding capabilities of the autoregressive model are improved using the discrete codes corresponding to the pure audio, and the model's ability to understand and generate original speech is refined, thereby improving the model's ability to capture emotions or specific timbres. The discrete codes corresponding to the text-audio pair are used to train the autoregressive model's ability to predict the next audio or text segment using the previous segment of text or audio.
[0088] The disclosed embodiments can adjust the order of text discrete codes and speech discrete codes in data entries according to the characteristics of the required model. If the required model is a speech synthesis model, the text discrete code is placed before the speech discrete code. If the required model is a speech recognition model, the speech discrete code is placed before the text discrete code. For different generation tasks, the model architecture does not need to be adjusted. Adjusting the front-to-back position relationship in the input training data can meet the requirements, improving the reuse characteristics of the model architecture and having the value of wide promotion. The disclosed embodiments use a mixed method of unsupervised and supervised data to train the autoregressive model, improving data utilization while eliminating the need for multi-stage training and reducing training difficulty. The disclosed embodiments also provide a speech synthesis method that uses discrete coding to represent audio. Compared with traditional mel-spectrograms, it has better context learning capabilities. When audio is generated using audio prompts, the generated audio can retain the ambient sound and speaker's emotions of the audio prompts. The speech recognition method in this technical solution can recognize the corresponding text by inputting the input audio into the model after discrete encoding, avoiding the use of mel-spectrograms to represent audio, and has better context learning capabilities, and the output text is more accurate.
[0089] Other optional features and technical effects of the embodiments of the present disclosure are partially described below, and partially can be understood by reading this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0090] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use. It should be noted that the drawings described below only cover some embodiments of the present disclosure. A person skilled in the art can derive other drawings based on these drawings without having to engage in creative work. The purpose of these drawings is to better illustrate the technical details to facilitate understanding of the embodiments of the present disclosure, including:
[0091] FIG1 is a schematic flow chart of a speech synthesis model training method according to an embodiment of the present disclosure;
[0092] FIG2 is a schematic flow chart showing discrete coding processing in the speech synthesis model training method according to an embodiment of the present disclosure;
[0093] FIG3 is a schematic diagram showing input data of an autoregressive model in a speech synthesis model training method according to an embodiment of the present disclosure;
[0094] FIG4 is a schematic flow chart showing a method for training a speech recognition model according to an embodiment of the present disclosure;
[0095] FIG5 is a schematic flow chart showing discrete coding processing in the speech recognition model training method according to an embodiment of the present disclosure;
[0096] FIG6 is a schematic diagram showing input data of an autoregressive model in a speech recognition model training method according to an embodiment of the present disclosure;
[0097] FIG7 shows a schematic flow chart of a speech synthesis method according to an embodiment of the present disclosure;
[0098] FIG8 is a schematic diagram showing input data of a speech synthesis model in a speech synthesis method according to an embodiment of the present disclosure;
[0099] FIG9 is a schematic flow chart of a speech recognition method according to an embodiment of the present disclosure;
[0100] FIG10 is a schematic diagram showing input data of a speech recognition model in a speech recognition method according to an embodiment of the present disclosure;
[0101] FIG11 shows an exemplary structural diagram of a speech synthesis model training device according to an embodiment of the present disclosure;
[0102] FIG12 shows an exemplary structural diagram of a speech recognition model training device according to an embodiment of the present disclosure;
[0103] FIG13 shows an exemplary structural diagram of a speech synthesis device according to an embodiment of the present disclosure;
[0104] FIG14 shows an exemplary structural diagram of a speech recognition device according to an embodiment of the present disclosure;
[0105] FIG15 shows an exemplary structural diagram of a speech synthesis model according to an embodiment of the present disclosure;
[0106] FIG16 shows an exemplary structural diagram of a speech recognition model according to an embodiment of the present disclosure;
[0107] FIG17 shows a schematic diagram of an exemplary structure of an electronic device capable of implementing the method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0108] The following detailed description of exemplary embodiments of the present disclosure is provided, and illustrations of the exemplary embodiments are shown in the accompanying drawings. When referring to the drawings, unless otherwise specified, identical numbers or symbols in different figures represent identical or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. Instead, they are merely examples of some aspects of the apparatus and methods covered by one or more embodiments of this specification, as detailed in the claims of this disclosure.
[0109] As used in this specification, the term "including" and its variations are intended to be broadly inclusive, meaning "including but not limited to" the listed items. Unless otherwise stated, the term "or" means "and / or," the term "based on" means relying on, or at least partially relying on, the terms "an example embodiment" and "an embodiment" refer to at least one example embodiment, and the term "another embodiment" refers to at least one different embodiment. The terms "first," "second," and so on may refer to different or the same items. Other explicit and implicit definitions may be included below.
[0110] In the embodiments of the present disclosure, speech synthesis can generate target audio by inputting text or text plus an audio segment into a speech synthesis model; speech recognition can generate target text by inputting audio into a speech recognition model. In the embodiments of the present disclosure, the speech synthesis model and speech recognition model can be autoregressive models or large language models. In some embodiments, they can be large language models with only an encoder. The architecture of the speech synthesis model and speech recognition model in the embodiments of the present disclosure can be designed with reference to the Vall-E model (https: / / arxiv.org / pdf / 2301.02111.pdf).
[0111] In the embodiments of the present disclosure, TTS (Text To Speech) refers to a speech synthesis technology, ASR (Auto Speech Recognition) refers to a speech recognition technology, LLM (large language model) refers to a large language model, and prompt refers to a prompt word, which, as part of the language model input, provides speaker-related information to guide the language model to perform tasks.
[0112] In the embodiments of the present disclosure, "timbre" is the fundamental characteristic that distinguishes one voice from another, and different people have different timbres. "Prosody" is a property of speech, including pitch and rhythm of musical speech. "Semantics" is a property of speech that identifies the content of the speech.
[0113] In the embodiments of the present disclosure, "model" has the conventional meaning in the field of machine learning. For example, the model can be a machine learning or deep learning model, such as a machine learning or deep learning model that includes the above-mentioned network or is composed of the above-mentioned network.
[0114] In the embodiments of the present disclosure, “loss function” and “loss value” have conventional meanings in the field of machine learning.
[0115] In the disclosed embodiments, "BERT (Bidirectional Encoder Representation from Transformers)" is a pre-trained language representation model. "Prompt" refers to a prompt word, which serves as part of the language model input and provides relevant information to guide the language model in its tasks. "Embedding" refers to the data processing step that converts discrete data such as text into continuous data vectors so that the model can perform calculations on them.
[0116] In the embodiments of the present disclosure, a "phoneme" is the smallest unit of speech divided according to the natural properties of speech, and one pronunciation action forms a phoneme. A "phoneme sequence" is a sequence composed of several phonemes.
[0117] In the embodiments of the present disclosure, "batch" refers to batch processing. During the model training process, "batch" refers to a batch of data to be processed before each model improvement, and the model parameters are adjusted according to the output results of this batch of training data; "token" refers to discrete coding, which can refer to word units or phonemes for text, and can refer to audio discrete coding for audio.
[0118] The speech synthesis model training method provided in the embodiments of this disclosure first discretizes text and speech into tokens. Then, during training, the LLM model is trained by mixing the three data types in a certain ratio within each batch, separating the text and audio using their start and stop tokens. The three data types are represented as follows.
[0119] Plain text data:<text start> <text><text end><audio start>null<audio end> .
[0120] Pure audio data:<text start> null<text end><audio start> <audio><audio end>.
[0121] Audio-text pairs:<text start> <text><text end><audio start> <audio><audio end>.
[0122] Input the three types of data into the LLM model or autoregressive model to train and generate a speech synthesis model.
[0123] If the audio token is placed before the text token, similar training can be performed, and this model can be used as a speech recognition model. The audio-text data representation is as follows.
[0124] Plain text data:<audio start> null<audio end><text start> <text><text end>.
[0125] Pure audio data:<audio start> <audio><audio end><text start>null<text end> .
[0126] Audio-text pairs:<audio start> <audio><audio end><text start> <text><text end>.
[0127] Input the three types of data into the LLM model or autoregressive model to train and generate a speech recognition model.
[0128] The embodiments of the present disclosure also provide a speech synthesis / recognition model training method and device, a speech synthesis / recognition method, a speech synthesis / recognition model, a synthesis / recognition device, a storage medium, and an electronic device. The method, device / model can be implemented with the aid of one or more computers. In some embodiments, the device / model can be implemented by software, hardware, or a combination of software and hardware. In some embodiments, the electronic device or computer can be implemented by the computer described herein or other electronic devices that can implement the corresponding functions.
[0129] The inventors of this disclosure discovered that current TTS training schemes require paired text and audio (labeled data), and that text data without audio or audio data without text cannot be effectively utilized. Furthermore, unsupervised training of text encoders and audio encoders, followed by fine-tuning of downstream tasks, complicates the training process and increases costs. Therefore, this disclosure proposes a method for generating a TTS model or ASR model by training a discrete coded mixed input autoregressive model corresponding to pure text, pure audio, and text-audio pair data. This simplifies the training process and improves efficiency.
[0130] As shown in FIG1 , the present disclosure provides a speech synthesis model training method, including the following steps:
[0131] S110: Acquire a training data set, where the training data set includes multiple data entries, and the types of data entries include plain text data entries, plain audio data entries, and text-audio pair data entries.
[0132] The disclosed embodiments reduce the requirements for the training data set, enabling the training method to be widely used.
[0133] S120: Select multiple data entries from the training data set to generate multiple batch data entry sets, where each batch data entry set contains plain text data entries, plain audio data entries, and text-audio pair data entries, and the ratio between the plain text data entries, plain audio data entries, and text-audio pair data entries in the batch data entry set meets the set ratio condition.
[0134] Optionally, each batch data entry set includes plain text data entries, pure audio data entries and text-audio pair data entries. For example, the batch data entry set batch is {{text 1, text 2…text P}, {audio 1, audio 2…audio N}, {text-audio pair 1, text-audio pair 2…text-audio pair M}}, where P is the number of plain text items in the batch, N is the number of pure audio items in the batch, and M is the number of text-audio pairs in the batch.
[0135] In the embodiment of the present disclosure, the set ratio conditions between the plain text data entries, the plain audio data entries and the text audio pair data entries can be: 20:1>P:M>10:1, 20:1>N:M>10:1.
[0136] S130: Perform discrete coding processing on the data entries in the batch data entry set to generate multiple data entry discrete codes, wherein each data entry discrete code includes a text discrete code and a voice discrete code, and the text discrete code is located before the voice discrete code in the data entry discrete code.
[0137] In some embodiments of the present disclosure, the text discrete code includes text discrete code content, and the speech discrete code may include speech discrete code content. As shown in FIG2 , discrete coding is performed on data entries in a batch data entry set to generate multiple data entry discrete codes, including:
[0138] S131: In response to the data entry being a plain text data entry, discretely encode the text content in the plain text data entry, obtain the text discrete encoding content, and set the voice discrete encoding content to be empty.
[0139] For example, a data entry corresponding to a plain text data entry is discretely encoded as {token(text content), token(NULL)}.
[0140] S132: In response to the data entry being a pure audio data entry, discretely encode the audio content in the pure audio data entry, obtain the discretely encoded speech content, and set the discretely encoded text content to be empty.
[0141] For example, a data entry corresponding to a pure audio data entry is discretely encoded as {token(NULL), token(audio content)}.
[0142] S133: In response to the data entry being a text-audio pair data entry, discretely encode the text content in the text-audio pair data entry to obtain text discretely encoded content, and discretely encode the audio content in the text-audio pair data entry to obtain voice discretely encoded content.
[0143] For example, the data entry corresponding to the text-audio pair data entry is discretely encoded as {token (text content), token (audio content)}.
[0144] In some embodiments of the present disclosure, in order to facilitate the identification of the starting position of text and audio, the text discrete coding may also include a text discrete coding start mark and a text discrete coding end mark, and the text discrete coding content is located between the text discrete coding start mark and the text discrete coding end mark; the voice discrete coding also includes a voice discrete coding start mark and a voice discrete coding end mark, and the voice discrete coding content is located between the voice discrete coding start mark and the voice discrete coding end mark.
[0145] For example, the text discrete encoding start flag can be set to token(text start), the text discrete encoding end flag can be set to token(text end), the voice discrete encoding start flag can be set to token(audio start), and the voice discrete encoding end flag can be set to token(audio end). Then the data entry corresponding to the plain text is discretely encoded as {token(text start), token(text content), token(text end), token(audio start), token(NULL), token(audio end)}; the data entry corresponding to the pure audio is discretely encoded as {token(text start), token(NULL), token(text end), token(audio start), token(audio content), token(audio end)}; the data entry corresponding to the text and audio pair is discretely encoded as {token(text start), token(text content), token(text end), token(audio start), token(audio content), token(audio end)}.
[0146] In some embodiments of the present disclosure, discrete encoding processing of text content to obtain discretely encoded text content includes:
[0147] Perform phoneme conversion on text content to generate text discrete coding content.
[0148] In some embodiments of the present disclosure, discrete encoding processing of audio content to obtain discrete speech encoding content includes:
[0149] The audio content is input into the audio encoder to obtain discrete speech coding content.
[0150] S140: Training an autoregressive model according to discrete encoding of multiple data entries in multiple batch data entry sets to generate a target speech synthesis model.
[0151] As shown in Figure 3, the autoregressive model has a first input, input1, and a second input, input2. During training, the data items are discretely encoded and fed into the autoregressive model. The text discrete encoding is fed into input1, and the speech discrete encoding is fed into input2. After calculation, the autoregressive model outputs the audio discrete encoding. This audio discrete encoding is then de-encoded to produce the output audio. This ultimately results in a trained speech synthesis model.
[0152] The training and parameter adjustment steps of the autoregressive model in this disclosure are similar to those of the LLM model.
[0153] The speech synthesis model training method of the embodiment of the present disclosure mixes pure text, pure audio, and text-audio pair data in a set ratio to form batch training data, performs discrete coding processing on the data items in the batch training data to form data item discrete codes, each data item discrete code contains a text discrete code and a speech discrete code, and inputs the data item discrete codes into an autoregressive model for training, ultimately obtaining a target speech synthesis model. By mixing and inputting the discrete codes corresponding to pure text, pure audio, and text-audio pairs into the autoregressive model, the model training speed is improved and the training difficulty is reduced. The language understanding and language generation capabilities of the autoregressive model are trained using the discrete codes corresponding to the pure text, and the acoustic understanding capabilities of the autoregressive model are improved using the discrete codes corresponding to the pure audio, and the model's ability to understand and generate original speech is refined, thereby improving the ability to capture emotions or specific timbres. The discrete codes corresponding to the text-audio pairs are used to train the autoregressive model's ability to predict the next audio segment using the previous segment of text and audio.
[0154] As shown in FIG4 , the present disclosure also provides a method for training a speech recognition model, including the following steps:
[0155] S210: Obtain a training data set, the training data set including multiple data entries, the types of which include pure text data entries, pure audio data entries, and text-audio pair data entries. The disclosed embodiment reduces the requirements for the training data set, making the training method widely applicable.
[0156] S220: Select multiple data entries from the training data set to generate multiple batch data entry sets, where each batch data entry set contains plain text data entries, plain audio data entries, and text-audio pair data entries, and the ratio between the plain text data entries, plain audio data entries, and text-audio pair data entries in the batch data entry set meets the set ratio condition.
[0157] Each batch data entry set in the present disclosure includes plain text data entries, pure audio data entries, and text-audio pair data entries. For example, the batch data entry set batch is {{text 1, text 2…text P}, {audio 1, audio 2…audio N}, {text-audio pair 1, text-audio pair 2…text-audio pair M}}, where P is the number of plain text items in the batch, N is the number of pure audio items in the batch, and M is the number of text-audio pairs in the batch. In the embodiment of the present disclosure, the set ratio conditions between the plain text data entries, the pure audio data entries, and the text-audio pair data entries can be: 20:1>P:M>10:1, 20:1>N:M>10:1.
[0158] S230: Perform discrete coding processing on the data entries in the batch data entry set to generate multiple data entry discrete codes, wherein each data entry discrete code includes a text discrete code and a voice discrete code, and the voice discrete code in the data entry discrete code is located before the text discrete code.
[0159] In some embodiments of the present disclosure, the text discrete code includes text discrete code content, and the speech discrete code includes speech discrete code content. As shown in FIG5 , discrete coding is performed on data entries in a batch data entry set to generate multiple data entry discrete codes, including:
[0160] S231: In response to the data entry being a plain text data entry, discretely encode the text content in the plain text data entry, obtain the text discrete encoding content, and set the voice discrete encoding content to be empty.
[0161] For example, a data entry corresponding to a plain text data entry is discretely encoded as {token(NULL), token(text content)}.
[0162] S232: In response to the data entry being a pure audio data entry, discretely encode the audio content in the pure audio data entry, obtain the discretely encoded speech content, and set the discretely encoded text content to be empty.
[0163] For example, a data entry corresponding to a pure audio data entry is discretely encoded as {token(audio content), token(NULL)}.
[0164] S233: In response to the data entry being a text-audio pair data entry, discretely encode the text content in the text-audio pair data entry to obtain text discretely encoded content, and discretely encode the audio content in the text-audio pair data entry to obtain voice discretely encoded content.
[0165] For example, a data entry corresponding to a text-audio pair data entry is discretely encoded as {token (audio content), token (text content)}.
[0166] In some embodiments of the present disclosure, in order to facilitate the identification of the starting position of text and audio, the text discrete coding also includes a text discrete coding start mark and a text discrete coding end mark, and the text discrete coding content is located between the text discrete coding start mark and the text discrete coding end mark; the voice discrete coding also includes a voice discrete coding start mark and a voice discrete coding end mark, and the voice discrete coding content is located between the voice discrete coding start mark and the voice discrete coding end mark.
[0167] For example, the text discrete encoding start flag can be set to token(text start), the text discrete encoding end flag can be set to token(text end), the voice discrete encoding start flag can be set to token(audio start), and the voice discrete encoding end flag can be set to token(audio end). Then the data entry corresponding to the plain text is discretely encoded as {token(audio start), token(NULL), token(audio end), token(text start), token(text content), token(text end)}; the data entry corresponding to the pure audio is discretely encoded as {token(audio start), token(audio content), token(audio end), token(text start), token(NULL), token(text end)}; the data entry corresponding to the text and audio pair is discretely encoded as {token(audio start), token(audio content), token(audio end), token(text start), token(text content), token(text end)}.
[0168] In some embodiments of the present disclosure, discrete encoding processing of text content to obtain discretely encoded text content includes:
[0169] Perform phoneme conversion on text content to generate text discrete coding content.
[0170] In some embodiments of the present disclosure, discrete encoding processing of audio content to obtain discrete speech encoding content includes:
[0171] The audio content is input into the audio encoder to obtain discrete speech coding content.
[0172] S240: Training an autoregressive model according to discrete encoding of multiple data items in multiple batch data item sets to generate a speech recognition model.
[0173] Optionally, as shown in Figure 6, the autoregressive model has a first input entry, input1, and a second input entry, input2. During training, the data item discrete encoding is input into the autoregressive model, where the text discrete encoding content is input into the second input entry, input2, and the speech discrete encoding content is input into the first input entry, input1. After calculation, the autoregressive model outputs the text discrete encoding, which is then de-encoded to obtain the output text. This ultimately results in a trained speech recognition model.
[0174] Optionally, the training and parameter adjustment steps of the autoregressive model in the embodiment of the present disclosure are similar to the training and parameter adjustment steps of the LLM model.
[0175] The speech recognition model training method of the embodiment of the present disclosure mixes pure text, pure audio, and text-audio pair data in a set ratio to form batch training data, performs discrete coding processing on the data items in the batch training data to form data item discrete codes, each data item discrete code contains a text discrete code and a voice discrete code, and inputs the data item discrete codes into the autoregressive model for training, thereby finally obtaining the target speech recognition model. By mixing and inputting the discrete codes corresponding to pure text, pure audio, and text-audio pairs into the autoregressive model, the model training speed is improved and the training difficulty is reduced. The language understanding and language generation capabilities of the autoregressive model are trained using the discrete codes corresponding to the pure text, and the acoustic understanding capabilities of the autoregressive model are improved using the discrete codes corresponding to the pure audio, and the model's ability to understand and generate original speech is refined, which can improve the ability to capture emotions or specific timbres. The discrete codes corresponding to the text-audio pairs are used to train the autoregressive model's ability to predict the next text segment using the previous text and audio segment.
[0176] As shown in FIG7 , the embodiment of the present disclosure provides a speech synthesis method, comprising the following steps:
[0177] S310: Acquire input data, where the input data includes a target text.
[0178] The embodiments of the present disclosure can directly generate target output audio based on the input target text. In some embodiments, in order to obtain output audio with a specific timbre or a specific ambient sound, the input data may further include a target audio segment with a set timbre and ambient sound.
[0179] S320: Perform discrete coding processing on the input data to obtain discrete coding of the input data. The discrete coding of the input data includes text discrete coding and voice discrete coding, and the text discrete coding is located before the voice discrete coding.
[0180] In some embodiments of the present disclosure, text discrete coding includes text discrete coding content, and speech discrete coding includes speech discrete coding content.
[0181] Optionally, when the input data is target text, discrete coding is performed on the input data to obtain discrete coding of the input data, including:
[0182] Perform discrete coding on the target text, obtain the text discrete coding content, and set the speech discrete coding content to blank.
[0183] For example, the input data is discretely encoded as {token(target text), token(NULL)}.
[0184] In some embodiments of the present disclosure, the target text is converted into phonemes to generate discrete coding content of the text.
[0185] Optionally, when the input data is a target text and a target audio segment, performing discrete encoding processing on the input data to obtain discrete encoding of the input data includes:
[0186] The target text is discretely coded to obtain the text discrete coding content; the target audio segment is discretely coded to obtain the speech discrete coding content.
[0187] In some embodiments of the present disclosure, the target audio segment is input into an audio encoder to obtain discrete speech encoding content.
[0188] For example, the input data is discretely encoded as {token (target text), token (target audio segment)}.
[0189] S330: Input the discrete code of the input data into the target speech synthesis model to obtain the discrete code of the output speech.
[0190] In some embodiments of the present disclosure, when the input data is a target text, the discrete coding of the input data is input into the target speech synthesis model to obtain the discrete coding of the output speech, including:
[0191] The discrete encoding of the input data is input into the target speech synthesis model, and after decoding calculation, the discrete encoding of the output speech is obtained.
[0192] The target speech synthesis model in the embodiment of the present disclosure may be an autoregressive model, which has multiple decoders. The multiple decoders are used to calculate the discrete coding of the input data to obtain the discrete coding of the output speech.
[0193] Optionally, when the input data is a target text and a target audio segment, inputting the discrete encoding of the input data into the target speech synthesis model to obtain the discrete encoding of the output speech includes:
[0194] The text discrete coding content input data is discretely coded and input into the target speech synthesis model, and the speech discrete coding content is input into the target speech synthesis model in the form of prompt words. After decoding calculation, the output speech discrete coding is obtained.
[0195] In the embodiments of the present disclosure, the model decoding calculation is adjusted by means of voice prompt words, so that the output audio can better have the timbre, ambient sound, and emotional characteristics of the target audio segment.
[0196] Optionally, as shown in FIG8 , the target speech synthesis model has a first input entry, input1, and a second input entry, input2. During training, discrete codes of data items need to be input into the target speech synthesis model. The text discrete codes are input into the first input entry, input1, and the speech discrete codes are input into the second input entry, input2. After calculation by the target speech synthesis model, the target speech synthesis model outputs the speech discrete codes.
[0197] In some embodiments of the present disclosure, the target speech synthesis model is trained by the speech synthesis model training method in any embodiment of the present disclosure.
[0198] S340: Input the output voice discrete code into the voice decoder for decoding to obtain the target output audio
[0199] As shown in FIG8 , in some embodiments of the present disclosure, decoding may be implemented using an audio decoder. After the output speech is discretely encoded and input into the audio decoder, the target output audio is obtained.
[0200] The speech synthesis method in the embodiment of the present disclosure uses discrete coding to represent audio. Compared with traditional mel-spectrograms, it has better context learning capabilities. When audio is generated using audio prompts, the generated audio can maintain the ambient sound of the audio prompts and the speaker's emotions.
[0201] In another embodiment, the disclosed inventive concept may also be applied to speech recognition, such as automatic speech recognition (ASR).
[0202] As shown in FIG9 , the embodiment of the present disclosure provides a speech recognition method, comprising the following steps:
[0203] S410: Acquire input data, where the input data includes target audio.
[0204] The embodiment of the present disclosure can directly generate a target text based on the input target audio.
[0205] S420: Perform discrete coding processing on the input data to obtain discrete coding of the input data. The discrete coding of the input data includes text discrete coding and voice discrete coding, and the voice discrete coding is located before the text discrete coding.
[0206] In some embodiments of the present disclosure, the text discrete coding may include text discrete coding content, and the speech discrete coding may include speech discrete coding content;
[0207] Perform discrete coding processing on the input data to obtain discrete coding of the input data, including:
[0208] Perform discrete coding on the target audio, obtain the discrete coding content of the voice, and set the discrete coding content of the text to be empty.
[0209] For example, the input data is discretely encoded as {token(target audio), token(NULL)}.
[0210] S430: Inputting the discrete code of the input data into the target speech recognition model to obtain the discrete code of the output text.
[0211] In some embodiments of the present disclosure, inputting discrete encoding of input data into a target speech recognition model to obtain discrete encoding of output text includes:
[0212] The discrete encoding of the input data is input into the target speech recognition model, and after decoding calculation, the discrete encoding of the output text is obtained.
[0213] Optionally, as shown in FIG10 , the target speech recognition model has a first input entry, input1, and a second input entry, input2. During training, the data item discrete codes need to be input into the target speech recognition model, wherein the text discrete codes are input into the second input entry, input2, and the speech discrete codes are input into the first input entry, input1. After calculation by the target speech recognition model, the target speech recognition model outputs the text discrete codes.
[0214] In some embodiments of the present disclosure, the target speech recognition model is trained by the speech recognition model training method in any embodiment of the present disclosure.
[0215] S440: Perform inverse discrete encoding processing on the discrete encoding of the output text to obtain the target output text.
[0216] As shown in FIG10 , in some embodiments of the present disclosure, the inverse discrete encoding process can be implemented using a phoneme-to-text machine, where the output text is discretely encoded and input into the phoneme-to-text machine to obtain the output target output text.
[0217] The speech recognition method in the embodiment of the present disclosure discretizes the input audio and inputs it into the model, thereby being able to recognize the corresponding text, avoiding the use of mel-spectrogram to represent the audio, having better context learning capabilities, and outputting more accurate text.
[0218] As shown in FIG11 , an embodiment of the present disclosure provides a speech synthesis model training device 1100 , which includes a training data set acquisition module 1110 , a batch data integration module 1120 , a discrete coding processing module 1130 , and a training module 1140 .
[0219] The training data set acquisition module 1110 is configured to acquire a training data set, the training data set including a plurality of data entries, the types of which include plain text data entries, plain audio data entries, and text-audio pair data entries;
[0220] The batch data integration module 1120 is configured to select multiple data entries from the training data set to generate multiple batch data entry sets, wherein each batch data entry set includes plain text data entries, pure audio data entries, and text-audio pair data entries, and the ratio between the plain text data entries, pure audio data entries, and text-audio pair data entries in the batch data entry set meets a set ratio condition;
[0221] The discrete code processing module 1130 is configured to perform discrete code processing on the data entries in the batch data entry set to generate a plurality of data entry discrete codes, wherein each data entry discrete code includes a text discrete code and a voice discrete code, and the text discrete code is located before the voice discrete code in the data entry discrete code;
[0222] The training module 1140 trains the autoregressive model according to the discrete encoding of multiple data entries in the multiple batch data entry sets to generate a target speech synthesis model.
[0223] As shown in FIG12 , the embodiment of the present disclosure provides a speech recognition model training device 1200, comprising a training data set acquisition module 1210, a batch data integration module 1220, a discrete coding processing module 1230, and a training module 1240, wherein:
[0224] The training data set acquisition module 1210 is configured to acquire a training data set, the training data set including a plurality of data entries, the types of which include plain text data entries, plain audio data entries, and text-audio pair data entries;
[0225] The batch data integration module 1220 is configured to select multiple data entries from the training data set to generate multiple batch data entry sets, wherein each batch data entry set includes plain text data entries, pure audio data entries, and text-audio pair data entries, and the ratio between the plain text data entries, pure audio data entries, and text-audio pair data entries in the batch data entry set meets a set ratio condition;
[0226] The discrete code processing module 1230 is configured to perform discrete code processing on the data entries in the batch data entry set to generate a plurality of data entry discrete codes, wherein each data entry discrete code includes a text discrete code and a voice discrete code, and the voice discrete code in the data entry discrete code is located before the text discrete code;
[0227] The training module 1240 is configured to train the autoregressive model according to the discrete encoding of multiple data items in the multiple batch data item sets to generate a speech recognition model.
[0228] As shown in FIG13 , the embodiment of the present disclosure provides a speech synthesis device 1300, which includes an input data acquisition module 1310, a discrete coding processing module 1320, a model calculation module 1330, and a speech decoding module 1340, wherein:
[0229] The input data acquisition module 1310 is configured to acquire input data, where the input data includes a target text;
[0230] The discrete coding processing module 1320 is configured to perform discrete coding processing on the input data to obtain discrete coding of the input data. The discrete coding of the input data includes text discrete coding and voice discrete coding, and the text discrete coding is located before the voice discrete coding.
[0231] The model calculation module 1330 is configured to input the discrete code of the input data into the target speech synthesis model and obtain the discrete code of the output speech;
[0232] The speech decoder module 1340 is configured to perform speech decoding on the output speech discrete code to obtain target output audio.
[0233] As shown in FIG14 , the embodiment of the present disclosure provides a speech recognition device 1400, which includes an input data acquisition module 1410, a discrete coding processing module 1420, a model calculation module 1430, and an inverse discrete coding processing module 1440, wherein:
[0234] The input data acquisition module 1410 is configured to acquire input data, where the input data includes target audio;
[0235] The discrete coding processing module 1420 is configured to perform discrete coding processing on the input data to obtain discrete coding of the input data. The discrete coding of the input data includes text discrete coding and voice discrete coding, and the voice discrete coding is located before the text discrete coding.
[0236] The model calculation module 1430 is configured to input the discrete code of the input data into the target speech recognition model and obtain the discrete code of the output text;
[0237] The de-discrete encoding module 1440 is configured to perform de-discrete encoding processing on the discrete encoding of the output text to obtain the target output text.
[0238] As shown in Figure 15, a speech synthesis model 1500 of an embodiment of the present disclosure is also shown, which may include a decoder 1510. The speech synthesis model 1500 can be trained by the speech synthesis model training method in any embodiment of the present disclosure. The speech synthesis model has a first input port input1 configured to receive discrete encoding of text and a second input port input2 configured to receive discrete encoding of speech.
[0239] As shown in FIG16 , a speech recognition model 1600 according to an embodiment of the present disclosure is also shown, which may include a decoder 1610. The speech synthesis model 1600 may be trained using the speech recognition model training method according to any embodiment of the present disclosure. The speech recognition model 1600 has a first input port input1 configured to receive discrete speech codes and a second input port inout2 configured to receive discrete text codes. Furthermore, the features of the apparatus, model, system, and components of the embodiments of the present disclosure may refer to the features of the methods, steps, and other aspects of the embodiments of the present disclosure. Furthermore, the apparatus, model, and system embodiments may combine the features of the method embodiments to obtain new embodiments, and vice versa, and these will not be repeated here.
[0240] In an embodiment of the present disclosure, an electronic device is provided, comprising: a processor and a memory storing a computer program, wherein the processor is configured to execute the speech synthesis model training method of any embodiment and the speech synthesis method of any embodiment when running the computer program.
[0241] FIG17 shows a schematic diagram of an electronic device 1700 that can implement a method or implement an embodiment of the present disclosure. In some embodiments, the method may include more or fewer electronic devices than shown. In some embodiments, the method may be implemented using a single or multiple electronic devices. In some embodiments, the method may be implemented using cloud-based or distributed electronic devices.
[0242] FIG17 shows a schematic diagram of an electronic device 1700 that can be used to implement the method or implement the embodiments of the present disclosure. In some embodiments, the number of electronic devices may be more or less than the number shown. In some embodiments, a single or multiple electronic devices may be used for implementation. In some embodiments, cloud-based or distributed electronic devices may also be used for implementation.
[0243] As shown in Figure 17, the electronic device 1700 includes a processor 1710 and a memory 1720. The processor is configured to execute programs stored in the memory, which can implement the methods, steps or functions described in the above embodiments when executed by a computer. The processor 1710 may include various types of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), a neural network processor (NPU), a digital signal processor (DSP), etc. The processor 1710 and the memory 1720 are interconnected via a bus 1730. An input / output (I / O) interface and the like can also be connected to the bus 1730.
[0244] The systems, devices, modules, or units described in the above embodiments may be implemented by a computer or its associated components. The computer may be, for example, a mobile terminal, a smartphone, a personal computer, a laptop computer, an in-vehicle human-computer interaction device, a personal digital assistant, a media player, a navigation device, a game console, a tablet computer, a wearable device, a smart TV, an Internet of Things system, a smart home, an industrial computer, a server, or a combination thereof.
[0245] Although not shown, in an embodiment of the present disclosure, a storage medium is provided, the storage medium storing a computer program, and the computer program is configured to perform the method of any embodiment when executed.
[0246] Storage media in embodiments of the present disclosure include permanent and non-permanent, removable and non-removable items that can be implemented by any method or technology to store information. Examples of storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmitting medium that can be configured to store information that can be accessed by a computing device.
[0247] The methods, programs, systems, and apparatuses of the embodiments of the present disclosure may be executed or implemented in a single or multiple networked computers, or may be practiced in a distributed computing environment. In the embodiments of this specification, in these distributed computing environments, tasks may be performed by remote processing devices connected via a communication network.
[0248] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, those skilled in the art will appreciate that the functional modules / units or controllers and related method steps described in the above embodiments may be implemented using software, hardware, or a combination of software / hardware.
[0249] Unless explicitly stated, the actions or steps of the methods, programs, and embodiments of the present disclosure do not have to be performed in a specific order and can still achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0250] In this document, multiple embodiments of the present disclosure are described, but for the sake of brevity, the description of each embodiment is not exhaustive, and the same or similar features or parts between the embodiments may be omitted. In this document, "one embodiment", "some embodiments", "example", "specific example", or "some examples" are intended to apply to at least one embodiment or example according to the present disclosure, not all embodiments. The above terms do not necessarily mean to refer to the same embodiment or example. Those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples without conflicting feelings.
[0251] While the exemplary systems and methods of the present disclosure have been specifically shown and described with reference to the foregoing embodiments, these are merely examples of the best modes for implementing the present systems and methods. Those skilled in the art will appreciate that various changes may be made to the embodiments of the systems and methods described herein when implementing the present systems and / or methods without departing from the spirit and scope of the present disclosure as defined in the appended claims. Industrial Applicability
[0252] The embodiments of the present disclosure provide a speech synthesis and speech recognition method, training method, device, electronic device, and storage medium, which can train an autoregressive model by using a mixture of a large amount of unsupervised data and supervised data, thereby improving data utilization, avoiding multi-stage training of the model, and reducing the difficulty of model training.< / text> < / audio> < / audio> < / text> < / audio> < / text> < / audio> < / text>
Claims
1. A method for training a speech synthesis model, characterized in that, It includes the following steps: Obtain a training data set, where the training data set includes multiple data entries, and the types of the data entries include pure text data entries, pure audio data entries, and text-audio pair data entries; Select multiple data entries from the training data set to generate multiple batch data entry sets. Among them, each batch data entry set contains pure text data entries, pure audio data entries, and text-audio pair data entries, and the ratio among the pure text data entries, pure audio data entries, and text-audio pair data entries in the batch data entry set meets the set ratio condition; Perform discrete encoding processing on the data entries in the batch data entry set to generate multiple data entry discrete encodings. Each data entry discrete encoding includes a text discrete encoding and a voice discrete encoding, and the text discrete encoding in the data entry discrete encoding is located before the voice discrete encoding; Train an autoregressive model according to the multiple data entry discrete encodings in the multiple batch data entry sets to generate a target speech synthesis model.
2. The method according to claim 1, wherein The text discrete encoding includes text discrete encoding content, and the voice discrete encoding includes voice discrete encoding content; performing discrete encoding processing on the data entries in the batch data entry set to generate multiple data entry discrete encodings includes: In response to the data entry being a pure text data entry, perform discrete encoding processing on the text content in the pure text data entry to obtain the text discrete encoding content, and set the voice discrete encoding content to be empty; In response to the data entry being a pure audio data entry, perform discrete encoding processing on the audio content in the pure audio data entry to obtain the voice discrete encoding content, and set the text discrete encoding content to be empty; In response to the data entry being a text-audio pair data entry, perform discrete encoding processing on the text content in the text-audio pair data entry to obtain the text discrete encoding content, and perform discrete encoding processing on the audio content in the text-audio pair data entry to obtain the voice discrete encoding content.
3. The method according to claim 2, characterized in that, The text discrete encoding further includes a text discrete encoding start flag and a text discrete encoding end flag, and the text discrete encoding content is located between the text discrete encoding start flag and the text discrete encoding end flag; the voice discrete encoding further includes a voice discrete encoding start flag and a voice discrete encoding end flag, and the voice discrete encoding content is located between the voice discrete encoding start flag and the voice discrete encoding end flag.
4. The method according to any one of claims 2-3, characterized in that, Performing discrete encoding processing on the text content to obtain the text discrete encoding content includes: Perform phoneme conversion on the text content to generate the text discrete encoding content.
5. The method according to any one of claims 2 to 4, characterized in that, Performing discrete encoding processing on the audio content to obtain the voice discrete encoding content includes: Input the audio content into an audio encoder to obtain the voice discrete encoding content.
6. A voice synthesis method, characterized in that, It includes the following steps: Obtain input data, where the input data includes a target text; Perform discrete encoding processing on the input data to obtain an input data discrete encoding, where the input data discrete encoding includes a text discrete encoding and a voice discrete encoding, and the text discrete encoding is located before the voice discrete encoding; Discretely encode the input data and input it into the target speech synthesis model to obtain the output speech discrete encoding; Input the output speech discrete encoding into a speech decoder for decoding to obtain the target output audio.
7. The method according to claim 6, characterized in that, The input data further includes a target audio segment with a set timbre and ambient sound.
8. The method according to any one of claims 6-7, characterized in that The text discrete encoding includes text discrete encoding content, and the speech discrete encoding includes speech discrete encoding content; When the input data is target text, the discrete encoding process of the input data to obtain the input data discrete encoding includes: Perform a discrete encoding process on the target text to obtain the text discrete encoding content, and set the speech discrete encoding content to empty; When the input data is target text and a target audio segment, the discrete encoding process of the input data to obtain the input data discrete encoding includes: Perform a discrete encoding process on the target text to obtain the text discrete encoding content, and perform a discrete encoding process on the target audio segment to obtain the speech discrete encoding content.
9. The method according to any one of claims 6-8, wherein When the input data is target text, the step of inputting the discrete encoding of the input data into the target speech synthesis model to obtain the output speech discrete encoding includes: Input the discrete encoding of the input data into the target speech synthesis model, and after decoding and calculation, obtain the output speech discrete encoding; When the input data is target text and a target audio segment, the step of inputting the discrete encoding of the input data into the target speech synthesis model to obtain the output speech discrete encoding includes: Input the text discrete encoding content into the discrete encoding of the input data into the target speech synthesis model, and input the speech discrete encoding content into the target speech synthesis model in the form of a prompt word. After decoding and calculation, obtain the output speech discrete encoding.
10. The method according to any one of claims 6-9, characterized in that, The target speech synthesis model is trained by the method according to any one of claims 1-5.
11. A method for training a speech recognition model, characterized in that, It includes the following steps: Obtain a training data set, the training data set includes multiple data entries, and the types of the data entries include pure text data entries, pure audio data entries, and text-audio pair data entries; Select multiple data entries from the training data set to generate multiple batch data entry sets. Among them, each batch data entry set contains pure text data entries, pure audio data entries, and text-audio pair data entries, and the ratio between the pure text data entries, pure audio data entries, and text-audio pair data entries in the batch data entry set satisfies the set ratio condition; Perform a discrete encoding process on the data entries in the batch data entry set to generate multiple data entry discrete encodings. Each data entry discrete encoding includes text discrete encoding and speech discrete encoding, and the speech discrete encoding in the data entry discrete encoding is located before the text discrete encoding; Train an autoregressive model according to the multiple data entry discrete encodings in the multiple batch data entry sets to generate a speech recognition model.
12. The method according to claim 11, wherein The text discrete encoding includes text discrete encoding content, and the speech discrete encoding includes speech discrete encoding content; the discrete encoding process for the data entries in the batch processing data entry set to generate multiple data entry discrete encodings includes: In response to the data entry being a pure text data entry, performing discrete encoding processing on the text content in the pure text data entry to obtain the text discrete encoding content, and setting the speech discrete encoding content to be empty; In response to the data entry being a pure audio data entry, performing discrete encoding processing on the audio content in the pure audio data entry to obtain the speech discrete encoding content, and setting the text discrete encoding content to be empty; In response to the data entry being a text-audio pair data entry, performing discrete encoding processing on the text content in the text-audio pair data entry to obtain the text discrete encoding content, and performing discrete encoding processing on the audio content in the text-audio pair data entry to obtain the speech discrete encoding content.
13. The method according to claim 12, characterized in that, The text discrete encoding further includes a text discrete encoding start flag and a text discrete encoding end flag, and the text discrete encoding content is located between the text discrete encoding start flag and the text discrete encoding end flag; the speech discrete encoding further includes a speech discrete encoding start flag and a speech discrete encoding end flag, and the speech discrete encoding content is located between the speech discrete encoding start flag and the speech discrete encoding end flag.
14. The method according to any one of claims 12-13, characterized in that, Performing discrete encoding processing on the text content to obtain the text discrete encoding content includes: Performing phoneme conversion on the text content to generate the text discrete encoding content.
15. The method according to any one of claims 12 - 14, characterized in that, Performing discrete encoding processing on the audio content to obtain the speech discrete encoding content includes: Inputting the audio content into an audio encoder to obtain the speech discrete encoding content.
16. A voice recognition method, characterized in that, Including the following steps: Obtaining input data, where the input data includes target audio; Performing discrete encoding processing on the input data to obtain input data discrete encoding, where the input data discrete encoding includes text discrete encoding and speech discrete encoding, and the speech discrete encoding is located before the text discrete encoding; Inputting the input data discrete encoding into a target speech recognition model to obtain output text discrete encoding; Performing inverse discrete encoding processing on the output text discrete encoding to obtain a target output text.
17. The method according to claim 16, characterized in that, The text discrete encoding includes text discrete encoding content, and the speech discrete encoding includes speech discrete encoding content; The discrete encoding processing for the input data to obtain input data discrete encoding includes: Performing discrete encoding processing on the target audio to obtain the speech discrete encoding content, and setting the text discrete encoding content to be empty.
18. The method according to any one of claims 16 - 17, characterized in that, The inputting the input data discrete encoding into a target speech recognition model to obtain output text discrete encoding includes: Inputting the input data discrete encoding into a target speech recognition model, and after decoding and calculating, obtaining the output text discrete encoding.
19. The method according to any one of claims 16 - 18, characterized in that, The target speech recognition model is trained by the method according to any one of claims 11 to 15.
20. A voice synthesis model training device, characterized in that, Including a training data set acquisition module, a batch processing data integration module, a discrete encoding processing module, and a training module, where, The training dataset acquisition module is configured to acquire a training dataset, where the training dataset includes multiple data entries, and the types of the data entries include pure text data entries, pure audio data entries, and text-audio pair data entries; The batch data integration module is configured to select multiple data entries from the training dataset to generate multiple batch data entry sets. Among them, each batch data entry set contains pure text data entries, pure audio data entries, and text-audio pair data entries, and the ratio among the pure text data entries, pure audio data entries, and text-audio pair data entries in the batch data entry set meets the set ratio condition; The discrete coding processing module is configured to perform discrete coding processing on the data entries in the batch data entry set to generate multiple data entry discrete codings. Each data entry discrete coding includes a text discrete coding and a speech discrete coding, and the text discrete coding in the data entry discrete coding is located before the speech discrete coding; The training module trains an autoregressive model according to the multiple data entry discrete codings in the multiple batch data entry sets to generate a target speech synthesis model.
21. A speech synthesis model, characterized in that, The speech synthesis model is trained by the method according to any one of claims 1 to 5.
22. A voice synthesis device, characterized in that, It includes an input data acquisition module, a discrete coding processing module, a model calculation module, and a speech decoder module, where, The input data acquisition module is configured to acquire input data, and the input data includes a target text; The discrete coding processing module is configured to perform discrete coding processing on the input data to obtain an input data discrete coding. The input data discrete coding includes a text discrete coding and a speech discrete coding, and the text discrete coding is located before the speech discrete coding; The model calculation module is configured to input the input data discrete coding into the target speech synthesis model to obtain an output speech discrete coding; The speech decoder module is configured to perform speech decoding on the output speech discrete coding to obtain a target output audio.
23. A voice recognition model training device, characterized in that It includes a training dataset acquisition module, a batch data integration module, a discrete coding processing module, and a training module, where, The training dataset acquisition module is configured to acquire a training dataset, where the training dataset includes multiple data entries, and the types of the data entries include pure text data entries, pure audio data entries, and text-audio pair data entries; The batch data integration module is configured to select multiple data entries from the training dataset to generate multiple batch data entry sets. Among them, each batch data entry set contains pure text data entries, pure audio data entries, and text-audio pair data entries, and the ratio among the pure text data entries, pure audio data entries, and text-audio pair data entries in the batch data entry set meets the set ratio condition; The discrete coding processing module is configured to perform discrete coding processing on the data entries in the batch data entry set, generating multiple discrete codings of data entries, where each discrete coding of data entry includes a text discrete coding and a voice discrete coding, and in the discrete coding of data entry, the voice discrete coding is located before the text discrete coding; The training module is configured to train an autoregressive model according to the multiple discrete codings of data entries in the multiple batch data entry sets, generating a speech recognition model.
24. A voice recognition model, characterized in that, The speech recognition model is trained by the method according to any one of claims 11 to 15.
25. A voice recognition device, characterized in that, It includes an input data acquisition module, a discrete coding processing module, a model calculation module, and an inverse discrete coding processing module, where, The input data acquisition module is configured to acquire input data, and the input data includes a target audio; The discrete coding processing module is configured to perform discrete coding processing on the input data, acquiring a discrete coding of input data, The discrete coding of input data includes a text discrete coding and a voice discrete coding, and the voice discrete coding is located before the text discrete coding; The model calculation module is configured to input the discrete coding of input data into a target speech recognition model, acquiring an output text discrete coding; The inverse discrete coding module is configured to perform inverse discrete coding processing on the output text discrete coding, acquiring a target output text.
26. An electronic device, characterized in that, It includes a processor and a memory storing a computer program, and the processor is configured to execute the method according to any one of claims 1 to 19 when running the computer program.
27. A storage medium, characterized in that, The storage medium stores a computer program, and the computer program is configured to execute the method according to any one of claims 1 to 19 when being run.
Citation Information
Patent Citations
Voice synthesis method and device based on attention mechanism
CN109767752A
Speech recognition method and device, equipment and storage medium
CN115512695A
Speech recognition method and device based on pre-training feature representation and electronic equipment
CN115831102A
Speech synthesis method, speech synthesis device, electronic equipment and storage medium
CN116343747A
Speech synthesis method and device, speech recognition method and device, training method and device, electronic equipment and storage medium
CN117953857A