Sound cloning model training method, sound cloning method and product

By preprocessing the audio data and training the sound cloning model, the existing zero-sample sound cloning model in stream generation is solved, and high-quality streaming sound cloning effect is achieved.

CN119920232AActive Publication Date: 2025-05-02CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

Patent Information

Application Number
CN202411975949.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-02
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

The existing zero-sample sound cloning model has problems in streaming generation. It requires synthesis of complete sentence audio to output, and high-quality streaming sound cloning cannot be achieved.

Method used

By preprocessing the audio data, a training sample set is constructed, including text sequence samples, Mel spectral samples and target Mel spectral tags, and a preset sound cloning model is trained. The model contains text encoder and decoder. The decoder adopts a Transformer structure based on U-net form, supporting non-stream and streaming computing modes.

Benefits of technology

A high-quality streaming sound cloning model is realized, which can output the predicted Mel spectrum corresponding to the part of the input when receiving an input that meets the block size, without waiting for all inputs to be received.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920232A_ABST
    Figure CN119920232A_ABST
Patent Text Reader

Abstract

The invention provides a training method of a sound cloning model, a sound cloning method and a product, and belongs to the technical field of speech synthesis. The method comprises the steps that multiple pieces of acquired audio data are preprocessed, a training sample set is constructed, a preset sound clone model is trained according to the training sample set and a preset block size, a text encoder in the sound clone model is used for extracting a text vector according to a text sequence sample of any training sample, and the text vector is used for generating a voice sequence; the decoder is used for outputting a predicted Mel spectrum according to the text vector of any training sample, the Mel spectrum sample and the speaker vector, and the block size is used for adjusting the calculation mode of the sound clone model, and the calculation mode comprises a non-streaming mode and a streaming mode; and performing iterative training on a preset sound clone model according to the target Mel spectrum label of any training sample, the predicted Mel spectrum and the loss function of the optimal transmission stream. The invention aims to provide a streaming sound clone model capable of realizing high quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of speech synthesis, and in particular, to a training method for a voice cloning model, a voice cloning method, and a product. Background Art

[0002] With the development of technology, speech synthesis technology, namely TTS technology, plays an increasingly important role in daily life, and has a significant impact on improving efficiency and convenience. For example, TTS technology is widely used in customer service systems to provide 24-hour uninterrupted service. TTS technology also plays an important role in human-computer interaction in smartphones or smart home systems. Users' demand for voice systems is not limited to providing clear and natural voice output, but also includes more diverse functions, an important one of which is high-quality zero-sample voice cloning capability.

[0003] Zero-shot voice cloning, also known as zero-shot speaker adaptation or zero-shot TTS, aims to synthesize the voice of a speaker that did not appear during the training process based on any reference speech.

[0004] In order to achieve better zero-sample cloning effects, there has been research on the application of large language models in the TTS field and the application of diffusion models based on flow matching technology in the TTS field. However, the speech models obtained based on the current methods do not solve the problem of streaming generation and require the synthesis of complete sentence audio for output.

[0005] Therefore, there is an urgent need for a streaming sound cloning model with better performance to achieve high-quality. Summary of the invention

[0006] The embodiments of the present application provide a training method for a sound cloning model, a sound cloning method, and a product, aiming to provide a high-quality streaming sound cloning model.

[0007] In a first aspect, an embodiment of the present application provides a method for training a sound cloning model, the method comprising: Preprocessing the acquired multiple audio data to construct a training sample set, wherein any training sample in the training sample set includes a text sequence sample, a Mel spectrum sample, and a target Mel spectrum label corresponding to any audio data; According to the training sample set and the preset block size, a preset voice cloning model is trained, the text encoder in the preset voice cloning model is used to extract a text vector according to a text sequence sample of any training sample, the decoder in the preset voice cloning model is used to output a predicted mel spectrum according to the text vector, mel spectrum sample and speaker vector of any training sample, and the block size is used to adjust the calculation mode of the preset voice cloning model to non-streaming or streaming; According to the target mel spectrum label, predicted mel spectrum and the loss function of the optimal transmission stream of any training sample, the preset sound cloning model is iteratively trained to obtain a trained sound cloning model.

[0008] Optionally, the obtained multiple audio data are preprocessed to construct a training sample set, including: For any audio data obtained, generate a text sequence and Mel spectrum corresponding to the audio data; Perform a masking operation on the Mel spectrum corresponding to any audio data to obtain a masked Mel spectrum and an unmasked Mel spectrum, use the unmasked Mel spectrum as a Mel spectrum sample, and use the masked Mel spectrum as a target Mel spectrum label; Insert characters into the text sequence corresponding to the target Mel spectrum label in the audio data so that the length of the text sequence corresponding to the target Mel spectrum label is consistent with the number of frames of the target Mel spectrum label, and use the text sequence corresponding to the Mel spectrum sample as the text sequence sample; The text sequence sample, Mel spectrum sample and target Mel spectrum label corresponding to any audio data are taken as a training sample.

[0009] Optionally, performing a mask operation on the Mel spectrum corresponding to any audio data includes: For any audio data, a mask operation is performed on a preset mask range of the Mel spectrum corresponding to the audio data, or a mask operation is performed on all Mel spectrums corresponding to the audio data.

[0010] Optionally, the method further comprises: Build preset sound cloning models; Among them, the preset sound cloning model includes a text encoder and a decoder, the text encoder includes a Conformer layer for extracting a text vector of any text sequence sample; the decoder is based on a Transformer structure in the form of U-net, including a downsampling module, an intermediate module and an upsampling module, any module among the downsampling module, the intermediate module and the upsampling module includes a residual module based on a convolutional layer and a Transformer module based on an attention mechanism, and any convolutional layer in the preset sound cloning model is a block convolution, and the attention mechanism of the Transformer module is an attention mechanism based on a block mask.

[0011] Optionally, the method further comprises: In response to a preset operation, determining a range of variation of a block size of a preset sound cloning model during training, wherein the block size is used to adjust a calculation mode of the preset sound cloning model to a non-streaming mode or a streaming mode; In response to the training strategy selection operation, in the process of training a preset sound cloning model based on the training sample set, executing a first training strategy or a second training strategy; Among them, when executing the first training strategy, in each iterative training process of the preset sound cloning model, a block size is randomly selected from the variation range of the block size as the block size of this iteration process; when executing the second training strategy, the loss function in the non-streaming calculation mode and the loss function in the streaming calculation mode are determined respectively, and in the back propagation of the preset sound cloning model, the model parameters of the preset sound cloning model are updated based on the loss function in the non-streaming calculation mode and the loss function in the streaming calculation mode at the same time.

[0012] In a second aspect, an embodiment of the present application provides a sound cloning method, the method comprising: Acquire a target prompt speech and a text sequence to be synthesized, and generate a prompt mel spectrum and a prompt text sequence of the target prompt speech; Preprocessing to obtain a text sequence to be input according to the text sequence to be synthesized and the prompt text sequence; Inputting the prompt mel-spectrogram of the target prompt speech and the text sequence to be input into a sound cloning model, and outputting the mel-spectrogram result corresponding to the text sequence to be synthesized through the sound cloning model, wherein the sound cloning model is obtained by the sound cloning model training method described in the first aspect of the embodiment; According to the mel-spectrogram result, audio data corresponding to the text sequence to be input is converted.

[0013] Optionally, preprocessing to obtain a text sequence to be input according to the text sequence to be synthesized and the prompt text sequence includes: According to the text sequence to be synthesized, estimating the number of frames of the mel spectrum result corresponding to the text sequence to be synthesized; Characters are inserted into a text sequence composed of the text sequence to be synthesized and the prompt text sequence to obtain a text sequence to be input, wherein the length of the text sequence to be input is equal to the sum of the number of frames of the prompt mel-spectrogram and the number of frames of the mel-spectrogram result corresponding to the text sequence to be synthesized.

[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the training method of the sound cloning model as described in the first aspect of the embodiment, or implement the sound cloning method as described in the second aspect of the embodiment.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program / instruction stored thereon. When the computer program / instruction is executed by a processor, the training method of the sound cloning model as described in the first aspect of the embodiment is implemented, or the sound cloning method as described in the second aspect of the embodiment is implemented.

[0016] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the training method of the sound cloning model as described in the first aspect, or implements the sound cloning method as described in the second aspect.

[0017] Beneficial effects: When training the sound cloning model, the obtained multiple audio data are preprocessed to construct a training sample set, wherein any training sample in the training sample set includes a text sequence sample, a Mel spectrum sample and a target Mel spectrum label corresponding to any audio data; according to the training sample set and a preset block size, the preset sound cloning model is trained, the text encoder in the preset sound cloning model is used to extract a text vector according to the text sequence sample of any training sample, and the decoder in the preset sound cloning model is used to output a predicted Mel spectrum according to the text vector, the Mel spectrum sample and the speaker vector of any training sample, and the block size is used to adjust the calculation mode of the preset sound cloning model to non-streaming or streaming; according to the target Mel spectrum label of any training sample, the predicted Mel spectrum and the loss function of the optimal transmission stream, the preset sound cloning model is iteratively trained to obtain a trained sound cloning model.

[0018] The sound cloning model trained by the method can output the predicted Mel spectrum corresponding to the part of the input in time when the sound cloning model receives the input that meets the block size, without having to wait until all the inputs are received before outputting the overall predicted Mel spectrum, thereby providing a high-quality streaming sound cloning model. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0020] Figure 1 is a flowchart of the steps of a method for training a sound cloning model proposed in one embodiment of the present application; Figure 2 is a training diagram of a sound cloning model proposed in an embodiment of the present application; Figure 3 is a schematic diagram of the structure of a decoder provided in an embodiment of the present application; Figure 4 is a schematic diagram of a block convolution proposed in an embodiment of the present application; Figure 5 is a schematic diagram of a chunk mask proposed in an embodiment of the present application; Figure 6 is a flowchart of the steps of a sound cloning method provided by an embodiment of the present application; Figure 7 is a reasoning diagram of a sound cloning model proposed in an embodiment of the present application; Figure 8 It is a functional module diagram of a training device for a sound cloning model provided in one embodiment of the present application; Fig. 9 1 is a functional module diagram of a sound cloning device according to an embodiment of the present application. DETAILED DESCRIPTION

[0021] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0022] TTS: text to speech, speech synthesis; Flow matching technology: a training method for continuous normalized flow models, which trains the model by learning the vector field associated with the conditional probability path; it can be used to train diffusion models, and can be combined with optimal transmission technology to further speed up the training speed. During inference, ordinary differential equations can be used to solve the generation of new samples; Mel spectrum: mel spectrogram, mel spectrum; mel spectrum refers to the spectrum converted from frequency to mel scale. Studies have shown that the human ear's perception of frequency is not linear. The mel scale is a nonlinear scale unit defined based on frequency, which indicates the equidistant changes in pitch to the human ear. ‌Encodec‌ is a neural network-based audio coding scheme that aims to achieve efficient and high-quality audio compression. By using AutoEncoder and Residual Vector Quantization (RVQ) technology, Encodec can significantly reduce the size of audio files while maintaining high sound quality. Matcha-TTS is an efficient text-to-speech (TTS) architecture based on conditional stream matching. Its main goal is to provide an efficient non-autoregressive neural network TTS model that uses conditional stream matching technology to accelerate speech synthesis based on ordinary differential equations (ODE).

[0023] With the development of technology, speech synthesis technology, namely TTS technology, plays an increasingly important role in daily life, and has a significant impact on improving efficiency and convenience. For example, TTS technology is widely used in customer service systems to provide 24-hour uninterrupted service. TTS technology also plays an important role in human-computer interaction in smartphones or smart home systems.

[0024] With the rapid development of information technology, users' demands for voice systems are not limited to providing clear and natural voice output, but also include more diverse functions. One of the important ones is high-quality zero-sample voice cloning capability; zero-sample voice cloning, also known as zero-sample speaker adaptation or zero-shot speech synthesis (zero-shot TTS), aims to synthesize the voice of a speaker that did not appear in the training process based on any reference speech.

[0025] The traditional research route mainly explores how to use the data of the target speaker to fine-tune the overall or partial parameters of the existing multi-person TTS model. Although this type of fine-tuning method can achieve a certain synthesis effect, this method has some obvious disadvantages: additional training time is required for each speaker outside the training set, and the synthesis effect is subject to the length of the speaker's corpus. Corpus with less than 10 sentences will often seriously reduce the quality of the synthesized audio.

[0026] In order to achieve better zero-sample cloning effects, there is a trend of studying the application of large language models in the field of TTS. For example, by using a speech encoder Encodec to discretize speech, and then using a large language model to train the "next token prediction" task on large-scale audio data, the model is able to extract information such as the timbre and rhythm of the target speaker from a 3-second audio prompt, thereby completing zero-sample voice cloning. However, the synthesis effect of such models is very susceptible to the quality of speech discretization, and the large language model uses an autoregressive reasoning method, which can only generate sample points one by one, with a slow generation speed, and is often difficult to meet actual business needs. In comparison, non-autoregressive models can be modeled in a continuous vector space and use parallel processing to improve the reasoning speed, so they are more popular in actual business.

[0027] The diffusion model based on flow matching technology is one of the mainstream methods in the current non-autoregressive generative model. The diffusion model can capture and fit complex data distribution by controlling the process of adding noise and denoising, thereby generating high-quality data samples. For example, Matcha-TTS combines the diffusion model based on flow matching as a decoder with the text encoder, and exceeds the traditional speech synthesis model in terms of word error rate and subjective average score. It also proposes to use the diffusion model to model the quantized hidden vector of the discrete encoder, and inject the information of the reference speech into the model through the attention mechanism, so as to achieve zero-sample generation capability. The design of the mask mechanism of Voicebox enables the model to have the ability of contextual learning, and the overall design scheme is more concise and elegant.

[0028] However, all of these research directions need to rely on third-party duration alignment tools to obtain duration information at the phoneme level. Furthermore, E2 TTS proposed a simpler method, which is to insert specified characters into the character sequence input to the model, that is, insert filler tokens, so that the character sequence input to the model and the dimension of the target audio feature are consistent, and then let the model learn to align the text content and audio. However, none of the above models solves the problem of streaming generation, and requires the synthesis of complete sentence audio for output.

[0029] In summary, the current zero-sample sound cloning model still has different problems, and there is an urgent need for a sound cloning model with better performance that can achieve high-quality.

[0030] Reference Figure 1, shows a flowchart of the steps of a method for training a sound cloning model in an embodiment of the present application, and the method may specifically include the following steps: S101: Preprocessing the acquired multiple audio data to construct a training sample set, wherein any training sample in the training sample set includes a text sequence sample, a mel spectrum sample and a target mel spectrum label corresponding to any audio data.

[0031] Specifically, the process of preprocessing the acquired multiple audio data and constructing a training sample set includes: A1: For any acquired audio data, generate a text sequence and mel spectrum corresponding to the audio data.

[0032] First, for each piece of audio data, a text sequence in the audio data and a Mel-spectrogram of the audio data are generated.

[0033] A2: Perform a mask operation on the Mel spectrum corresponding to any audio data to obtain a masked Mel spectrum and an unmasked Mel spectrum. The unmasked Mel spectrum is used as a Mel spectrum sample, and the masked Mel spectrum is used as a target Mel spectrum label.

[0034] Specifically, for any audio data, a preset mask range of the Mel spectrum corresponding to the audio data is masked, or all Mel spectrums corresponding to the audio data are masked. The preset mask range can be customized according to the needs of the actual application. For example, the range of X1%~X2% of the Mel spectrum of an audio data can be masked.

[0035] For example, for each piece of audio data, 70% to 90% of its Mel spectrum can be masked with a probability of 50%, and the entire Mel spectrum of the audio data can be masked with a probability of 50% to conceal it.

[0036] In the actual implementation process, during the mask operation, a binary mask matrix is ​​generated and multiplied with the Mel spectrum of the audio data. The formula is as follows:

[0037] in, represents the Hadamard product, is the Mel spectrum of the audio data, is the Mel spectrum sample of the audio data, M is the binary mask matrix, , D and T are the dimensions of the mask matrix, and the dimension of the mask matrix is ​​consistent with the dimension of the Mel spectrum of the audio data.

[0038] During the training process of the sound cloning model, the loss function of the model only considers the loss of the masked part, that is, the model predicts the masked Mel spectrum based on the Mel spectrum sample, that is, the unmasked Mel spectrum, and iteratively trains the sound cloning model based on the loss of the predicted masked Mel spectrum and the target Mel spectrum label, so that the masked Mel spectrum predicted by the sound cloning model is closer to the target Mel spectrum label.

[0039] A3: Insert characters into the text sequence corresponding to the target Mel spectrum label in the audio data so that the length of the text sequence corresponding to the target Mel spectrum label is consistent with the number of frames of the target Mel spectrum label, and use the text sequence corresponding to the Mel spectrum sample as the text sequence sample.

[0040] The target mel spectrum label, that is, the mel spectrum part of the mask, needs to be predicted by the model. However, since the length of text content is much shorter than that of audio for speech, the length of the two often needs to be aligned in the TTS system.

[0041] In this method, in order to reduce the dependence on third-party time alignment tools, the length of the text sequence corresponding to the target Mel-spectrogram label can be made consistent with the number of frames of the target Mel-spectrogram label by inserting characters, that is, inserting filler token 〈F〉.

[0042] For example, assume that the text sequence corresponding to the target Mel spectrum label in the audio data is , then by inserting characters, the text sequence corresponding to the expanded target Mel spectrum label is obtained as follows:

[0043] in, Indicates filler token. The number of is equal to TM, where M is the number of characters in the text sequence corresponding to the target Mel spectrum label.

[0044] Thus, the text sequence corresponding to the expanded target Mel spectrum label is The length of is consistent with the number of frames T of the target Mel spectrum label.

[0045] In the actual implementation process, characters can be inserted at the end of the text sequence, and the inserted characters cannot conflict with the original characters in the text sequence. For example, if the text sequence includes text, punctuation characters can be inserted; if the text sequence includes text and punctuation, English words can be inserted, etc.

[0046] A4: Take the text sequence sample, Mel spectrum sample and target Mel spectrum label corresponding to any audio data as a training sample.

[0047] In actual implementation, text sequence samples can include two formats: In the first format, the text sequence samples consist of only text and punctuation marks, and the text part includes two languages, Chinese and English. For example, English can be represented by words, and Chinese can be represented by pinyin; In the second format, the text sequence samples are represented by characters, phonemes, and punctuation marks. That is, phoneme representations are added to the characters. Therefore, during the training process, a 15% probability is set to convert the characters into phoneme representations. Brackets are added before and after the phoneme representations as separators between phonemes and characters.

[0048] The purpose of setting two text sequence sample formats is to deal with the situation where some text has special pronunciation in actual business. If only the first format is used, that is, only represented by text and punctuation marks, then for some special texts in actual business, such as specific names, place names and abbreviations, it is difficult to intervene in the corresponding pronunciation, and the training data for special texts is often scarce or even non-existent, so the model's pronunciation prediction for this part will be strange or inconsistent with the context. Then, by introducing phoneme representation, a means of intervening in the pronunciation information in the model's synthesized audio is provided, thereby improving the model's adaptability to actual complex scenarios.

[0049] S102: According to the training sample set and the preset block size, a preset voice cloning model is trained, the text encoder in the preset voice cloning model is used to extract a text vector according to a text sequence sample of any training sample, the decoder in the preset voice cloning model is used to output a predicted Mel spectrum according to the text vector, Mel spectrum sample and speaker vector of any training sample, and the block size is used to adjust the calculation mode of the preset voice cloning model to non-streaming or streaming.

[0050] Reference Figure 2 , shows a training schematic diagram of the sound cloning model provided by an embodiment of the present application, and pre-builds a model structure of a preset sound cloning model, wherein the preset sound cloning model includes a text encoder and a decoder, and after the text sequence samples, mel spectrum samples and target mel spectrum labels of the training samples are input into the model, since the text encoder extracts a text vector according to the text sequence samples of any training sample, the decoder in the sound cloning model is used to output a predicted mel spectrum according to the text vector, mel spectrum sample and speaker vector of any training sample.

[0051] The text encoder includes a Conformer layer for extracting a text vector of any text sequence sample.

[0052] Reference Figure 3, showing a schematic diagram of the structure of a decoder provided in an embodiment of the present application, wherein the decoder is based on a Transformer structure in the form of a U-net, including down-sampling blocks, midblocks and up-sampling blocks, wherein any of the down-sampling modules, the midblocks and the up-sampling modules includes a residual module based on a convolutional layer and a Transformer module based on an attention mechanism.

[0053] The input of the decoder is the Mel spectrum sample with part of the Mel spectrum masked after the mask operation. , as well as the speaker vector V extracted by the voice cloning model based on the training sample, the text vector Y output by the text encoder, the time step t and the intermediate features corresponding to the time t .

[0054] In a feasible implementation, the loss function of the sound cloning model may be constructed based on an optimal-transport conditional flow matching model (OT-CFM).

[0055] OT-CFM is an extended form of flow matching technology. The purpose of OT-CFM is to use a prior distribution To predict the data distribution of the target Mel spectrum label .

[0056] First, use the time-dependent vector field To define the probability density path, , the flow can be generated by the following ordinary differential equation : (1) in, , , the prior distribution Normal distribution By solving the initial value problem of equation (1), the target distribution Can be used to approximate and sample from it.

[0057] To fit the vector field , construct a loss function based on the optimal transmission flow to train the sound cloning model:

[0058] in:

[0059] .

[0060] Furthermore, in order to support both non-streaming and streaming reasoning for the trained sound cloning model, in this embodiment, the sound cloning model adopts chunk-based streaming computing, which requires the model to calculate the corresponding output when receiving input that satisfies a chunk size, without having to wait until all inputs are received before performing the calculation.

[0061] Therefore, in the sound cloning model structure of the present embodiment, any convolution layer in the preset sound cloning model is set to a chunk convolution, and the attention mechanism of the Transformer module is a chunk mask attention mechanism.

[0062] Reference Figure 4 , shows a schematic diagram of the block convolution provided in this embodiment. There are usually three ways of convolution calculation: conventional convolution, causal convolution and block convolution.

[0063] Assuming the chunk size is 7, the sequence corresponding to the first chunk is , the convolution kernel size is 5, the convolution step size is 1, then the conventional convolution will fill two 0s at the beginning of the sequence, and it is necessary to use This obviously does not satisfy streaming computing, because in streaming computing Information belonging to the next chunk.

[0064] The causal convolution will first fill 4 zeros at the beginning of the sequence to ensure that the output corresponding to each time point only sees the information of the current moment and the past moment, and does not use the information of the future moment. Although this meets the requirements of streaming computing, it will obviously reduce the effect of the model.

[0065] The difference between block convolution and regular convolution is that two zeros are filled at the end of the sequence, while the rest remains unchanged. This ensures that the model does not use future information when calculating each chunk, while at the same time, information about future time points can still be seen at each time point in a chunk.

[0066] Therefore, the sound cloning model of this embodiment adopts chunk-based streaming computing in structure, so that it can support both non-streaming and streaming reasoning.

[0067] Furthermore, in the attention-based Transformer layer of the sound cloning model, a chunkmask is used to control the input of the sound cloning model.

[0068] Reference Figure 5 , shows a schematic diagram of the chunk mask provided in this embodiment. A binary mask matrix is ​​generated as the chunk mask according to the chunk size. The value is 1 only within each chunk range and 0 otherwise. When the chunk size is equal to the total length of the input sequence, the chunk mask is equivalent to the full mask, that is, a matrix of all 1s. If the chunk size is 2, then Figure 5 As shown in the chunk mask.

[0069] Since the block-based streaming computing allows the sound cloning model to adjust the computing mode to non-streaming or streaming by setting the chunk size, two training strategies are also adopted in the training phase.

[0070] For example, in response to a preset operation, a variation range of a block size of a preset sound cloning model during training is determined, where the block size is used to adjust a calculation mode of the preset sound cloning model to a non-streaming mode or a streaming mode.

[0071] Then, in response to the training strategy selection operation, in the process of training the preset sound cloning model based on the training sample set, the first training strategy or the second training strategy is executed; When executing the first training strategy, during each iterative training process of the preset sound cloning model, a block size is randomly selected from the range of block sizes as the block size of this iterative process, so that the chunk size of the sound cloning model changes dynamically within a certain range during the training process. In this way, the model can adapt to chunk sizes of different sizes during inference, thereby improving the wide adaptability of the sound cloning model in streaming computing.

[0072] When executing the second training strategy, the distillation loss function is used to guide the streaming calculation results with the non-streaming calculation results. During the training process, when the model is calculated in streaming mode, in addition to calculating the loss function between the accurate result, it also calculates the loss function between the model and the result calculated in non-streaming mode.

[0073] Specifically, the loss function in the non-streaming computing mode and the loss function in the streaming computing mode are respectively determined, and in the back propagation of the preset sound cloning model, the model parameters of the preset sound cloning model are updated based on the loss function in the non-streaming computing mode and the loss function in the streaming computing mode.

[0074] S103: Iteratively train the preset sound cloning model according to the target mel spectrum label, the predicted mel spectrum and the loss function of the optimal transmission stream of any training sample to obtain a trained sound cloning model.

[0075] The voice cloning model trained by this method has the ability of zero-sample voice cloning. That is, during training, the mask design is used to allow the model to learn to predict the masked part based on the unmasked Mel spectrum, and the model is trained with large-scale data, so that the model has strong contextual learning ability and can achieve zero-sample voice reproduction for speakers outside the set.

[0076] In addition, flow matching technology is used to train the diffusion process to achieve refined control of data distribution, thereby improving the quality of synthesized speech. Based on the model's strong contextual learning ability, speaker vectors are introduced to further improve the timbre similarity of voice cloning.

[0077] The model structure and reasoning method are constructed based on the block size calculation method, so that the model can achieve non-streaming generation or streaming generation by adjusting the chunk size. The reasoning method is more flexible and can effectively reduce the first packet delay and response time.

[0078] Reference Figure 6 , shows a flow chart of a sound cloning method provided by an embodiment of the present application, the method comprising the following steps: S201: Acquire a target prompt speech and a text sequence to be synthesized, and generate a prompt mel spectrum and a prompt text sequence of the target prompt speech.

[0079] S202: Preprocessing to obtain a text sequence to be input according to the text sequence to be synthesized and the prompt text sequence.

[0080] Assume that the target prompt speech's prompt Mel spectrum is denoted by , the prompt text sequence corresponding to the target prompt voice is , the text sequence to be synthesized is .

[0081] Specifically, the process of preprocessing to obtain the text sequence to be input according to the text sequence to be synthesized and the prompt text sequence includes the following steps: B1: according to the text sequence to be synthesized, estimate the number of frames of the mel spectrum result corresponding to the text sequence to be synthesized.

[0082] Since the number of frames of the mel-spectrogram results corresponding to the text sequence to be synthesized is unknown, it can be estimated by the average duration corresponding to each text in the training set and the prompt speech. The estimation method can be selected according to the needs of the actual application.

[0083] For example, if the text sequence with a length of L1 corresponds to a Mel-spectrogram frame number of Z1, the Mel-spectrogram frame number Z corresponding to the text sequence of unit length can be calculated, and thus the number of frames of the Mel-spectrogram result corresponding to the text sequence to be synthesized can be calculated based on the length of the text sequence to be synthesized and the Mel-spectrogram frame number Z corresponding to the text sequence of unit length.

[0084] It is also possible to make an estimation based on the number of Mel-spectrogram frames of each text in the text sequence. For example, if the text sequence to be synthesized includes the text "weather", and the number of Mel-spectrogram frames corresponding to the text "weather" in the training set or the prompt voice is Z2, then the corresponding number of Mel-spectrogram frames for the text sequence to be synthesized including the text "weather" can be determined as Z2, thereby determining the number of Mel-spectrogram frames of each text in the text sequence to be synthesized, and taking the number of Mel-spectrogram frames of each text as the number of frames of the Mel-spectrogram result corresponding to the text sequence to be synthesized.

[0085] It is worth noting that when estimating the number of frames of the Mel spectrum result corresponding to the text sequence to be synthesized, it must be satisfied ,in, The number of frames representing the target prompt speech's prompt mel-spectrogram. Represents the number of Mel spectrum frames corresponding to the text sequence to be synthesized, Represents the length of the prompt text sequence, Represents the length of the text sequence to be synthesized.

[0086] B2: inserting characters into the text sequence composed of the text sequence to be synthesized and the prompt text sequence to obtain a text sequence to be input, wherein the length of the text sequence to be input is equal to the sum of the number of frames of the prompt mel-spectrogram and the number of frames of the mel-spectrogram result corresponding to the text sequence to be synthesized.

[0087] Specifically, a filler token can be inserted at the end of a text sequence. , and the inserted characters cannot conflict with the original characters in the text sequence. For example, if the text sequence includes text, punctuation characters can be inserted; if the text sequence includes text and punctuation, English words can be inserted, etc.

[0088] By inserting characters, the text sequence to be input is:

[0089] in, The number is equal to .

[0090] Thus, the length of the text sequence to be synthesized is approximated to the number of frames of the Mel spectrum result corresponding to the text sequence to be synthesized, so that the text length is aligned with the audio length.

[0091] S203: Inputting the prompt mel-spectrogram of the target prompt speech and the text sequence to be input into the sound cloning model, and outputting the mel-spectrogram result corresponding to the text sequence to be synthesized through the sound cloning model.

[0092] The sound cloning model is obtained by the sound cloning model training method described in this embodiment.

[0093] Reference Figure 7 , shows a reasoning schematic diagram of the sound cloning model provided in an embodiment of the present application. After the text sequence to be input and the prompt mel-spectrogram are input into the sound cloning model, the mel-spectrogram corresponding to the text sequence to be synthesized can be regarded as the masked part after the mask operation, so that the sound cloning model generates and outputs the mel-spectrogram corresponding to the masked text sequence to be synthesized.

[0094] The text encoder of the voice cloning model generates a text vector of the text sequence to be input, and the voice cloning model extracts the speaker vector based on the prompt Mel spectrum. Finally, the decoder of the voice cloning model generates and outputs the Mel spectrum corresponding to the text sequence to be synthesized based on the text vector, speaker vector and prompt Mel spectrum of the text sequence to be input.

[0095] S204: Convert the mel-spectrogram result to obtain audio data corresponding to the text sequence to be input.

[0096] For example, the mel spectrum result can be converted into audio data corresponding to the text sequence to be input according to a vocoder tool, such as a vocoder tool such as HiFiGAN.

[0097] The sound cloning model trained in this embodiment is a high-quality streaming zero-sample sound cloning model based on stream matching technology. The overall model architecture is designed by a diffusion model based on stream matching, and the text sequence is optimized to reduce the need for third-party time alignment tools. It also supports setting the pronunciation of specific text sequences, introduces mask design, and constructs a streaming inference architecture based on block calculation, so that the model can achieve high-quality sound cloning.

[0098] Specifically, the sound cloning model of this embodiment has at least the following beneficial effects: 1. Enable zero-sample voice cloning capabilities; use masks to design the model's loss function, train the model based on large-scale corpus training data, and improve the model's contextual learning capabilities, so that the model can extract the timbre and rhythm information of speakers outside the training set based on the given prompt voice, and then synthesize the target speaker's audio to achieve voice cloning.

[0099] 2. Improve the quality and timbre similarity of synthesized speech. Design the model architecture based on the diffusion model of flow matching technology to capture the details of complex data distribution and achieve high-quality speech synthesis. Introduce speaker vectors to further improve the timbre similarity of voice cloning.

[0100] 3. Support streaming voice interaction; the voice cloning model adopts a block sample-based calculation method. It can support non-streaming reasoning or streaming reasoning by adjusting the chunk size, realizing flexible and diverse generation methods. Streaming reasoning can also effectively reduce the first packet delay and response time.

[0101] In a feasible implementation, the hifitts dataset, aishell3 dataset and a manually annotated dataset are used, which include two languages ​​​​in Chinese and English, and the length of audio data is about 10,000 hours; among them, 1 hour of audio data is randomly selected as a verification set, and the remaining audio data are used to construct a training sample set. In the actual implementation process, all audio data are resampled to 24k, the frame length of the Mel spectrum is 1024, the frame shift is 256, and the dimension is 80 dimensions.

[0102] In the model structure of the constructed sound cloning model, the number of Conformer layers of the text decoder is set to 6, the downsampling module and upsampling module of the decoder are both 2 layers, each layer contains 1 layer of ResNet block and 4 layers of Transformers layer, and the middle module consists of 12 layers of Transformer layers.

[0103] In streaming computing, the chunk size is set to 52 frames, and the history cache length of the attention mechanism is set to 10 frames.

[0104] It has been verified that the voice cloning model can reproduce the speaker's relevant timbre, style and rhythm with high quality for speakers outside the training sample set, provided that a prompt voice of more than 3 seconds is provided.

[0105] Reference Figure 8 , shows a functional module diagram of a training device for a sound cloning model provided in an embodiment of the present application, the training device comprising: A training sample set building module 101 is used to pre-process the acquired multiple audio data to build a training sample set, wherein any training sample in the training sample set includes a text sequence sample, a Mel spectrum sample and a target Mel spectrum label corresponding to any audio data; The training module 102 is used to train a preset voice cloning model according to the training sample set and a preset block size, wherein the text encoder in the preset voice cloning model is used to extract a text vector according to a text sequence sample of any training sample, and the decoder in the preset voice cloning model is used to output a predicted Mel spectrum according to the text vector, Mel spectrum sample and speaker vector of any training sample, and the block size is used to adjust the calculation mode of the preset voice cloning model to non-streaming or streaming; The iteration module 103 is used to iteratively train the preset sound cloning model according to the target mel spectrum label, the predicted mel spectrum and the loss function of the optimal transmission stream of any training sample to obtain a trained sound cloning model.

[0106] Optionally, the training sample set building module includes a training sample set building unit, which is used to: For any audio data obtained, generate a text sequence and Mel spectrum corresponding to the audio data; Perform a masking operation on the Mel spectrum corresponding to any audio data to obtain a masked Mel spectrum and an unmasked Mel spectrum, use the unmasked Mel spectrum as a Mel spectrum sample, and use the masked Mel spectrum as a target Mel spectrum label; Insert characters into the text sequence corresponding to the target Mel spectrum label in the audio data so that the length of the text sequence corresponding to the target Mel spectrum label is consistent with the number of frames of the target Mel spectrum label, and use the text sequence corresponding to the Mel spectrum sample as the text sequence sample; The text sequence sample, Mel spectrum sample and target Mel spectrum label corresponding to any audio data are taken as a training sample.

[0107] Optionally, the training sample set construction unit includes a mask unit, which is used to: For any audio data, a mask operation is performed on a preset mask range of the Mel spectrum corresponding to the audio data, or a mask operation is performed on all Mel spectrums corresponding to the audio data.

[0108] Optionally, the training device further comprises: Model building module, used to build preset sound cloning models; Among them, the preset sound cloning model includes a text encoder and a decoder, the text encoder includes a Conformer layer, which is used to extract a text vector of any text sequence sample; the decoder is based on a Transformer structure in the form of U-net, including a downsampling module, an intermediate module and an upsampling module, any module among the downsampling module, the intermediate module and the upsampling module includes a residual module based on a convolutional layer and a Transformer module based on an attention mechanism, and any convolutional layer in the preset sound cloning model is a block convolution, and the attention mechanism of the Transformer module is an attention mechanism based on a block mask.

[0109] Optionally, the training device further comprises: A block size setting module, for determining, in response to a preset operation, a range of variation of a block size of a preset sound cloning model during training, wherein the block size is used to adjust a calculation mode of the preset sound cloning model to a non-streaming mode or a streaming mode; a training strategy setting module, configured to execute a first training strategy or a second training strategy in response to a training strategy selection operation in a process of training a preset sound cloning model based on the training sample set; Among them, when executing the first training strategy, in each iterative training process of the preset sound cloning model, a block size is randomly selected from the variation range of the block size as the block size of this iteration process; when executing the second training strategy, the loss function in the non-streaming calculation mode and the loss function in the streaming calculation mode are determined respectively, and in the back propagation of the preset sound cloning model, the model parameters of the preset sound cloning model are updated based on the loss function in the non-streaming calculation mode and the loss function in the streaming calculation mode at the same time.

[0110] Reference Fig. 9 , shows a functional module diagram of a sound cloning device provided in an embodiment of the present application, the sound cloning device comprising: The acquisition module 201 is used to acquire the target prompt speech and the text sequence to be synthesized, and generate the prompt mel spectrum and prompt text sequence of the target prompt speech; A preprocessing module 202, configured to preprocess the text sequence to be input according to the text sequence to be synthesized and the prompt text sequence; a mel-spectrogram generating module 203, configured to input the prompt mel-spectrogram of the target prompt speech and the text sequence to be input into a sound cloning model, and output a mel-spectrogram result corresponding to the text sequence to be synthesized through the sound cloning model, wherein the sound cloning model is obtained by the training method of the sound cloning model according to any one of claims 1 to 5; The audio data conversion module 204 is used to convert the audio data corresponding to the text sequence to be input according to the mel spectrum result.

[0111] Optionally, the preprocessing module includes a preprocessing unit, configured to: According to the text sequence to be synthesized, estimating the number of frames of the mel spectrum result corresponding to the text sequence to be synthesized; Characters are inserted into a text sequence composed of the text sequence to be synthesized and the prompt text sequence to obtain a text sequence to be input, wherein the length of the text sequence to be input is equal to the sum of the number of frames of the prompt mel-spectrogram and the number of frames of the mel-spectrogram result corresponding to the text sequence to be synthesized.

[0112] An embodiment of the present application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the training method of the sound cloning model as described in the embodiment, or implements the sound cloning method as described in the embodiment.

[0113] An embodiment of the present application further provides a computer-readable storage medium having a computer program / instruction stored thereon. When the computer program / instruction is executed by a processor, the training method of the sound cloning model as described in the embodiment or the sound cloning method as described in the embodiment is implemented.

[0114] The embodiments of the present application also provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the training method of the sound cloning model as described in the embodiments, or implements the sound cloning method as described in the embodiments.

[0115] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0116] It should be understood by those skilled in the art that the embodiments of the present application can be provided as methods, devices or computer program products. Therefore, the embodiments of the present application can be in the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the embodiments of the present application can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0117] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0118] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable terminal device. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0120] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.

[0121] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0122] Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A method for training a sound cloning model, characterized in that: The method comprises: Preprocessing the acquired multiple audio data to construct a training sample set, wherein any training sample in the training sample set includes a text sequence sample, a Mel spectrum sample, and a target Mel spectrum label corresponding to any audio data; According to the training sample set and the preset block size, a preset voice cloning model is trained, the text encoder in the preset voice cloning model is used to extract a text vector according to a text sequence sample of any training sample, the decoder in the preset voice cloning model is used to output a predicted mel spectrum according to the text vector, mel spectrum sample and speaker vector of any training sample, and the block size is used to adjust the calculation mode of the preset voice cloning model to non-streaming or streaming; According to the target mel spectrum label, predicted mel spectrum and the loss function of the optimal transmission stream of any training sample, the preset sound cloning model is iteratively trained to obtain a trained sound cloning model.

2. The method according to claim 1, characterized in that Preprocess the acquired multiple audio data to construct a training sample set, including: For any audio data obtained, generate a text sequence and Mel spectrum corresponding to the audio data; Perform a masking operation on the Mel spectrum corresponding to any audio data to obtain a masked Mel spectrum and an unmasked Mel spectrum, use the unmasked Mel spectrum as a Mel spectrum sample, and use the masked Mel spectrum as a target Mel spectrum label; Insert characters into the text sequence corresponding to the target Mel spectrum label in the audio data so that the length of the text sequence corresponding to the target Mel spectrum label is consistent with the number of frames of the target Mel spectrum label, and use the text sequence corresponding to the Mel spectrum sample as the text sequence sample; The text sequence sample, Mel spectrum sample and target Mel spectrum label corresponding to any audio data are taken as a training sample.

3. The method according to claim 2, characterized in that The masking operation of the Mel spectrum corresponding to any audio data includes: For any audio data, a mask operation is performed on a preset mask range of the Mel spectrum corresponding to the audio data, or a mask operation is performed on all Mel spectrums corresponding to the audio data.

4. The method according to claim 1, characterized in that The method further comprises: Build preset sound cloning models; Among them, the preset sound cloning model includes a text encoder and a decoder, the text encoder includes a Conformer layer for extracting a text vector of any text sequence sample; the decoder is based on a Transformer structure in the form of U-net, including a downsampling module, an intermediate module and an upsampling module, any module among the downsampling module, the intermediate module and the upsampling module includes a residual module based on a convolutional layer and a Transformer module based on an attention mechanism, and any convolutional layer in the preset sound cloning model is a block convolution, and the attention mechanism of the Transformer module is an attention mechanism based on a block mask.

5. The method according to claim 4, characterized in that The method further comprises: In response to a preset operation, determining a range of variation of a block size of a preset sound cloning model during training, wherein the block size is used to adjust a calculation mode of the preset sound cloning model to a non-streaming mode or a streaming mode; In response to the training strategy selection operation, in the process of training a preset sound cloning model based on the training sample set, executing a first training strategy or a second training strategy; Among them, when executing the first training strategy, in each iterative training process of the preset sound cloning model, a block size is randomly selected from the variation range of the block size as the block size of this iteration process; when executing the second training strategy, the loss function in the non-streaming calculation mode and the loss function in the streaming calculation mode are determined respectively, and in the back propagation of the preset sound cloning model, the model parameters of the preset sound cloning model are updated based on the loss function in the non-streaming calculation mode and the loss function in the streaming calculation mode at the same time.

6. A sound cloning method, characterized in that: The method comprises: Acquire a target prompt speech and a text sequence to be synthesized, and generate a prompt mel spectrum and a prompt text sequence of the target prompt speech; Preprocessing to obtain a text sequence to be input according to the text sequence to be synthesized and the prompt text sequence; Inputting the prompt mel-spectrogram of the target prompt speech and the text sequence to be input into a sound cloning model, and outputting the mel-spectrogram result corresponding to the text sequence to be synthesized through the sound cloning model, wherein the sound cloning model is obtained by the training method of the sound cloning model according to any one of claims 1 to 5; According to the mel-spectrogram result, audio data corresponding to the text sequence to be input is converted.

7. The method according to claim 6, characterized in that Preprocessing the text sequence to be input according to the text sequence to be synthesized and the prompt text sequence includes: According to the text sequence to be synthesized, estimating the number of frames of the mel spectrum result corresponding to the text sequence to be synthesized; Characters are inserted into a text sequence composed of the text sequence to be synthesized and the prompt text sequence to obtain a text sequence to be input, wherein the length of the text sequence to be input is equal to the sum of the number of frames of the prompt mel-spectrogram and the number of frames of the mel-spectrogram result corresponding to the text sequence to be synthesized.

8. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the training method of the sound cloning model according to any one of claims 1 to 5, or implements the sound cloning method according to any one of claims 6 to 7.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instruction is executed by a processor, the training method of the sound cloning model as described in any one of claims 1 to 5 is implemented, or the sound cloning method as described in any one of claims 6 to 7 is implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the training method of the sound cloning model as described in any one of claims 1 to 5 is implemented, or the sound cloning method as described in any one of claims 6 to 7 is implemented.

Citation Information

Patent Citations

  • System and method for training cloned tone and rhythm based on Bottleneck features

    CN111210803A

  • Voice cloning method and system based on cross-domain consistency loss

    CN116229932A

  • Acoustic model training method and device, speech synthesis method and device and computer equipment

    CN116312458A

  • Audio generation network training method, audio generation method and device

    CN116959464A

  • Speech synthesis model training method, speech synthesis method and speech synthesis device

    CN118430509A

Cited By

  • Control method and device based on task understanding representation, equipment and medium

    CN120871706A