Voice decoupling method, device, electronic device, storage medium and program product

Through reconstruction loss training of the timbre encoder and prosody encoder, the problem of insufficient speech decoupling in the existing technology is solved, and high-quality speech synthesis and widely used speech decoupling effects are achieved.

CN119785775BActive Publication Date: 2025-09-19IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411940443.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-09-19
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

The existing speech decoupling model requires label annotation due to the supervised training method, which leads to limited training samples and cannot effectively decouple rhythm and timbre. In addition, the unsupervised training effect is poor and high-quality speech synthesis cannot be achieved.

Method used

A timbre encoder and a prosody encoder are used to extract timbre and prosody information respectively through the first reconstruction loss and second reconstruction loss training. The model is trained using sample audio data from different speakers to improve the adequacy of decoupling and generalization.

Benefits of technology

It achieves full decoupling of speech content, timbre and rhythm, improves the quality and generalization ability of speech synthesis, and can be trained based on audio data from massive speakers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785775B_ABST
    Figure CN119785775B_ABST
Patent Text Reader

Abstract

The present invention provides a speech decoupling method, device, electronic device, storage medium, and program product, relating to the field of audio processing technology. The method comprises: inputting the speech data to be decoupled into a timbre encoder and a prosody encoder respectively, obtaining decoupled timbre information output by the timbre encoder, and decoupled prosody information output by the prosody encoder; wherein a first reconstruction loss is determined based on sample audio data of a first speaker and reconstructed audio data of the first speaker, and the reconstructed audio data of the first speaker is reconstructed based on target timbre information corresponding to the first speaker and target prosody information corresponding to the first speaker. The present invention can constrain the timbre preservation capability of the timbre encoder through the first reconstruction loss, thereby improving the adequacy of timbre decoupling, and can constrain the prosody preservation capability of the prosody encoder, thereby improving the adequacy of prosody decoupling; and the present invention can also improve the generalization of speech decoupling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio processing technology, and in particular to a speech decoupling method, device, electronic device, storage medium and program product. Background Art

[0002] With the rapid development of artificial intelligence, speech synthesis is finding an increasingly broad range of applications, including personalized speech synthesis, timbre transfer, and prosody transfer. Speech synthesis often requires speech decoupling, which involves breaking down human speech into its three components: content, prosody, and timbre. To achieve high-quality speech synthesis, effective decoupling of speech content, timbre, and prosody is crucial while maintaining naturalness.

[0003] Currently, speech decoupling models are trained using supervised training, and speech decoupling is performed based on the trained speech decoupling model. However, supervised training requires labeling, which prevents training from using too many training samples, resulting in an inability to effectively decouple prosody and timbre. Furthermore, it cannot be trained on massive amounts of speaker audio data, resulting in low generalization of speech decoupling. Even when training speech decoupling models using unsupervised training, the decoupling effect achieved by this training method often falls short of expectations, resulting in an inability to effectively decouple prosody and timbre. Summary of the Invention

[0004] The present invention provides a speech decoupling method, device, electronic device, storage medium and program product, which are used to solve the defect of incomplete timbre and rhythm decoupling in the prior art and achieve full speech decoupling.

[0005] The present invention provides a speech decoupling method, comprising:

[0006] Inputting the to-be-decoupled speech data into a timbre encoder and a prosody encoder respectively, obtaining decoupled timbre information output by the timbre encoder and decoupled prosody information output by the prosody encoder;

[0007] The timbre encoder and the prosody encoder are trained based on a first reconstruction loss and a second reconstruction loss;

[0008] The first reconstruction loss is determined based on sample audio data of a first speaker and reconstructed audio data of the first speaker, the reconstructed audio data of the first speaker is reconstructed based on target timbre information corresponding to the first speaker and target prosody information corresponding to the first speaker, the target timbre information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker by the timbre encoder, and the target prosody information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker by the prosody encoder; the second reconstruction loss is determined based on sample audio data of a second speaker and reconstructed audio data of the second speaker, the reconstructed audio data of the second speaker is reconstructed based on target timbre information corresponding to the second speaker and target prosody information corresponding to the second speaker, the target timbre information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker by the timbre encoder, and the target prosody information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker by the prosody encoder;

[0009] The first speaker and the second speaker are different speakers.

[0010] According to a speech decoupling method provided by the present invention, the timbre encoder and the prosody encoder are trained based on the following method:

[0011] Inputting the sample audio data of the first speaker and the sample audio data of the second speaker into a timbre encoder in the acoustic model respectively to obtain target timbre information output by the timbre encoder respectively, and inputting the sample audio data of the first speaker and the sample audio data of the second speaker into a prosody encoder in the acoustic model respectively to obtain target prosody information output by the prosody encoder respectively;

[0012] Inputting the target timbre information corresponding to the first speaker and the target prosody information corresponding to the second speaker into an audio decoder in the acoustic model to obtain first reconstructed audio data output by the audio decoder; the audio decoder is configured to reconstruct audio data based on the input information;

[0013] Inputting the first reconstructed audio data into the timbre encoder to obtain reconstructed timbre information output by the timbre encoder, and inputting the first reconstructed audio data into the prosody encoder to obtain reconstructed prosody information output by the prosody encoder;

[0014] Inputting the reconstructed timbre information and the target prosody information corresponding to the first speaker into the audio decoder to obtain second reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the second speaker into the audio decoder to obtain third reconstructed audio data output by the audio decoder;

[0015] determining a first reconstruction loss based on the sample audio data of the first speaker and the second reconstructed audio data, and determining a second reconstruction loss based on the sample audio data of the second speaker and the third reconstructed audio data;

[0016] The acoustic model is trained based on the first reconstruction loss and the second reconstruction loss.

[0017] According to a speech decoupling method provided by the present invention, the sample audio data of the first speaker includes different first audio data and second audio data, and the sample audio data of the second speaker includes different third audio data and fourth audio data;

[0018] The method comprises: inputting the sample audio data of the first speaker and the sample audio data of the second speaker into a timbre encoder in the acoustic model to obtain target timbre information output by the timbre encoder, and inputting the sample audio data of the first speaker and the sample audio data of the second speaker into a prosody encoder in the acoustic model to obtain target prosody information output by the prosody encoder, including:

[0019] inputting the second audio data and the fourth audio data into the timbre encoder respectively to obtain target timbre information outputted by the timbre encoder respectively, and inputting the first audio data and the third audio data into the prosody encoder respectively to obtain target prosody information outputted by the prosody encoder respectively;

[0020] The step of inputting the target timbre information corresponding to the first speaker and the target prosody information corresponding to the second speaker into an audio decoder in the acoustic model to obtain first reconstructed audio data output by the audio decoder includes:

[0021] Inputting the target timbre information corresponding to the second audio data of the first speaker and the target prosody information corresponding to the third audio data of the second speaker into the audio decoder to obtain first reconstructed audio data output by the audio decoder;

[0022] The step of inputting the reconstructed timbre information and the target prosody information corresponding to the first speaker into the audio decoder to obtain second reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the second speaker into the audio decoder to obtain third reconstructed audio data output by the audio decoder includes:

[0023] inputting the reconstructed timbre information and the target prosody information corresponding to the first audio data of the first speaker into the audio decoder to obtain second reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the fourth audio data of the second speaker into the audio decoder to obtain third reconstructed audio data output by the audio decoder;

[0024] The determining of a first reconstruction loss based on the sample audio data of the first speaker and the second reconstructed audio data, and determining a second reconstruction loss based on the sample audio data of the second speaker and the third reconstructed audio data, comprises:

[0025] A first reconstruction loss is determined based on the first audio data of the first speaker and the second reconstructed audio data, and a second reconstruction loss is determined based on the third audio data of the second speaker and the third reconstructed audio data.

[0026] According to a speech decoupling method provided by the present invention, the second audio data and the fourth audio data are respectively inputted into the timbre encoder to obtain target timbre information respectively outputted by the timbre encoder, and the first audio data and the third audio data are respectively inputted into the prosody encoder to obtain target prosody information respectively outputted by the prosody encoder, including:

[0027] inputting the second audio data and the fourth audio data into the timbre encoder, respectively, to obtain target timbre information outputted by the timbre encoder, inputting the first audio data and the third audio data into the prosody encoder, respectively, to obtain target prosody information outputted by the prosody encoder, and inputting first text data corresponding to the first audio data and third text data corresponding to the third audio data into a content encoder in the acoustic model, respectively, to obtain target content information outputted by the content encoder;

[0028] The step of inputting the target timbre information corresponding to the second audio data of the first speaker and the target prosody information corresponding to the third audio data of the second speaker into the audio decoder to obtain first reconstructed audio data output by the audio decoder includes:

[0029] Inputting the target timbre information corresponding to the second audio data of the first speaker, the target prosody information corresponding to the third audio data of the second speaker, and the target content information corresponding to the third audio data of the second speaker into the audio decoder, thereby obtaining first reconstructed audio data output by the audio decoder;

[0030] The step of inputting the reconstructed timbre information and the target prosody information corresponding to the first audio data of the first speaker into the audio decoder to obtain second reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the fourth audio data of the second speaker into the audio decoder to obtain third reconstructed audio data output by the audio decoder includes:

[0031] The reconstructed timbre information and the target prosody information corresponding to the first audio data of the first speaker, as well as the target content information corresponding to the first audio data of the first speaker are input into the audio decoder to obtain second reconstructed audio data output by the audio decoder, and the target timbre information and the reconstructed prosody information corresponding to the fourth audio data of the second speaker, as well as the target content information corresponding to the third audio data of the second speaker are input into the audio decoder to obtain third reconstructed audio data output by the audio decoder.

[0032] According to a speech decoupling method provided by the present invention, the acoustic model is trained based on the first reconstruction loss and the second reconstruction loss, comprising:

[0033] Inputting the reconstructed timbre information and the target prosody information corresponding to the second speaker into the audio decoder to obtain fourth reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the first speaker into the audio decoder to obtain fifth reconstructed audio data output by the audio decoder;

[0034] determining a third reconstruction loss based on the first reconstructed audio data and the fourth reconstructed audio data, and determining a fourth reconstruction loss based on the first reconstructed audio data and the fifth reconstructed audio data;

[0035] The acoustic model is trained based on the first reconstruction loss, the second reconstruction loss, the third reconstruction loss, and the fourth reconstruction loss.

[0036] According to a speech decoupling method provided by the present invention, the sample audio data of the first speaker includes different first audio data and second audio data, and the sample audio data of the second speaker includes different third audio data and fourth audio data;

[0037] The method comprises: inputting the sample audio data of the first speaker and the sample audio data of the second speaker into a timbre encoder in the acoustic model to obtain target timbre information output by the timbre encoder, and inputting the sample audio data of the first speaker and the sample audio data of the second speaker into a prosody encoder in the acoustic model to obtain target prosody information output by the prosody encoder, including:

[0038] inputting the second audio data and the fourth audio data into the timbre encoder, respectively, to obtain target timbre information outputted by the timbre encoder, inputting the first audio data and the third audio data into the prosody encoder, respectively, to obtain target prosody information outputted by the prosody encoder, and inputting first text data corresponding to the first audio data and third text data corresponding to the third audio data into a content encoder in the acoustic model, respectively, to obtain target content information outputted by the content encoder;

[0039] The step of inputting the target timbre information corresponding to the first speaker and the target prosody information corresponding to the second speaker into an audio decoder in the acoustic model to obtain first reconstructed audio data output by the audio decoder includes:

[0040] Inputting the target timbre information corresponding to the second audio data of the first speaker, the target prosody information corresponding to the third audio data of the second speaker, and the target content information corresponding to the third audio data of the second speaker into the audio decoder, thereby obtaining first reconstructed audio data output by the audio decoder;

[0041] The step of inputting the reconstructed timbre information and the target prosody information corresponding to the second speaker into the audio decoder to obtain fourth reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the first speaker into the audio decoder to obtain fifth reconstructed audio data output by the audio decoder includes:

[0042] The reconstructed timbre information and the target prosody information corresponding to the third audio data of the second speaker, as well as the target content information corresponding to the third audio data of the second speaker are input into the audio decoder to obtain fourth reconstructed audio data output by the audio decoder, and the target timbre information and the reconstructed prosody information corresponding to the second audio data of the first speaker, as well as the target content information corresponding to the third audio data of the second speaker are input into the audio decoder to obtain fifth reconstructed audio data output by the audio decoder.

[0043] According to a speech decoupling method provided by the present invention, the acoustic model is trained based on the first reconstruction loss, the second reconstruction loss, the third reconstruction loss, and the fourth reconstruction loss, including:

[0044] Inputting the first reconstructed audio data, the second reconstructed audio data, the third reconstructed audio data, and the fourth reconstructed audio data into a discriminator respectively, and obtaining four discrimination losses output by the discriminator;

[0045] The acoustic model is trained based on the first reconstruction loss, the second reconstruction loss, the third reconstruction loss, the fourth reconstruction loss, and the four discriminant losses.

[0046] According to a speech decoupling method provided by the present invention, the method of inputting sample audio data of a first speaker and sample audio data of a second speaker into a prosody encoder in an acoustic model respectively to obtain target prosody information outputted by the prosody encoder respectively includes:

[0047] Performing timbre change processing on the sample audio data of the first speaker and the sample audio data of the second speaker, respectively, to obtain the sample audio data of the first speaker after timbre change and the sample audio data of the second speaker after timbre change;

[0048] The sample audio data of the first speaker after the timbre change and the sample audio data of the second speaker after the timbre change are respectively inputted into a prosody encoder in the acoustic model to obtain target prosody information respectively outputted by the prosody encoder.

[0049] The present invention also provides a speech decoupling device, comprising:

[0050] A speech decoupling module is used to input the speech data to be decoupled into the timbre encoder and the prosody encoder respectively, to obtain the decoupled timbre information output by the timbre encoder and the decoupled prosody information output by the prosody encoder;

[0051] The timbre encoder and the prosody encoder are trained based on a first reconstruction loss and a second reconstruction loss;

[0052] The first reconstruction loss is determined based on sample audio data of a first speaker and reconstructed audio data of the first speaker, the reconstructed audio data of the first speaker is reconstructed based on target timbre information corresponding to the first speaker and target prosody information corresponding to the first speaker, the target timbre information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker by the timbre encoder, and the target prosody information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker by the prosody encoder; the second reconstruction loss is determined based on sample audio data of a second speaker and reconstructed audio data of the second speaker, the reconstructed audio data of the second speaker is reconstructed based on target timbre information corresponding to the second speaker and target prosody information corresponding to the second speaker, the target timbre information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker by the timbre encoder, and the target prosody information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker by the prosody encoder;

[0053] The first speaker and the second speaker are different speakers.

[0054] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the speech decoupling method described above is implemented.

[0055] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described speech decoupling methods.

[0056] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned speech decoupling methods.

[0057] The speech decoupling method, device, electronic device, storage medium and program product provided by the present invention are obtained by training a timbre encoder and a prosody encoder based on a first reconstruction loss and a second reconstruction loss, and the first reconstruction loss is determined based on the sample audio data of the first speaker and the reconstructed audio data of the first speaker, the reconstructed audio data of the first speaker is reconstructed based on the target timbre information corresponding to the first speaker and the target prosody information corresponding to the first speaker, and the target timbre information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker by the timbre encoder, and the target prosody information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker by the prosody encoder, so that the first reconstruction loss can constrain the timbre maintenance ability of the timbre encoder for the first speaker, thereby improving the adequacy of timbre decoupling, and can constrain the prosody maintenance ability of the prosody encoder for the first speaker, thereby improving the adequacy of prosody decoupling; at the same time, the second reconstruction loss The loss is determined based on the sample audio data of the second speaker and the reconstructed audio data of the second speaker, the reconstructed audio data of the second speaker is reconstructed based on the target timbre information corresponding to the second speaker and the target prosody information corresponding to the second speaker, and the target timbre information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker based on the timbre encoder, and the target prosody information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker based on the prosody encoder, so that the second reconstruction loss can constrain the timbre retention ability of the timbre encoder for the second speaker, thereby improving the adequacy of timbre decoupling, and can constrain the prosody retention ability of the prosody encoder for the second speaker, thereby improving the adequacy of prosody decoupling; in addition, the first speaker and the second speaker are different speakers, so that the model training can be based on the sample audio data of different speakers, that is, it can be trained based on the audio data of a large number of speakers to improve the generalization of speech decoupling. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0059] Figure 1 This is one of the flow charts of the speech decoupling method provided by the present invention.

[0060] Figure 2 This is the second flow chart of the speech decoupling method provided by the present invention.

[0061] Figure 3This is the third flow chart of the speech decoupling method provided by the present invention.

[0062] Figure 4 This is the fourth flow chart of the speech decoupling method provided by the present invention.

[0063] Figure 5 It is a structural schematic diagram of the speech decoupling device provided by the present invention.

[0064] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0065] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0066] With the rapid development of artificial intelligence, speech synthesis technology has made significant progress, and its application scenarios are becoming increasingly diverse, including personalized speech synthesis, timbre transfer, prosody transfer, and sound editing. Speech synthesis often requires speech decomposition, which involves breaking down human speech into its three components: content, prosody, and timbre. To achieve high-quality speech synthesis, it is necessary to effectively decouple speech content, timbre, and prosody while maintaining naturalness.

[0067] However, these three attributes are closely intertwined in speech signals and difficult to separate directly. Traditional speech decoupling methods often rely on manual feature engineering, which is not only time-consuming and labor-intensive, but also difficult to adapt to different languages ​​and accents.

[0068] It's important to note that speech is a crucial means of conveying information and expressing emotion in human communication. Speech signals contain a wealth of information, including not only the content of speech (i.e., textual information), but also the speaker's timbre (i.e., the characteristics of the voice) and prosody (i.e., the rhythm and emphasis of speech). These attributes collectively contribute to the uniqueness and diversity of human speech.

[0069] Currently, speech decoupling models are trained using supervised training, and speech decoupling is performed based on the trained speech decoupling model. However, supervised training requires labeling, which prevents training from using too many training samples, resulting in an inability to effectively decouple prosody and timbre. Furthermore, it cannot be trained on massive amounts of speaker audio data, resulting in low generalization of speech decoupling. Even when training speech decoupling models using unsupervised training, the decoupling effect achieved by this training method often falls short of expectations, resulting in an inability to effectively decouple prosody and timbre.

[0070] For example, supervised approaches use phoneme prediction to decouple content information, speaker classification to decouple timbre information, and fundamental frequency prediction to decouple prosody information. Furthermore, gradient reversal schemes are used to reduce the leakage of other attributes within each branch. However, supervised speech decoupling schemes often require careful design of the proportions of various losses, and model training stability is poor. Speaker prediction tasks cannot accept massive amounts of speaker data, resulting in unsatisfactory timbre generalization.

[0071] In view of the above problems, the present invention proposes the following embodiments. Figures 1-4 The speech decoupling method of the present invention is described.

[0072] Figure 1 This is one of the flow charts of the speech decoupling method provided by the present invention, such as Figure 1 As shown, the speech decoupling method includes the following step 110.

[0073] Step 110 : Input the speech data to be decoupled into a timbre encoder and a prosody encoder respectively, to obtain decoupled timbre information output by the timbre encoder and decoupled prosody information output by the prosody encoder.

[0074] Here, the speech data to be decoupled is audio data to be decoupled, and the speech data to be decoupled may be speech data collected from a speaker. Further, the speech data to be decoupled may be mel spectrum data.

[0075] Here, the timbre encoder is used to decouple the timbre information from the speech data to be decoupled to obtain decoupled timbre information. The prosody encoder is used to decouple the prosody information from the speech data to be decoupled to obtain decoupled prosody information.

[0076] The timbre information decoupled by the timbre encoder includes the speaker's global voiceprint information. For example, the timbre encoder can derive a global timbre vector from the mel-spectrogram data (audio data). This means that timbre is a global attribute that does not vary with the length of the audio. Therefore, the timbre information extracted from two voices of the same speaker should represent the timbre of the same speaker and should be sufficiently similar.

[0077] The prosody information decoupled by the prosody encoder may include, but is not limited to, information such as the speaker's speaking style, rhythm, intonation, and emotions.

[0078] The timbre encoder and the prosody encoder are trained based on the first reconstruction loss and the second reconstruction loss, that is, the timbre encoder and the prosody encoder are trained based on the first reconstruction loss and the second reconstruction loss.

[0079] Wherein, the first reconstruction loss is determined based on the sample audio data of the first speaker and the reconstructed audio data of the first speaker, the reconstructed audio data of the first speaker is reconstructed based on the target timbre information corresponding to the first speaker and the target prosody information corresponding to the first speaker, the target timbre information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker based on the timbre encoder, and the target prosody information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker based on the prosody encoder.

[0080] Here, the sample audio data of the first speaker may be voice data collected from the first speaker.

[0081] The reconstructed audio data of the first speaker is reconstructed based on the target timbre information corresponding to the first speaker and the target prosody information corresponding to the first speaker, and the target timbre information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker based on the timbre encoder, and the target prosody information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker based on the prosody encoder. Therefore, if the speech decoupling of the timbre encoder and the prosody encoder is more accurate and sufficient, the first reconstruction loss should be as small as possible. Based on this, the first reconstruction loss can constrain the timbre encoder's ability to maintain the timbre of the first speaker, thereby improving the adequacy of the timbre decoupling, and can constrain the prosody encoder's ability to maintain the prosody of the first speaker, thereby improving the adequacy of the prosody decoupling.

[0082] The second reconstruction loss is determined based on the sample audio data of the second speaker and the reconstructed audio data of the second speaker, the reconstructed audio data of the second speaker is reconstructed based on the target timbre information corresponding to the second speaker and the target prosody information corresponding to the second speaker, the target timbre information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker based on the timbre encoder, and the target prosody information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker based on the prosody encoder.

[0083] Here, the sample audio data of the second speaker may be voice data collected from the second speaker.

[0084] The reconstructed audio data of the second speaker is reconstructed based on the target timbre information corresponding to the second speaker and the target prosody information corresponding to the second speaker, and the target timbre information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker through the timbre encoder, and the target prosody information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker through the prosody encoder. Therefore, if the speech decoupling of the timbre encoder and the prosody encoder is more accurate and sufficient, the second reconstruction loss should be as small as possible. Based on this, the second reconstruction loss can constrain the timbre encoder's ability to maintain the timbre of the second speaker, thereby improving the adequacy of the timbre decoupling, and can constrain the prosody encoder's ability to maintain the prosody of the second speaker, thereby improving the adequacy of the prosody decoupling.

[0085] Among them, the first speaker and the second speaker are different speakers, so the model training can be performed based on sample audio data of different speakers, that is, training can be performed based on audio data of a large number of speakers, thereby improving the generalization of speech decoupling, that is, the trained timbre encoder and prosody encoder can be transplanted to various scenarios for the purpose of decoupling.

[0086] Furthermore, the sample audio data of the first speaker includes different first audio data and second audio data, and the sample audio data of the second speaker includes different third audio data and fourth audio data. The first reconstruction loss is determined based on the first audio data of the first speaker and the reconstructed audio data of the first speaker, the reconstructed audio data of the first speaker is reconstructed based on the target timbre information corresponding to the first speaker and the target prosody information corresponding to the first speaker, the target timbre information corresponding to the first speaker is obtained by extracting the second audio data of the first speaker based on the timbre encoder, and the target prosody information corresponding to the first speaker is obtained by extracting the first audio data of the first speaker based on the prosody encoder. The second reconstruction loss is determined based on the third audio data of the second speaker and the reconstructed audio data of the second speaker, the reconstructed audio data of the second speaker is reconstructed based on the target timbre information corresponding to the second speaker and the target prosody information corresponding to the second speaker, the target timbre information corresponding to the second speaker is obtained by extracting the fourth audio data of the second speaker based on the timbre encoder, and the target prosody information corresponding to the second speaker is obtained by extracting the third audio data of the second speaker based on the prosody encoder.

[0087] Based on this, since the first audio data and the second audio data are audio data of the same speaker, the timbre information of the first audio data and the second audio data is the same, but the prosody information of the first audio data and the second audio data is different, and since the third audio data and the fourth audio data are audio data of the same speaker, the timbre information of the third audio data and the fourth audio data is the same, but the prosody information of the third audio data and the fourth audio data is different. Therefore, the first reconstruction loss and the second reconstruction loss can constrain the stability of the timbre encoder in extracting the timbre of different voices of the same person, thereby improving the stability of timbre decoupling, and can constrain the stability of the prosody encoder in extracting the prosody of the same prosodic audio of different people, thereby improving the stability of prosody decoupling.

[0088] Furthermore, the speech data to be decoupled is input into the timbre encoder, the prosody encoder and the content encoder respectively to obtain the decoupled timbre information output by the timbre encoder, the decoupled prosody information output by the prosody encoder, and the decoupled content information output by the content encoder.

[0089] Here, the content information decoupled by the content encoder includes the speech content of the speaker.

[0090] In one embodiment, the content encoder includes multiple Conformer layers for capturing the contextual relationship of text data and encoding it into content information.

[0091] The timbre encoder, prosody encoder, and content encoder are trained based on a first reconstruction loss and a second reconstruction loss. The first reconstruction loss is determined based on the sample audio data of the first speaker and the reconstructed audio data of the first speaker. The reconstructed audio data of the first speaker is reconstructed based on the target timbre information corresponding to the first speaker, the target prosody information corresponding to the first speaker, and the target content information corresponding to the first speaker. The target timbre information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker using the timbre encoder. The target prosody information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker using the prosody encoder. The target content information corresponding to the first speaker is obtained by extracting the text data corresponding to the sample audio data of the first speaker using the content encoder. The second reconstruction loss is determined based on the sample audio data of the second speaker and the reconstructed audio data of the second speaker, the reconstructed audio data of the second speaker is reconstructed based on the target timbre information corresponding to the second speaker, the target prosody information corresponding to the second speaker and the target content information corresponding to the second speaker, the target timbre information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker based on the timbre encoder, the target prosody information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker based on the prosody encoder, and the target content information corresponding to the second speaker is obtained by extracting the text data corresponding to the sample audio data of the second speaker based on the content encoder.

[0092] Based on this, the first reconstruction loss can also constrain the content encoder's ability to preserve the audio data, thereby improving the adequacy of content decoupling. The second reconstruction loss can also constrain the content encoder's ability to preserve the audio data, thereby improving the adequacy of content decoupling.

[0093] Furthermore, during the training process of the prosody encoder, the sample audio data of the first speaker and the sample audio data of the second speaker are respectively processed with timbre change to obtain the sample audio data of the first speaker after timbre change and the sample audio data of the second speaker after timbre change; the sample audio data of the first speaker after timbre change and the sample audio data of the second speaker after timbre change are respectively input into the prosody encoder to obtain the target prosody information output by the prosody encoder respectively; based on this, it is ensured that the prosody encoder focuses only on the extraction of prosody information as much as possible, thereby reducing the leakage of timbre information and content information.

[0094] In the speech decoupling method provided by the embodiment of the present invention, the timbre encoder and the prosody encoder are obtained by training based on the first reconstruction loss and the second reconstruction loss, and the first reconstruction loss is determined based on the sample audio data of the first speaker and the reconstructed audio data of the first speaker, the reconstructed audio data of the first speaker is reconstructed based on the target timbre information corresponding to the first speaker and the target prosody information corresponding to the first speaker, and the target timbre information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker based on the timbre encoder, and the target prosody information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker based on the prosody encoder, so that the first reconstruction loss can constrain the timbre maintenance ability of the timbre encoder for the first speaker, thereby improving the adequacy of timbre decoupling, and can constrain the prosody maintenance ability of the prosody encoder for the first speaker, thereby improving the adequacy of prosody decoupling; at the same time, the second reconstruction loss is determined based on the second reconstruction loss. The first speaker and the second speaker are determined by the sample audio data of the speaker and the reconstructed audio data of the second speaker, the reconstructed audio data of the second speaker is reconstructed based on the target timbre information corresponding to the second speaker and the target prosody information corresponding to the second speaker, and the target timbre information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker based on the timbre encoder, and the target prosody information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker based on the prosody encoder, so that the second reconstruction loss can constrain the timbre retention ability of the timbre encoder for the second speaker, thereby improving the adequacy of timbre decoupling, and can constrain the prosody retention ability of the prosody encoder for the second speaker, thereby improving the adequacy of prosody decoupling; in addition, the first speaker and the second speaker are different speakers, so that the model training can be performed based on the sample audio data of different speakers, that is, training can be performed based on the audio data of a large number of speakers, thereby improving the generalization of speech decoupling.

[0095] Based on any of the above embodiments, Figure 2 This is the second flow chart of the voice decoupling method provided by the present invention, such as Figure 2 As shown, the timbre encoder and the prosody encoder are trained based on the following method, which includes steps 210 to 260.

[0096] In step 210, the sample audio data of the first speaker and the sample audio data of the second speaker are respectively input into the timbre encoder in the acoustic model to obtain the target timbre information respectively output by the timbre encoder, and the sample audio data of the first speaker and the sample audio data of the second speaker are respectively input into the prosody encoder in the acoustic model to obtain the target prosody information respectively output by the prosody encoder.

[0097] Specifically, the sample audio data of the first speaker is input into the timbre encoder to obtain the target timbre information corresponding to the first speaker output by the timbre encoder; the sample audio data of the second speaker is input into the timbre encoder to obtain the target timbre information corresponding to the second speaker output by the timbre encoder; the sample audio data of the first speaker is input into the prosody encoder to obtain the target prosody information corresponding to the first speaker output by the prosody encoder; the sample audio data of the second speaker is input into the prosody encoder to obtain the target prosody information corresponding to the second speaker output by the prosody encoder.

[0098] Furthermore, the sample audio data of the first speaker and the sample audio data of the second speaker are respectively inputted into the content encoder in the acoustic model to obtain target content information respectively outputted by the content encoder.

[0099] Step 220 : Input the target timbre information corresponding to the first speaker and the target prosody information corresponding to the second speaker into the audio decoder in the acoustic model to obtain first reconstructed audio data output by the audio decoder.

[0100] The audio decoder is used to reconstruct audio data based on input information, that is, to reconstruct first reconstructed audio data based on target timbre information corresponding to the first speaker and target prosody information corresponding to the second speaker.

[0101] In one embodiment, the audio decoder comprises a plurality of residual convolution modules with skip connections.

[0102] Furthermore, the target timbre information corresponding to the first speaker, the target prosody information corresponding to the second speaker, and the target content information corresponding to the sample audio data of the second speaker are input into the audio decoder to obtain first reconstructed audio data output by the audio decoder.

[0103] Step 230: Input the first reconstructed audio data into the timbre encoder to obtain reconstructed timbre information output by the timbre encoder, and input the first reconstructed audio data into the prosody encoder to obtain reconstructed prosody information output by the prosody encoder.

[0104] It should be noted that, since the first reconstructed audio data corresponds to the timbre of the first speaker, the extracted reconstructed timbre information also corresponds to the timbre of the first speaker. Since the first reconstructed audio data corresponds to the rhythm of the sample audio data of the second speaker, the extracted reconstructed rhythm information also corresponds to the rhythm of the sample audio data of the second speaker.

[0105] Step 240: Input the reconstructed timbre information and the target prosody information corresponding to the first speaker into the audio decoder to obtain second reconstructed audio data output by the audio decoder, and input the target timbre information and the reconstructed prosody information corresponding to the second speaker into the audio decoder to obtain third reconstructed audio data output by the audio decoder.

[0106] Furthermore, the reconstructed timbre information, the target prosody information corresponding to the first speaker, and the target content information corresponding to the sample audio data of the first speaker are input into the audio decoder to obtain second reconstructed audio data output by the audio decoder, and the target timbre information, the reconstructed prosody information corresponding to the second speaker, and the target content information corresponding to the sample audio data of the second speaker are input into the audio decoder to obtain third reconstructed audio data output by the audio decoder.

[0107] Step 250 : Determine a first reconstruction loss based on the sample audio data of the first speaker and the second reconstructed audio data, and determine a second reconstruction loss based on the sample audio data of the second speaker and the third reconstructed audio data.

[0108] Based on this, the first reconstruction loss can constrain the timbre preservation capability of the timbre encoder for the first speaker, thereby improving the adequacy of timbre decoupling, and can constrain the prosody preservation capability of the prosody encoder for the first speaker, thereby improving the adequacy of prosody decoupling. The second reconstruction loss can constrain the timbre preservation capability of the timbre encoder for the second speaker, thereby improving the adequacy of timbre decoupling, and can constrain the prosody preservation capability of the prosody encoder for the second speaker, thereby improving the adequacy of prosody decoupling.

[0109] Step 260: Train the acoustic model based on the first reconstruction loss and the second reconstruction loss.

[0110] Specifically, based on the first reconstruction loss and the second reconstruction loss, each network layer in the acoustic model is optimized, that is, the timbre encoder, the prosody encoder and the audio decoder are optimized.

[0111] Furthermore, the first, second, and third reconstructed audio data are respectively input into a discriminator to obtain a discrimination loss output by the discriminator. The acoustic model is trained based on the first, second, and discrimination losses. This allows training based not only on the reconstruction loss but also on the discrimination loss, thereby improving the decoupling adequacy of the timbre encoder, improving the decoupling adequacy of the prosody encoder, and improving the audio reconstruction accuracy of the audio decoder.

[0112] It should be noted that the training process of the acoustic model can be performed on the execution body of the speech decoupling method, or on other training devices, and then the trained timbre encoder and prosody encoder can be deployed to the execution body of the speech decoupling method.

[0113] The speech decoupling method provided by the embodiment of the present invention, through the above-mentioned manner, the first reconstruction loss can constrain the timbre retention capability of the timbre encoder for the first speaker, thereby improving the adequacy of timbre decoupling, and can constrain the prosody retention capability of the prosody encoder for the first speaker, thereby improving the adequacy of prosody decoupling; at the same time, the second reconstruction loss can constrain the timbre retention capability of the timbre encoder for the second speaker, thereby improving the adequacy of timbre decoupling, and can constrain the prosody retention capability of the prosody encoder for the second speaker, thereby improving the adequacy of prosody decoupling; in addition, the first speaker and the second speaker are different speakers, so the model training can be performed based on the sample audio data of different speakers, that is, the training can be performed based on the audio data of a large number of speakers, thereby improving the generalization of speech decoupling; in addition, the sample audio data of the first speaker and the sample audio data of the second speaker are first audio-trained. The timbre information is extracted and the rhythm information is extracted, and then the target timbre information corresponding to the first speaker and the target rhythm information corresponding to the second speaker are reconstructed to obtain the first reconstructed audio data, and then the timbre information and rhythm information are extracted again on the first reconstructed audio data, and then the second reconstructed audio data is obtained based on the reconstructed timbre information and the target rhythm information corresponding to the first speaker, and the third reconstructed audio data is obtained based on the target timbre information and the reconstructed rhythm information corresponding to the second speaker. Thus, the timbre encoder performs timbre extraction twice, and then the first reconstruction loss can constrain the timbre encoder's ability to maintain the timbre of the first speaker twice, thereby further improving the adequacy of timbre decoupling. At the same time, the rhythm encoder performs rhythm extraction twice, and the second reconstruction loss can constrain the rhythm encoder's ability to maintain the rhythm of the second speaker twice, thereby further improving the adequacy of rhythm decoupling.

[0114] Based on any of the above embodiments, in this method, the sample audio data of the first speaker includes different first audio data and second audio data, and the sample audio data of the second speaker includes different third audio data and fourth audio data.

[0115] In other words, two audio data sets are obtained from two different speakers. Furthermore, each audio data set may be mel-spectrometric data, i.e., the first audio data set, the second audio data set, the third audio data set, and the fourth audio data set are all mel-spectrometric data. For example, the mel-spectrometric data is obtained by performing a short-time Fourier transform on the audio data to obtain an amplitude spectrum, which is then converted to a mel-scale.

[0116] Since the first audio data and the second audio data are audio data of the same speaker, the timbre information of the first audio data and the second audio data is the same, but the prosody information of the first audio data and the second audio data is different. Since the third audio data and the fourth audio data are audio data of the same speaker, the timbre information of the third audio data and the fourth audio data is the same, but the prosody information of the third audio data and the fourth audio data is different.

[0117] Accordingly, if Figure 3 As shown, the above-mentioned step 210 includes: inputting the second audio data and the fourth audio data into the timbre encoder respectively to obtain the target timbre information output by the timbre encoder respectively, and inputting the first audio data and the third audio data into the rhythm encoder respectively to obtain the target rhythm information output by the rhythm encoder respectively.

[0118] Specifically, the second audio data of the first speaker is input into the timbre encoder to obtain the target timbre information corresponding to the second audio data of the first speaker output by the timbre encoder; the fourth audio data of the second speaker is input into the timbre encoder to obtain the target timbre information corresponding to the fourth audio data of the second speaker output by the timbre encoder; the first audio data of the first speaker is input into the rhythm encoder to obtain the target rhythm information corresponding to the first audio data of the first speaker output by the rhythm encoder; the third audio data of the second speaker is input into the rhythm encoder to obtain the target rhythm information corresponding to the third audio data of the second speaker output by the rhythm encoder.

[0119] Accordingly, if Figure 3 As shown, the above step 220 includes: inputting the target timbre information corresponding to the second audio data of the first speaker and the target prosody information corresponding to the third audio data of the second speaker into the audio decoder to obtain the first reconstructed audio data output by the audio decoder.

[0120] Based on this, the timbre information of the first reconstructed data is the timbre information of the first speaker, and the prosody information of the first reconstructed data is the prosody information of the third audio data of the second speaker.

[0121] Accordingly, if Figure 3 As shown, the above-mentioned step 240 includes: inputting the reconstructed timbre information and the target prosody information corresponding to the first audio data of the first speaker into the audio decoder to obtain second reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the fourth audio data of the second speaker into the audio decoder to obtain third reconstructed audio data output by the audio decoder.

[0122] Accordingly, if Figure 3 As shown, the above step 250 includes: determining a first reconstruction loss based on the first audio data of the first speaker and the second reconstructed audio data, and determining a second reconstruction loss based on the third audio data of the second speaker and the third reconstructed audio data.

[0123] The speech decoupling method provided by the embodiment of the present invention has the following characteristics: since the first audio data and the second audio data are audio data of the same speaker, the timbre information of the first audio data and the second audio data are the same, but the prosody information of the first audio data and the second audio data are different. Therefore, the second reconstructed audio data is obtained by reconstructing the timbre information of the second audio data of the first speaker and the prosody information of the first audio data of the first speaker in the above manner, and then the first reconstruction loss is determined based on the first audio data and the second reconstructed audio data of the first speaker. Therefore, the first reconstruction loss can constrain the stability of the timbre extraction of different voices of the same person by the timbre encoder, thereby improving the timbre decomposition. coupling stability; at the same time, since the third audio data and the fourth audio data are audio data of the same speaker, the timbre information of the third audio data and the fourth audio data is the same, but the prosody information of the third audio data and the fourth audio data is different. Therefore, the third reconstructed audio data is obtained by reconstructing based on the timbre information of the fourth audio data of the second speaker and the prosody information of the third audio data of the second speaker in the above manner, and then the second reconstruction loss is determined based on the third audio data of the second speaker and the third reconstructed audio data. Therefore, the second reconstruction loss can constrain the prosody extraction stability of the prosody encoder for the same prosody audio of different people, thereby improving the stability of prosody decoupling.

[0124] Based on any of the above embodiments, Figure 4 As shown, the above step 210 includes: inputting the second audio data and the fourth audio data into the timbre encoder respectively to obtain the target timbre information output by the timbre encoder respectively, and inputting the first audio data and the third audio data into the prosody encoder respectively to obtain the target prosody information output by the prosody encoder respectively, and inputting the first text data corresponding to the first audio data and the third text data corresponding to the third audio data into the content encoder in the acoustic model respectively to obtain the target content information output by the content encoder respectively.

[0125] Here, the first text data or the third text data may be text transcribed by a speech recognition system or manually annotated text.

[0126] Specifically, the first text data corresponding to the first audio data is input into the content encoder to obtain the target content information corresponding to the first audio data of the first speaker output by the content encoder; the third text data corresponding to the third audio data is input into the content encoder to obtain the target content information corresponding to the third audio data of the second speaker output by the content encoder.

[0127] Accordingly, if Figure 4 As shown, the above-mentioned step 220 includes: inputting the target timbre information corresponding to the second audio data of the first speaker, the target prosody information corresponding to the third audio data of the second speaker, and the target content information corresponding to the third audio data of the second speaker into the audio decoder to obtain the first reconstructed audio data output by the audio decoder.

[0128] Based on this, the timbre information of the first reconstructed data is the timbre information of the first speaker, the prosody information of the first reconstructed data is the prosody information of the third audio data of the second speaker, and the content information of the first reconstructed data is the content information of the third audio data of the second speaker. This ensures that the timbre information and prosody information correspond to different speakers, but the prosody information and content information correspond to the same speaker and audio data.

[0129] Accordingly, if Figure 4 As shown, the above-mentioned step 240 includes: inputting the reconstructed timbre information and the target prosody information corresponding to the first audio data of the first speaker, and the target content information corresponding to the first audio data of the first speaker into the audio decoder to obtain second reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the fourth audio data of the second speaker, and the target content information corresponding to the third audio data of the second speaker into the audio decoder to obtain third reconstructed audio data output by the audio decoder.

[0130] It should be understood that the second reconstructed audio data should maintain the timbre of the first reconstructed audio data, that is, the timbre of the first speaker, while maintaining the prosody and content of the first audio data of the first speaker. The third reconstructed audio data should maintain the timbre of the fourth audio data of the second speaker, that is, the timbre of the second speaker, while maintaining the prosody of the third audio data, that is, the prosody of the third audio data of the second speaker, while maintaining the content of the third audio data of the second speaker.

[0131] Accordingly, based on the first reconstruction loss and the second reconstruction loss, each network layer in the acoustic model is optimized, that is, the timbre encoder, the prosody encoder, the audio decoder and the audio decoder are optimized.

[0132] Accordingly, the speech data to be decoupled is input into the trained content encoder to obtain the decoupled content information output by the content encoder.

[0133] In the speech decoupling method provided by an embodiment of the present invention, the timbre encoder, prosody encoder, and content encoder are trained based on a first reconstruction loss and a second reconstruction loss. Through the above-mentioned method, the first reconstruction loss and the second reconstruction loss can constrain the content encoder's ability to retain the content of the audio data, thereby improving the adequacy of content decoupling.

[0134] Based on any of the above embodiments, in the method, step 260 includes:

[0135] Inputting the reconstructed timbre information and the target prosody information corresponding to the second speaker into the audio decoder to obtain fourth reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the first speaker into the audio decoder to obtain fifth reconstructed audio data output by the audio decoder;

[0136] determining a third reconstruction loss based on the first reconstructed audio data and the fourth reconstructed audio data, and determining a fourth reconstruction loss based on the first reconstructed audio data and the fifth reconstructed audio data;

[0137] The acoustic model is trained based on the first reconstruction loss, the second reconstruction loss, the third reconstruction loss, and the fourth reconstruction loss.

[0138] It should be understood that the fourth reconstructed audio data should maintain the timbre of the first reconstructed audio data, that is, the timbre of the first speaker, while maintaining the prosody and content of the third audio data of the second speaker. The fifth reconstructed audio data should maintain the timbre of the first speaker, while maintaining the prosody of the third audio data, that is, the prosody of the third audio data of the second speaker, while maintaining the content of the third audio data of the second speaker.

[0139] Based on this, the third reconstruction loss can constrain the timbre preservation capability of the timbre encoder for the first speaker, thereby improving the adequacy of timbre decoupling, and can constrain the prosody preservation capability of the prosody encoder for the second speaker, thereby improving the adequacy of prosody decoupling. The fourth reconstruction loss can constrain the timbre preservation capability of the timbre encoder for the first speaker, thereby improving the adequacy of timbre decoupling, and can constrain the prosody preservation capability of the prosody encoder for the second speaker, thereby improving the adequacy of prosody decoupling.

[0140] Specifically, based on the first reconstruction loss, the second reconstruction loss, the third reconstruction loss and the fourth reconstruction loss, each network layer in the acoustic model is optimized, that is, the timbre encoder, the prosody encoder and the audio decoder are optimized.

[0141] Furthermore, the sample audio data of the first speaker includes different first audio data and second audio data, and the sample audio data of the second speaker includes different third audio data and fourth audio data. The step of inputting the sample audio data of the first speaker and the sample audio data of the second speaker into a timbre encoder in the acoustic model to obtain target timbre information output by the timbre encoder, and inputting the sample audio data of the first speaker and the sample audio data of the second speaker into a prosody encoder in the acoustic model to obtain target prosody information output by the prosody encoder comprises: inputting the second audio data and the fourth audio data into the timbre encoder to obtain target timbre information output by the timbre encoder, and inputting the first audio data and the third audio data into the prosody encoder to obtain target prosody information output by the prosody encoder. Inputting the target timbre information corresponding to the first speaker and the target prosody information corresponding to the second speaker into the audio decoder in the acoustic model to obtain first reconstructed audio data output by the audio decoder includes: inputting the target timbre information corresponding to the second audio data of the first speaker and the target prosody information corresponding to the third audio data of the second speaker into the audio decoder to obtain first reconstructed audio data output by the audio decoder. Inputting the reconstructed timbre information and the target prosody information corresponding to the second speaker into the audio decoder to obtain fourth reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the first speaker into the audio decoder to obtain fifth reconstructed audio data output by the audio decoder includes: inputting the reconstructed timbre information and the target prosody information corresponding to the third audio data of the second speaker into the audio decoder to obtain fourth reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the second audio data of the first speaker into the audio decoder to obtain fifth reconstructed audio data output by the audio decoder.

[0142] Furthermore, the reconstructed timbre information and the target prosody information corresponding to the second speaker, as well as the target content information corresponding to the sample audio data of the second speaker are input into the audio decoder to obtain the fourth reconstructed audio data output by the audio decoder, and the target timbre information and the reconstructed prosody information corresponding to the first speaker, as well as the target content information corresponding to the sample audio data of the second speaker are input into the audio decoder to obtain the fifth reconstructed audio data output by the audio decoder.

[0143] The speech decoupling method provided by an embodiment of the present invention, through the above-mentioned method, the third reconstruction loss can constrain the timbre encoder's ability to maintain the timbre of the first speaker, thereby improving the adequacy of timbre decoupling, and compared with the first reconstruction loss, the reconstruction target is changed from real audio data to first reconstructed audio data, and the transmission path of the gradient flow is shorter, thereby improving the stability of model training, that is, improving the stability of timbre decoupling; at the same time, the fourth reconstruction loss can constrain the prosody encoder's ability to maintain the prosody of the second speaker, thereby improving the adequacy of prosody decoupling, and compared with the second reconstruction loss, the reconstruction target is changed from real audio data to first reconstructed audio data, and the transmission path of the gradient flow is shorter, thereby improving the stability of model training, that is, improving the stability of prosody decoupling.

[0144] Based on any of the above embodiments, in this method, the sample audio data of the first speaker includes different first audio data and second audio data, and the sample audio data of the second speaker includes different third audio data and fourth audio data.

[0145] Accordingly, the above step 210 includes: inputting the second audio data and the fourth audio data into the timbre encoder respectively to obtain target timbre information respectively output by the timbre encoder, inputting the first audio data and the third audio data into the prosody encoder respectively to obtain target prosody information respectively output by the prosody encoder, and inputting the first text data corresponding to the first audio data and the third text data corresponding to the third audio data into the content encoder in the acoustic model respectively to obtain target content information respectively output by the content encoder;

[0146] Correspondingly, the above-mentioned step 220 includes: inputting the target timbre information corresponding to the second audio data of the first speaker, the target rhythm information corresponding to the third audio data of the second speaker, and the target content information corresponding to the third audio data of the second speaker into the audio decoder to obtain the first reconstructed audio data output by the audio decoder.

[0147] Correspondingly, the step of inputting the reconstructed timbre information and the target prosody information corresponding to the second speaker into the audio decoder to obtain fourth reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the first speaker into the audio decoder to obtain fifth reconstructed audio data output by the audio decoder includes: inputting the reconstructed timbre information and the target prosody information corresponding to the third audio data of the second speaker, and the target content information corresponding to the third audio data of the second speaker into the audio decoder to obtain fourth reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the second audio data of the first speaker, and the target content information corresponding to the third audio data of the second speaker into the audio decoder to obtain fifth reconstructed audio data output by the audio decoder.

[0148] In the speech decoupling method provided by an embodiment of the present invention, the timbre encoder, prosody encoder, and content encoder are trained based on a first reconstruction loss, a second reconstruction loss, a third reconstruction loss, and a fourth reconstruction loss. Through the above-mentioned method, the first reconstruction loss, the second reconstruction loss, the third reconstruction loss, and the fourth reconstruction loss can constrain the content encoder's ability to retain the content of the audio data, thereby improving the adequacy of content decoupling.

[0149] To facilitate understanding of the above embodiments, a specific embodiment is described here. First, obtain the first audio data A1 of the first speaker and the second audio data A2 of the first speaker, as well as the third audio data B1 of the second speaker and the fourth audio data B2 of the second speaker, and the first text data corresponding to the first audio data A1. The third text data corresponding to the third audio data B1 Secondly, the second audio data A2 and the fourth audio data B2 are input to the timbre encoder, and the timbre encoder outputs the target timbre information corresponding to the second audio data A2. The target timbre information corresponding to the fourth audio data B2 The first audio data A1 and the third audio data B1 are respectively input to the prosody encoder, and the target prosody information corresponding to the first audio data A1 output by the prosody encoder is obtained. Target prosody information corresponding to the third audio data B1 , and the first text data corresponding to the first audio data A1 The third text data corresponding to the third audio data B1 Input them into the content encoder respectively, and obtain the target content information corresponding to the first audio data A1 output by the content encoder respectively Target content information corresponding to the third audio data B1 Afterwards, the target timbre information corresponding to the second audio data A2 of the first speaker Target prosody information corresponding to the third audio data B1 of the second speaker , and the target content information corresponding to the third audio data B1 of the second speaker Input the first reconstructed audio data G1 to the audio decoder to obtain the first reconstructed audio data G1 output by the audio decoder; then, input the first reconstructed audio data G1 to the timbre encoder to obtain the reconstructed timbre information output by the timbre encoder The first reconstructed audio data G1 is input to the prosody encoder to obtain the reconstructed prosody information output by the prosody encoder. ; Afterwards, the timbre information will be reconstructed Target prosody information corresponding to the first audio data A1 of the first speaker , and the target content information corresponding to the first audio data A1 Input to the audio decoder, obtain the second reconstructed audio data G2 output by the audio decoder, and convert the target timbre information corresponding to the fourth audio data B2 of the second speaker into the target timbre information and reconstructing prosodic information Input to the audio decoder, obtain the third reconstructed audio data G3 output by the audio decoder, and reconstruct the timbre information Target prosody information corresponding to the third audio data B1 of the second speaker , and the target content information corresponding to the third audio data B1 of the second speaker The audio data is input to the audio decoder to obtain the fourth reconstructed audio data G4 output by the audio decoder, and the target timbre information corresponding to the second audio data A2 of the first speaker is converted into the target timbre information and reconstructing prosodic information , and the target content information corresponding to the third audio data B1 of the second speaker The audio data is input to the audio decoder to obtain the fifth reconstructed audio data G5 output by the audio decoder.

[0150] After determining the first reconstruction loss, the second reconstruction loss, the third reconstruction loss and the fourth reconstruction loss based on the above, a total reconstruction loss is determined based on the first reconstruction loss, the second reconstruction loss, the third reconstruction loss and the fourth reconstruction loss, and the acoustic model is trained based on the total reconstruction loss.

[0151] Exemplarily, the total reconstruction loss is as follows:

[0152] ;

[0153] Where, represents the total reconstruction loss, represents the first reconstruction loss, represents the second reconstruction loss, represents the third reconstruction loss, represents the fourth reconstruction loss.

[0154] Among them, the reconstruction losses are as follows:

[0155] ;

[0156] ;

[0157] ;

[0158] ;

[0159] Where, first audio data representing a first speaker, represents the second reconstructed audio data, Represents the reconstruction of timbre information, represents the target prosody information corresponding to the first audio data A1, Indicates the target content information corresponding to the first audio data A1, Represents the audio decoder, third audio data representing a second speaker, represents the third reconstructed audio data, represents the target timbre information corresponding to the fourth audio data B2, Represents the reconstructed prosodic information, Indicates the target content information corresponding to the third audio data B1, the first reconstructed audio data , represents the target timbre information corresponding to the second audio data A2, represents the target prosody information corresponding to the third audio data B1, represents the fourth reconstructed audio data, represents the fifth reconstructed audio data.

[0160] Among them, the first reconstruction loss Constraining the first reconstruction audio data G1 to maintain the timbre of the first speaker A, showing the stability of the model in extracting the timbre of different voices of the same person; the second reconstruction loss Constraining the ability of the first reconstructed audio data G1 to maintain the rhythm of the third audio data B1 of the second speaker B, showing the stability of the model's rhythm extraction of the same rhythmic audio of different speakers; the third reconstruction loss It also constrains the timbre preservation ability of the first reconstructed audio data G1 for the first speaker A, and Compared with the original audio data, the reconstruction target is changed from real audio data to reconstructed audio data. The transfer path of the gradient flow is shorter, which helps to improve the stability of training. The fourth reconstruction loss It is also a constraint on the ability to preserve the prosody of the third audio data B1 of the second speaker B in the first reconstructed audio data G1, and In comparison, the reconstruction target is changed from real audio data to reconstructed audio data, and the gradient flow transmission path is shorter, which helps to improve the stability of training.

[0161] Based on any of the foregoing embodiments, in the method, training the acoustic model based on the first reconstruction loss, the second reconstruction loss, the third reconstruction loss, and the fourth reconstruction loss includes:

[0162] Inputting the first reconstructed audio data, the second reconstructed audio data, the third reconstructed audio data, the fourth reconstructed audio data, and the fifth reconstructed audio data into a discriminator respectively, and obtaining five discrimination losses output by the discriminator;

[0163] The acoustic model is trained based on the first reconstruction loss, the second reconstruction loss, the third reconstruction loss, the fourth reconstruction loss, and the five discriminant losses.

[0164] Here, the discriminator is used to discriminate the audio reconstruction accuracy of the audio decoder, thereby training the acoustic model based on five discriminant losses to improve the robustness of each network layer in the acoustic model, that is, to improve the decoupling adequacy of the timbre encoder, to improve the decoupling adequacy of the prosody encoder, and to improve the audio reconstruction accuracy of the audio decoder, and to further improve the training effect of the acoustic model while improving the audio reconstruction accuracy.

[0165] The discriminator can be pre-trained, that is, its optimized parameters are used to update the discriminator, and then the audio decoder (generator) is optimized based on the discriminant loss (adversarial loss). In other words, the training of the acoustic model includes not only the reconstruction loss but also the discriminant loss (adversarial loss).

[0166] It should be understood that since the audio decoder is reconstructed five times in one training round, five discriminant losses can be obtained based on this, thereby improving the training effect while ensuring training efficiency.

[0167] The speech decoupling method provided in an embodiment of the present invention, through the above-mentioned method, is trained not only based on the reconstruction loss, but also based on the discrimination loss, thereby improving the decoupling adequacy of the timbre encoder, improving the decoupling adequacy of the prosody encoder, and improving the audio reconstruction accuracy of the audio decoder.

[0168] Based on any of the above embodiments, in the method, in the above step 120, the sample audio data of the first speaker and the sample audio data of the second speaker are respectively input into the prosody encoder in the acoustic model to obtain the target prosody information respectively output by the prosody encoder, including:

[0169] Performing timbre change processing on the sample audio data of the first speaker and the sample audio data of the second speaker, respectively, to obtain the sample audio data of the first speaker after timbre change and the sample audio data of the second speaker after timbre change;

[0170] The sample audio data of the first speaker after the timbre change and the sample audio data of the second speaker after the timbre change are respectively inputted into a prosody encoder in the acoustic model to obtain target prosody information respectively outputted by the prosody encoder.

[0171] In one embodiment, the timbre change processing method is random pitch conversion to obtain the pitch-shifted audio data.

[0172] Correspondingly, in the above step 230, the first reconstructed audio data is input into the prosody encoder to obtain the reconstructed prosody information output by the prosody encoder, including: performing timbre change processing on the first reconstructed audio data to obtain the first reconstructed audio data after the timbre change; inputting the first reconstructed audio data after the timbre change into the prosody encoder to obtain the reconstructed prosody information output by the prosody encoder.

[0173] It should be noted that the timbre of the input audio data needs to be changed only during the training of the prosody encoder. This is because the encoder has learned to focus only on the extraction of prosody information during training, thereby reducing the leakage of timbre and content information.

[0174] Furthermore, the audio data is represented using mel-spectrogram data. The lower 20 dimensions of the audio data from the first speaker after timbre modification and the lower 20 dimensions of the audio data from the second speaker after timbre modification are input into the prosody encoder, respectively, to obtain the target prosody information output by the prosody encoder. This is because the lower 20 dimensions of the mel-spectrogram data retain the complete fundamental frequency and some harmonics, representing the prosody features while also compressing information, thereby improving the decoupling efficiency of the prosody encoder.

[0175] The speech decoupling method provided by an embodiment of the present invention performs timbre change processing on sample audio data of a first speaker and sample audio data of a second speaker during the training process of a prosody encoder to obtain the sample audio data of the first speaker after the timbre change and the sample audio data of the second speaker after the timbre change; the sample audio data of the first speaker after the timbre change and the sample audio data of the second speaker after the timbre change are respectively input into the prosody encoder to obtain target prosody information respectively output by the prosody encoder, thereby ensuring that the prosody encoder focuses only on the extraction of prosody information as much as possible, thereby reducing the leakage of timbre information and content information.

[0176] Based on any of the above embodiments, in this method, the audio data is represented by mel-spectrogram data. When the audio data needs to be input into a prosody encoder, the lower 20 dimensions of the input audio data are input into the prosody encoder. Based on this, considering that the lower 20 dimensions of the mel-spectrogram data retain the complete fundamental frequency and some harmonics, they can represent prosodic features while also compressing information, thereby improving the decoupling efficiency of the prosody encoder.

[0177] Based on any of the above embodiments, in the method, the prosody encoder includes a first residual module, a downsampling module, a second residual module and a vector quantization module connected in sequence.

[0178] In one embodiment, the first residual module and the second residual module are each composed of a plurality of residual convolution modules with jump connections, such as a ResBlock module or a ConvNext module. Furthermore, each convolutional layer in the residual convolution module is composed of a plurality of convolution units, and the parameters of each convolution unit are optimized by a back-propagation algorithm. The purpose of the convolution operation is to extract different features of the input. The first convolution layer may only be able to extract some low-level features such as edges, lines, and corners. A network with more layers can iteratively extract more complex features from low-level features.

[0179] In one embodiment, the output of the first residual module is fed into a first downsampling module, which includes several convolutional and pooling layers to further compress the audio data in the temporal dimension. The output of the downsampling module is then fed into a second residual module for further processing of hidden features.

[0180] Here, the output of the second residual module is processed by the vector quantization module to quantize the prosody features into several discrete vectors, further compressing the information and reducing the difficulty of prosody prediction. The quantized vectors can now be used as prosody information.

[0181] The speech decoupling method provided in the embodiment of the present invention can further improve the decoupling adequacy and accuracy of the prosody encoder through the structure of the above-mentioned prosody encoder.

[0182] Based on the above embodiments, the acoustic model trained using the above training method has speech decoupling capabilities. Specifically, the timbre encoder can extract speaker timbre information in addition to prosody and content from any audio; the prosody encoder can extract prosody information in addition to timbre and content from any audio; and the content encoder can extract content information from text. These three decoupled information can be arbitrarily combined to accomplish downstream speech synthesis tasks such as voice cloning, timbre transfer, and prosody transfer.

[0183] The speech decoupling device provided by the present invention is described below. The speech decoupling device described below and the speech decoupling method described above can be referenced to each other.

[0184] Figure 5 Schematic diagram of the structure of the speech decoupling device provided by the present invention, such as Figure 5 As shown, the speech decoupling device includes a speech decoupling module 510.

[0185] The speech decoupling module 510 is configured to input the speech data to be decoupled into the timbre encoder and the prosody encoder respectively, to obtain decoupled timbre information output by the timbre encoder and decoupled prosody information output by the prosody encoder.

[0186] The timbre encoder and the prosody encoder are trained based on a first reconstruction loss and a second reconstruction loss; the first reconstruction loss is determined based on sample audio data of a first speaker and the reconstructed audio data of the first speaker, the reconstructed audio data of the first speaker is reconstructed based on target timbre information corresponding to the first speaker and target prosody information corresponding to the first speaker, the target timbre information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker by the timbre encoder, and the target prosody information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker by the prosody encoder; the second reconstruction loss is determined based on sample audio data of a second speaker and the reconstructed audio data of the second speaker, the reconstructed audio data of the second speaker is reconstructed based on target timbre information corresponding to the second speaker and target prosody information corresponding to the second speaker, the target timbre information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker by the timbre encoder, and the target prosody information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker by the prosody encoder; the first speaker and the second speaker are different speakers.

[0187] Based on any of the above embodiments, the device further includes a training module, which includes:

[0188] a first input unit for inputting the sample audio data of the first speaker and the sample audio data of the second speaker into a timbre encoder in the acoustic model, respectively, to obtain target timbre information outputted by the timbre encoder, and for inputting the sample audio data of the first speaker and the sample audio data of the second speaker into a prosody encoder in the acoustic model, respectively, to obtain target prosody information outputted by the prosody encoder;

[0189] A second input unit is configured to input the target timbre information corresponding to the first speaker and the target prosody information corresponding to the second speaker into an audio decoder in the acoustic model to obtain first reconstructed audio data output by the audio decoder; the audio decoder is configured to reconstruct the audio data based on the input information;

[0190] a third input unit, configured to input the first reconstructed audio data into the timbre encoder to obtain reconstructed timbre information output by the timbre encoder, and input the first reconstructed audio data into the prosody encoder to obtain reconstructed prosody information output by the prosody encoder;

[0191] a fourth input unit, configured to input the reconstructed timbre information and the target prosody information corresponding to the first speaker into the audio decoder to obtain second reconstructed audio data output by the audio decoder, and input the target timbre information and the reconstructed prosody information corresponding to the second speaker into the audio decoder to obtain third reconstructed audio data output by the audio decoder;

[0192] a loss determining unit configured to determine a first reconstruction loss based on the sample audio data of the first speaker and the second reconstructed audio data, and to determine a second reconstruction loss based on the sample audio data of the second speaker and the third reconstructed audio data;

[0193] A model training unit is used to train the acoustic model based on the first reconstruction loss and the second reconstruction loss.

[0194] Figure 6 An example of a physical structure diagram of an electronic device is shown below. Figure 6As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630 and a communication bus 640, wherein the processor 610, the communications interface 620 and the memory 630 communicate with each other via the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute a speech decoupling method, which includes: inputting the speech data to be decoupled into a timbre encoder and a prosody encoder respectively, obtaining decoupled timbre information output by the timbre encoder and decoupled prosody information output by the prosody encoder; wherein the timbre encoder and the prosody encoder are trained based on a first reconstruction loss and a second reconstruction loss; the first reconstruction loss is determined based on sample audio data of a first speaker and reconstructed audio data of the first speaker, the reconstructed audio data of the first speaker is reconstructed based on target timbre information corresponding to the first speaker and target prosody information corresponding to the first speaker, and the target timbre information corresponding to the first speaker is based on the timbre encoder's training of the sample audio data of the first speaker. The target prosody information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker based on the prosody encoder; the second reconstruction loss is determined based on the sample audio data of the second speaker and the reconstructed audio data of the second speaker, the reconstructed audio data of the second speaker is reconstructed based on the target timbre information corresponding to the second speaker and the target prosody information corresponding to the second speaker, the target timbre information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker based on the timbre encoder, and the target prosody information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker based on the prosody encoder; the first speaker and the second speaker are different speakers.

[0195] Furthermore, the logic instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0196] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the speech decoupling method provided by the above methods, which includes: inputting the speech data to be decoupled into a timbre encoder and a prosody encoder respectively, and obtaining the decoupled timbre information output by the timbre encoder, and the decoupled prosody information output by the prosody encoder; wherein the timbre encoder and the prosody encoder are trained based on a first reconstruction loss and a second reconstruction loss; the first reconstruction loss is determined based on the sample audio data of the first speaker and the reconstructed audio data of the first speaker, and the reconstructed audio data of the first speaker is reconstructed based on the target timbre information corresponding to the first speaker and the target prosody information corresponding to the first speaker, and the first speaker The target timbre information corresponding to the speaker is obtained based on the timbre encoder extracting the sample audio data of the first speaker, and the target prosody information corresponding to the first speaker is obtained based on the prosody encoder extracting the sample audio data of the first speaker; the second reconstruction loss is determined based on the sample audio data of the second speaker and the reconstructed audio data of the second speaker, the reconstructed audio data of the second speaker is reconstructed based on the target timbre information corresponding to the second speaker and the target prosody information corresponding to the second speaker, the target timbre information corresponding to the second speaker is obtained based on the timbre encoder extracting the sample audio data of the second speaker, and the target prosody information corresponding to the second speaker is obtained based on the prosody encoder extracting the sample audio data of the second speaker; the first speaker and the second speaker are different speakers.

[0197] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the speech decoupling method provided by the above-mentioned methods, the method comprising: inputting the speech data to be decoupled into a timbre encoder and a prosody encoder respectively, obtaining decoupled timbre information output by the timbre encoder, and decoupled prosody information output by the prosody encoder; wherein the timbre encoder and the prosody encoder are trained based on a first reconstruction loss and a second reconstruction loss; the first reconstruction loss is determined based on the sample audio data of a first speaker and the reconstructed audio data of the first speaker, the reconstructed audio data of the first speaker is reconstructed based on the target timbre information corresponding to the first speaker and the target prosody information corresponding to the first speaker, and the target timbre information corresponding to the first speaker is based on the The timbre encoder extracts the sample audio data of the first speaker, and the target prosody information corresponding to the first speaker is obtained based on the sample audio data extracted by the prosody encoder from the first speaker; the second reconstruction loss is determined based on the sample audio data of the second speaker and the reconstructed audio data of the second speaker, the reconstructed audio data of the second speaker is reconstructed based on the target timbre information corresponding to the second speaker and the target prosody information corresponding to the second speaker, the target timbre information corresponding to the second speaker is obtained based on the sample audio data extracted by the timbre encoder from the second speaker, and the target prosody information corresponding to the second speaker is obtained based on the sample audio data extracted by the prosody encoder from the second speaker; the first speaker and the second speaker are different speakers.

[0198] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0199] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0200] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A speech decoupling method, characterized in that: include: Inputting the to-be-decoupled speech data into a timbre encoder and a prosody encoder respectively, obtaining decoupled timbre information output by the timbre encoder and decoupled prosody information output by the prosody encoder; The timbre encoder and the prosody encoder are trained based on a first reconstruction loss and a second reconstruction loss; The first reconstruction loss is determined based on sample audio data of a first speaker and reconstructed audio data of the first speaker, the reconstructed audio data of the first speaker is reconstructed based on target timbre information corresponding to the first speaker and target prosody information corresponding to the first speaker, the target timbre information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker by the timbre encoder, and the target prosody information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker by the prosody encoder; the second reconstruction loss is determined based on sample audio data of a second speaker and reconstructed audio data of the second speaker, the reconstructed audio data of the second speaker is reconstructed based on target timbre information corresponding to the second speaker and target prosody information corresponding to the second speaker, the target timbre information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker by the timbre encoder, and the target prosody information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker by the prosody encoder; The first speaker and the second speaker are different speakers.

2. The speech decoupling method according to claim 1, characterized in that: The timbre encoder and the prosody encoder are trained based on the following method: Inputting the sample audio data of the first speaker and the sample audio data of the second speaker into a timbre encoder in the acoustic model respectively to obtain target timbre information output by the timbre encoder respectively, and inputting the sample audio data of the first speaker and the sample audio data of the second speaker into a prosody encoder in the acoustic model respectively to obtain target prosody information output by the prosody encoder respectively; Inputting the target timbre information corresponding to the first speaker and the target prosody information corresponding to the second speaker into an audio decoder in the acoustic model to obtain first reconstructed audio data output by the audio decoder; the audio decoder is configured to reconstruct audio data based on the input information; Inputting the first reconstructed audio data into the timbre encoder to obtain reconstructed timbre information output by the timbre encoder, and inputting the first reconstructed audio data into the prosody encoder to obtain reconstructed prosody information output by the prosody encoder; Inputting the reconstructed timbre information and the target prosody information corresponding to the first speaker into the audio decoder to obtain second reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the second speaker into the audio decoder to obtain third reconstructed audio data output by the audio decoder; determining a first reconstruction loss based on the sample audio data of the first speaker and the second reconstructed audio data, and determining a second reconstruction loss based on the sample audio data of the second speaker and the third reconstructed audio data; The acoustic model is trained based on the first reconstruction loss and the second reconstruction loss.

3. The speech decoupling method according to claim 2, characterized in that: The sample audio data of the first speaker includes different first audio data and second audio data, and the sample audio data of the second speaker includes different third audio data and fourth audio data; The method comprises: inputting the sample audio data of the first speaker and the sample audio data of the second speaker into a timbre encoder in the acoustic model to obtain target timbre information output by the timbre encoder, and inputting the sample audio data of the first speaker and the sample audio data of the second speaker into a prosody encoder in the acoustic model to obtain target prosody information output by the prosody encoder, including: inputting the second audio data and the fourth audio data into the timbre encoder respectively to obtain target timbre information outputted by the timbre encoder respectively, and inputting the first audio data and the third audio data into the prosody encoder respectively to obtain target prosody information outputted by the prosody encoder respectively; The step of inputting the target timbre information corresponding to the first speaker and the target prosody information corresponding to the second speaker into an audio decoder in the acoustic model to obtain first reconstructed audio data output by the audio decoder includes: Inputting the target timbre information corresponding to the second audio data of the first speaker and the target prosody information corresponding to the third audio data of the second speaker into the audio decoder to obtain first reconstructed audio data output by the audio decoder; The step of inputting the reconstructed timbre information and the target prosody information corresponding to the first speaker into the audio decoder to obtain second reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the second speaker into the audio decoder to obtain third reconstructed audio data output by the audio decoder includes: inputting the reconstructed timbre information and the target prosody information corresponding to the first audio data of the first speaker into the audio decoder to obtain second reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the fourth audio data of the second speaker into the audio decoder to obtain third reconstructed audio data output by the audio decoder; The determining of a first reconstruction loss based on the sample audio data of the first speaker and the second reconstructed audio data, and determining a second reconstruction loss based on the sample audio data of the second speaker and the third reconstructed audio data, comprises: A first reconstruction loss is determined based on the first audio data of the first speaker and the second reconstructed audio data, and a second reconstruction loss is determined based on the third audio data of the second speaker and the third reconstructed audio data.

4. The speech decoupling method according to claim 3, characterized in that: The step of inputting the second audio data and the fourth audio data into the timbre encoder to obtain target timbre information output by the timbre encoder, and inputting the first audio data and the third audio data into the prosody encoder to obtain target prosody information output by the prosody encoder, respectively, includes: inputting the second audio data and the fourth audio data into the timbre encoder, respectively, to obtain target timbre information outputted by the timbre encoder, inputting the first audio data and the third audio data into the prosody encoder, respectively, to obtain target prosody information outputted by the prosody encoder, and inputting first text data corresponding to the first audio data and third text data corresponding to the third audio data into a content encoder in the acoustic model, respectively, to obtain target content information outputted by the content encoder; The step of inputting the target timbre information corresponding to the second audio data of the first speaker and the target prosody information corresponding to the third audio data of the second speaker into the audio decoder to obtain first reconstructed audio data output by the audio decoder includes: Inputting the target timbre information corresponding to the second audio data of the first speaker, the target prosody information corresponding to the third audio data of the second speaker, and the target content information corresponding to the third audio data of the second speaker into the audio decoder, thereby obtaining first reconstructed audio data output by the audio decoder; The step of inputting the reconstructed timbre information and the target prosody information corresponding to the first audio data of the first speaker into the audio decoder to obtain second reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the fourth audio data of the second speaker into the audio decoder to obtain third reconstructed audio data output by the audio decoder includes: The reconstructed timbre information and the target prosody information corresponding to the first audio data of the first speaker, as well as the target content information corresponding to the first audio data of the first speaker are input into the audio decoder to obtain second reconstructed audio data output by the audio decoder, and the target timbre information and the reconstructed prosody information corresponding to the fourth audio data of the second speaker, as well as the target content information corresponding to the third audio data of the second speaker are input into the audio decoder to obtain third reconstructed audio data output by the audio decoder.

5. The speech decoupling method according to claim 2, characterized in that: The step of training the acoustic model based on the first reconstruction loss and the second reconstruction loss includes: Inputting the reconstructed timbre information and the target prosody information corresponding to the second speaker into the audio decoder to obtain fourth reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the first speaker into the audio decoder to obtain fifth reconstructed audio data output by the audio decoder; determining a third reconstruction loss based on the first reconstructed audio data and the fourth reconstructed audio data, and determining a fourth reconstruction loss based on the first reconstructed audio data and the fifth reconstructed audio data; The acoustic model is trained based on the first reconstruction loss, the second reconstruction loss, the third reconstruction loss, and the fourth reconstruction loss.

6. The speech decoupling method according to claim 5, characterized in that: The sample audio data of the first speaker includes different first audio data and second audio data, and the sample audio data of the second speaker includes different third audio data and fourth audio data; The method comprises: inputting the sample audio data of the first speaker and the sample audio data of the second speaker into a timbre encoder in the acoustic model to obtain target timbre information output by the timbre encoder, and inputting the sample audio data of the first speaker and the sample audio data of the second speaker into a prosody encoder in the acoustic model to obtain target prosody information output by the prosody encoder, including: inputting the second audio data and the fourth audio data into the timbre encoder, respectively, to obtain target timbre information outputted by the timbre encoder, inputting the first audio data and the third audio data into the prosody encoder, respectively, to obtain target prosody information outputted by the prosody encoder, and inputting first text data corresponding to the first audio data and third text data corresponding to the third audio data into a content encoder in the acoustic model, respectively, to obtain target content information outputted by the content encoder; The step of inputting the target timbre information corresponding to the first speaker and the target prosody information corresponding to the second speaker into an audio decoder in the acoustic model to obtain first reconstructed audio data output by the audio decoder includes: Inputting the target timbre information corresponding to the second audio data of the first speaker, the target prosody information corresponding to the third audio data of the second speaker, and the target content information corresponding to the third audio data of the second speaker into the audio decoder, thereby obtaining first reconstructed audio data output by the audio decoder; The step of inputting the reconstructed timbre information and the target prosody information corresponding to the second speaker into the audio decoder to obtain fourth reconstructed audio data output by the audio decoder, and inputting the target timbre information and the reconstructed prosody information corresponding to the first speaker into the audio decoder to obtain fifth reconstructed audio data output by the audio decoder includes: The reconstructed timbre information and the target prosody information corresponding to the third audio data of the second speaker, as well as the target content information corresponding to the third audio data of the second speaker are input into the audio decoder to obtain fourth reconstructed audio data output by the audio decoder, and the target timbre information and the reconstructed prosody information corresponding to the second audio data of the first speaker, as well as the target content information corresponding to the third audio data of the second speaker are input into the audio decoder to obtain fifth reconstructed audio data output by the audio decoder.

7. The speech decoupling method according to claim 5, characterized in that: The step of training the acoustic model based on the first reconstruction loss, the second reconstruction loss, the third reconstruction loss, and the fourth reconstruction loss includes: Inputting the first reconstructed audio data, the second reconstructed audio data, the third reconstructed audio data, and the fourth reconstructed audio data into a discriminator respectively, and obtaining four discrimination losses output by the discriminator; The acoustic model is trained based on the first reconstruction loss, the second reconstruction loss, the third reconstruction loss, the fourth reconstruction loss, and the four discriminant losses.

8. The speech decoupling method according to claim 2, characterized in that: The step of inputting the sample audio data of the first speaker and the sample audio data of the second speaker into the prosody encoder in the acoustic model respectively, and obtaining the target prosody information outputted by the prosody encoder respectively, comprises: Performing timbre change processing on the sample audio data of the first speaker and the sample audio data of the second speaker, respectively, to obtain the sample audio data of the first speaker after timbre change and the sample audio data of the second speaker after timbre change; The sample audio data of the first speaker after the timbre change and the sample audio data of the second speaker after the timbre change are respectively inputted into a prosody encoder in the acoustic model to obtain target prosody information respectively outputted by the prosody encoder.

9. A speech decoupling device, characterized in that: include: A speech decoupling module is used to input the speech data to be decoupled into the timbre encoder and the prosody encoder respectively, to obtain the decoupled timbre information output by the timbre encoder and the decoupled prosody information output by the prosody encoder; The timbre encoder and the prosody encoder are trained based on a first reconstruction loss and a second reconstruction loss; The first reconstruction loss is determined based on sample audio data of a first speaker and reconstructed audio data of the first speaker, the reconstructed audio data of the first speaker is reconstructed based on target timbre information corresponding to the first speaker and target prosody information corresponding to the first speaker, the target timbre information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker by the timbre encoder, and the target prosody information corresponding to the first speaker is obtained by extracting the sample audio data of the first speaker by the prosody encoder; the second reconstruction loss is determined based on sample audio data of a second speaker and reconstructed audio data of the second speaker, the reconstructed audio data of the second speaker is reconstructed based on target timbre information corresponding to the second speaker and target prosody information corresponding to the second speaker, the target timbre information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker by the timbre encoder, and the target prosody information corresponding to the second speaker is obtained by extracting the sample audio data of the second speaker by the prosody encoder; The first speaker and the second speaker are different speakers.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the speech decoupling method according to any one of claims 1 to 8 is implemented.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the speech decoupling method according to any one of claims 1 to 8 is implemented.

12. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the speech decoupling method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Speech synthesis method and device, electronic equipment and storage medium

    CN112786012A

  • Voice conversion model training method, voice conversion method, device and medium

    CN115171666A