Audio synthesis method and model training method, apparatus, device, medium and product

By converting reference audio into a specified timbre and extracting text semantics and expressive features, the problem of insufficient timbre transfer capability in traditional speech synthesis technology is solved, achieving high-quality, highly natural speech synthesis and enhancing the expressiveness of emotion and rhythm.

CN122392481APending Publication Date: 2026-07-14TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2026-06-12
Publication Date
2026-07-14

Smart Images

  • Figure CN122392481A_ABST
    Figure CN122392481A_ABST
Patent Text Reader

Abstract

An audio synthesis method, model training method, apparatus, device, medium, and product are disclosed. The method includes: converting reference audio into a first audio with a specified timbre, wherein the timbre of the reference audio differs from the specified timbre, and the specified timbre is selected from a variety of candidate timbres; obtaining first text, extracting semantic features from the first text, and converting the first text into a first phoneme, wherein the first text describes the content of the audio to be synthesized; extracting first timbre features and first expression features from the first audio, wherein the first expression features include at least one of a first prosodic feature or a first emotional feature; concatenating the first text semantic features, the first phoneme, and the first expression features to predict the audio content, thereby obtaining a first audio content representation; and fusing the first audio content representation, the first phoneme, and the first timbre features to obtain a target audio with the specified timbre. This method can improve the controllability of timbre and the accuracy of audio synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio synthesis technology, and in particular to an audio synthesis method, model training method, apparatus, computer device, storage medium, and computer program product. Background Technology

[0002] Speech synthesis technology is increasingly widely used in modern information technology. Traditional speech synthesis technology mainly focuses on accurate pronunciation and natural intonation, but it still has significant shortcomings in expressing emotions and styles in speech, as well as in timbre transfer. In order to meet users' needs for emotionally rich and diverse speech, speech emotion control and style and timbre transfer have become urgent problems to be solved. Summary of the Invention

[0003] Therefore, it is necessary to provide an audio synthesis method, model training method, apparatus, computer equipment, storage medium, and computer program product to address the aforementioned technical problems, thereby improving the accuracy and reliability of audio synthesis.

[0004] Firstly, this application provides an audio synthesis method. The method includes:

[0005] Obtain a reference audio, convert the reference audio into a first audio with a specified timbre, wherein the timbre of the reference audio is different from the specified timbre, and the specified timbre is a timbre selected from a variety of candidate timbres;

[0006] Obtain the first text, extract the semantic features of the first text, and convert the first text into the first phoneme. The first text is used to describe the content in the audio to be synthesized.

[0007] Extract a first timbre feature and a first expression feature from the first audio, wherein the first expression feature includes at least one of a first prosodic feature or a first emotional feature;

[0008] The first text semantic features, the first phoneme and the first expression features are concatenated and then audio content prediction is performed to obtain the first audio content representation;

[0009] The first audio content representation, the first phoneme, and the first timbre feature are fused together to obtain the target audio corresponding to the specified timbre.

[0010] Secondly, this application also provides an audio synthesis apparatus. The apparatus includes:

[0011] The first conversion module is used to acquire reference audio and convert the reference audio into a first audio with a corresponding specified timbre. The timbre of the reference audio is different from the specified timbre, and the specified timbre is a timbre selected from a variety of candidate timbres.

[0012] The first feature extraction module is used to acquire first text, extract semantic features from the first text, and convert the first text into first phonemes. The first text is used to describe the content in the audio to be synthesized. The module also extracts first timbre features and first expression features from the first audio. The first expression features include at least one of first prosodic features or first emotional features.

[0013] The first prediction module is used to concatenate the first text semantic features, the first phoneme and the first expression features to predict the audio content and obtain the first audio content representation.

[0014] The first synthesis module is used to fuse the first audio content representation, the first phoneme, and the first timbre feature to obtain the target audio corresponding to the specified timbre.

[0015] Thirdly, this application provides a model training method. The method includes:

[0016] A sample audio is acquired and converted into a second audio with a specified timbre through a conversion network in an initial audio synthesis model. The timbre of the sample audio is different from the specified timbre, which is selected from a variety of candidate timbres.

[0017] The second text is obtained, and its semantic features are extracted through the feature extraction network in the initial audio synthesis model. The second text is then converted into a second phoneme. The second timbre features and second expression features are extracted from the second audio. The second text is used to describe the content in the audio to be synthesized. The second expression features include at least one of the second prosodic features or the second emotional features.

[0018] The second audio content representation is obtained by concatenating the second text semantic features, the second phoneme and the second expression features through the semantic prediction network in the initial audio synthesis model and then predicting the audio content.

[0019] By fusing the second audio content representation, the second phoneme, and the second timbre feature through the synthesis network in the initial audio synthesis model, a third audio corresponding to the specified timbre is obtained;

[0020] The audio conversion loss is determined based on the second audio and the first audio tag, and the audio synthesis loss is determined based on the third audio and the second audio tag;

[0021] The initial audio synthesis model is trained based on the audio conversion loss and the audio synthesis loss to obtain the audio synthesis model.

[0022] Fourthly, this application also provides a model training apparatus. The apparatus includes:

[0023] The second conversion module is used to acquire sample audio and convert it into a second audio with a specified timbre through the conversion network in the initial audio synthesis model. The timbre of the sample audio is different from the specified timbre, which is selected from a variety of candidate timbres.

[0024] The second feature extraction module is used to acquire the second text, extract the semantic features of the second text through the feature extraction network in the initial audio synthesis model, convert the second text into a second phoneme, extract the second timbre features and the second expression features from the second audio, the second text is used to describe the content in the audio to be synthesized, and the second expression features include at least one of the second prosodic features or the second emotional features.

[0025] The second prediction module is used to concatenate the second text semantic features, the second phoneme, and the second expression features through the semantic prediction network in the initial audio synthesis model to predict the audio content and obtain the second audio content representation.

[0026] The second synthesis module is used to fuse the second audio content representation, the second phoneme, and the second timbre feature through the synthesis network in the initial audio synthesis model to obtain a third audio corresponding to the specified timbre;

[0027] The loss determination module is used to determine audio conversion loss based on the second audio and the first audio tag, and to determine audio synthesis loss based on the third audio and the second audio tag;

[0028] The training module is used to train the initial audio synthesis model based on the audio conversion loss and the audio synthesis loss to obtain the audio synthesis model.

[0029] Fifthly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of any of the methods described above.

[0030] Sixthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of any of the above methods.

[0031] Seventhly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of any of the methods described above.

[0032] The aforementioned audio synthesis methods, model training methods, devices, equipment, media, and products convert reference audio into a first audio with a specified timbre, selected from multiple candidate timbres. This allows for the conversion of the reference audio's timbre into the user-selected timbre, achieving timbre controllability. After conversion, first timbre features and first expressive features are extracted. These timbre features are derived from the specified timbre, while the expressive features are at least one of the prosodic or emotional features represented by the reference audio. This allows for controllability of the expressed prosodic and emotional content during audio synthesis, enhancing the expressiveness of prosodic and emotional content. First semantic features and first phonemes are extracted from the first text. Focusing on pronunciation within the text through phonemes improves the accuracy of subsequent content prediction and timbre fusion. A two-stage generation architecture is employed. In the first stage, the first text semantic features, first phonemes, and first expressive features (prosodic and / or emotional features) are concatenated to predict and generate a first audio content representation. The generated audio content representation is a feature representation that integrates text semantics, text pronunciation (provided by phonemes), and expressed prosodic and emotional content. In the second stage, the first audio content representation generated in the first stage is fused with the first phoneme and first timbre features to obtain the target timbre corresponding to the specified timbre. This accurately transfers the user-selected timbre to the content and pronunciation of the first text, while effectively preserving the rhythmic and emotional expression styles of the reference audio. This makes the emotional and / or rhythmic expression richer and more natural, improving the expressiveness of the audio. Phoneme fusion ensures that each pronunciation can be accurately integrated with the timbre, achieving high-quality, highly natural speech synthesis. Furthermore, by introducing joint modeling of the semantic features of the first text and the first phoneme, the pronunciation accuracy and semantic coherence of complex texts are improved, the dependence on large-scale labeled data is reduced, and the generalization ability and robustness of the audio synthesis model under different speakers and contexts are enhanced. Attached Figure Description

[0033] Figure 1 This is a diagram illustrating the application environment of an audio synthesis method in one embodiment;

[0034] Figure 2 This is a flowchart illustrating an audio synthesis method in one embodiment;

[0035] Figure 3 This is a schematic diagram of an audio synthesis method in one embodiment;

[0036] Figure 4 This is a schematic diagram of an audio synthesis method in another embodiment;

[0037] Figure 5 This is a schematic diagram of the audio synthesis method in one embodiment;

[0038] Figure 6 This is a logic diagram in one embodiment for converting a reference audio into a first audio corresponding to a specified timbre;

[0039] Figure 7 This is a schematic diagram of the audio synthesis method in another embodiment;

[0040] Figure 8 This is a schematic diagram of the audio synthesis method in yet another embodiment;

[0041] Figure 9 This is a schematic diagram of the architecture of the audio synthesis method in another embodiment;

[0042] Figure 10 This is a flowchart illustrating a model training method in one embodiment;

[0043] Figure 11 This is a schematic diagram illustrating the principle of the model training method in another embodiment;

[0044] Figure 12 This is a schematic diagram of the processing logic of a model training method in one embodiment;

[0045] Figure 13 This is an architecture diagram of the conversion network in the initial audio synthesis model of one embodiment;

[0046] Figure 14 This is a schematic diagram of the processing in the conversion network in one embodiment;

[0047] Figure 15 This is a structural block diagram of an audio synthesis device in one embodiment;

[0048] Figure 16 This is a structural block diagram of an audio synthesis device in one embodiment;

[0049] Figure 17 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0051] It should be noted that in the following description, the terms "first, second, third, fourth, fifth and sixth" are merely used to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third, fourth, fifth and sixth" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0052] Figure 1 This is a schematic diagram of an application environment in one embodiment, such as... Figure 1 As shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on other servers. Terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart TVs, smart in-vehicle devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. Server 104 can be a standalone physical server, a service node in a blockchain system, or a server cluster consisting of at least two physical servers. The server cluster can be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.

[0053] Both the terminal and the server can be used independently to execute the methods provided in the embodiments of this application. For example, the server obtains reference audio, converts the reference audio into a first audio corresponding to a specified timbre, the timbre of the reference audio being different from the specified timbre, the specified timbre being a timbre selected from a variety of candidate timbres. The server obtains first text, extracts semantic features from the first text, and converts the first text into a first phoneme, the first text being used to describe the content in the audio to be synthesized. First timbre features and first expression features are extracted from the first audio, the first expression features including at least one of a first prosodic feature or a first emotional feature. The first text semantic features, the first phoneme, and the first expression features are concatenated and then used to predict the audio content to obtain a first audio content representation. The server fuses the first audio content representation, the first phoneme, and the first timbre features to obtain the target audio corresponding to the specified timbre.

[0054] In one embodiment, such as Figure 2 As shown, an audio synthesis method is provided. This method can be executed by a server or terminal alone, or by both a server and a terminal, to be applied to... Figure 1 Taking a terminal as an example, this audio synthesis method may include the following steps:

[0055] Step S202: Obtain reference audio and convert the reference audio into the first audio corresponding to the specified timbre. The timbre of the reference audio is different from the specified timbre, which is selected from a variety of candidate timbres.

[0056] Among them, the reference audio is the audio used to provide expressive information, including the style and emotion of the provided audio, etc. The specified timbre is the timbre required in the audio to be synthesized. For example, there are 5 selectable timbres, and the user specifies to synthesize an audio with the 3rd timbre. The first audio is the audio obtained by migrating the specified timbre to the reference timbre.

[0057] The reference audio can be obtained, and the specified timbre is migrated to the reference audio to convert the reference audio into the first audio corresponding to the specified timbre.

[0058] Exemplarily, the first audio corresponding to the specified timbre can be an audio with timbre similarity to the specified timbre. For example, the similarity between the timbre of the first audio and the specified timbre reaches a preset similarity. The preset similarity can be, for example, 80%, 90%, 100%, etc.

[0059] Exemplarily, the first audio corresponding to the specified timbre can specifically be the first audio with the specified timbre. The reference audio can be obtained, and the specified timbre is migrated to the reference audio to obtain the first audio with the specified timbre.

[0060] Step S204: Obtain the first text, extract the semantic features of the first text, and convert the first text into the first phoneme. The first text is used to describe the content in the audio to be synthesized.

[0061] The first text is the text that determines the content of the synthesized audio and can provide the semantic content and pronunciation features of the audio to be synthesized. For example, if the first text is "The weather is really nice today and it's suitable for outings", then the content "The weather is really nice today and it's suitable for outings" is included in the finally generated target audio.

[0062] Exemplarily, a semantic feature sequence of the first text can be extracted. The semantic feature sequence of the first text is a sequence composed of at least two semantic features of the first text, and each semantic feature of the first text is used to express the semantic feature of a part of the content in the first text. For example, each semantic feature of the text is used to express the semantic feature of a segment in the first text, or is used to express the semantic feature of a sentence in the first text, or is used to express the semantic feature of a word and / or character in the first text.

[0063] Exemplarily, the first text can be converted into a first phoneme sequence. The first phoneme sequence is a sequence composed of at least two first phonemes. A phoneme is the smallest speech unit that constitutes pronunciation. For example, the Chinese character "啊" is pronounced as "ā" and has only one phoneme. The Chinese character "爱" is pronounced as "ài" and is composed of 2 phonemes.

[0064] Obtain the first text, perform feature extraction on the first text, and obtain the semantic features of the first text.

[0065] For example, the first text can be segmented to obtain at least two segments, each segment may include at least one sentence. Feature extraction is performed on each segment to obtain the first text semantic features of each segment.

[0066] For example, a first text is obtained, and a first text semantic feature sequence is extracted from the first text. The first text semantic feature sequence includes at least two first text semantic features. First text semantic features are extracted from each segment separately, and the first text semantic features are concatenated in sequence to obtain the first text semantic feature sequence.

[0067] For example, the first text can be divided into words, and features can be extracted from each word to obtain semantic features of the first text. The semantic features of the first text can be concatenated in sequence to obtain a sequence of semantic features of the first text.

[0068] For example, the first text can be divided into characters, and features can be extracted from each character to obtain the semantic features of the first text. The semantic features of the first text can be concatenated in sequence to obtain the semantic feature sequence of the first text.

[0069] It can determine the pronunciation of each character in the first text and determine the phoneme (i.e., the first phoneme) corresponding to the pronunciation of each character.

[0070] For example, the first text can be converted into a first phoneme sequence, which includes at least two first phonemes. The first phoneme corresponding to the pronunciation of each word in the first text is determined, and the first phonemes are concatenated sequentially to obtain the first phoneme sequence.

[0071] For example, the first text can be divided into characters, the phoneme sequence of each character can be found in the pronunciation dictionary, and the phoneme sequences of each character can be concatenated in order to obtain the first phoneme sequence.

[0072] Step S206: Extract a first timbre feature and a first expression feature from the first audio. The first expression feature includes at least one of a first prosodic feature or a first emotional feature.

[0073] The first timbre feature is a feature used to characterize the timbre of the first audio audio. The first timbre feature is a timbre feature extracted from a frame of the first audio audio.

[0074] The first expressive feature is an expressive feature extracted from a frame of the first audio. The first expressive feature is a feature that expresses the rhythm and / or emotion of a frame in the first audio. That is, the first expressive feature includes a first prosodic feature and / or a first emotional feature. The first prosodic feature is a feature in the first audio used to express rhythm. The first emotional feature is a feature in the first audio used to express emotion.

[0075] For example, a first timbre feature sequence and a first expression feature sequence can be extracted from the first audio. The first timbre feature sequence is a sequence formed by concatenating at least two first timbre features. The first expression feature sequence is a sequence formed by concatenating at least two first expression features. There is a one-to-one correspondence between at least two first timbre features in the first timbre feature sequence and at least two first expression features in the first expression feature sequence.

[0076] The first audio can be divided into at least two frames. For each frame, the first timbre feature and the first expression feature are extracted respectively.

[0077] For example, the first timbre features of each frame are sequentially spliced ​​into a first timbre feature sequence, and the first expression features of each frame are sequentially spliced ​​into a first expression feature sequence.

[0078] For example, when the first expression feature sequence includes a first prosodic feature sequence, for each segmented frame, the first timbre feature and the first prosodic feature are extracted for each frame, and the first prosodic features of each frame are concatenated in sequence to form the first prosodic feature sequence.

[0079] For example, when the first expression feature sequence includes the first emotion feature sequence, for each segmented frame, the first timbre feature and the first emotion feature are extracted for each frame respectively, and the first emotion features of each frame are concatenated into the first emotion feature sequence in sequence.

[0080] Step S208: After concatenating the first text semantic features, the first phoneme, and the first expression features, perform audio content prediction to obtain the first audio content representation.

[0081] The first audio content representation is an audio content representation obtained by fusing semantics and phonemes from the first text and incorporating expressive features from the first audio.

[0082] The first text semantic features, the first phoneme, and the first expression features can be concatenated, and the concatenated features can be used to predict audio content to obtain the first audio content representation.

[0083] For example, the first audio content representation is a feature that characterizes the content of a frame in the audio to be synthesized.

[0084] For example, the first audio content representation sequence is a representation sequence obtained by fusing semantics and phonemes from the first text and fusing expressive features from the first audio.

[0085] The first text semantic feature sequence, the first phoneme sequence, and the first expression feature sequence can be concatenated, and the concatenated feature sequence can be used to predict audio content to obtain the first audio content representation sequence.

[0086] For example, the first text semantic feature sequence, the first phoneme sequence, and the first expression feature sequence can be concatenated, and the concatenated feature sequence can be subjected to autoregressive prediction to obtain the first audio content representation sequence.

[0087] Step S210: The first audio content representation, the first phoneme, and the first timbre feature are fused to obtain the target audio with the specified timbre.

[0088] The target audio corresponding to the specified timbre can be audio with timbre similarity to the specified timbre. For example, the similarity between the timbre of the target audio and the specified timbre reaches a preset similarity. The preset similarity can be, for example, 80%, 90%, 100%, etc.

[0089] The first audio content representation, the first phoneme, and the first timbre feature are fused to obtain the target audio with the specified timbre. The principle diagram of this audio synthesis method is shown below. Figure 3 As shown.

[0090] For example, the first audio content representation, the first phoneme, and the first timbre feature can be fused to obtain a fused feature, which can then be converted into an audio waveform to obtain the target audio corresponding to the specified timbre. Further, the fused feature can be converted into a spectrogram feature, and then the spectrogram feature can be converted into an audio waveform. This audio waveform is the waveform of the target audio.

[0091] For example, the target audio corresponding to the specified timbre can specifically be target audio with the specified timbre. Audio synthesis can be performed based on the first audio content representation, the first phoneme, and the first timbre feature to obtain the target audio with the specified timbre.

[0092] For example, audio synthesis can be performed based on a first audio content representation sequence, a first phoneme sequence, and a first timbre feature sequence to obtain a target audio with a specified timbre.

[0093] The first audio content representation sequence, the first phoneme sequence, and the first timbre feature sequence are fused together. The fused feature sequence is then converted into a spectrogram, and the spectrogram is then waveform-converted to obtain the target audio with the specified timbre.

[0094] For example, such as Figure 4The diagram illustrates the principle of an audio synthesis method, comprising: acquiring reference audio and converting it into a first audio with a specified timbre; acquiring first text, extracting a semantic feature sequence from the first text, and converting the first text into a first phoneme sequence; extracting a first timbre feature sequence and a first expression feature sequence from the first audio; concatenating the first text semantic feature sequence, the first phoneme sequence, and the first expression feature sequence to predict audio content, obtaining a first audio content representation sequence; and performing audio synthesis based on the first audio content representation sequence, the first phoneme sequence, and the first timbre feature sequence to obtain a target audio with a specified timbre. The first expression feature includes at least one of a first prosodic feature or a first emotional feature.

[0095] In the above embodiments, by converting the reference audio into a first audio with a specified timbre selected from multiple candidate timbres, the timbre of the reference audio can be converted into the user-selected specified timbre, achieving controllability of the timbre. After conversion, a first timbre feature and a first expression feature are extracted. The timbre feature is actually derived from the specified timbre, and the expression feature is at least one of the prosodic or emotional features represented by the reference audio. This allows for controllability of the expressed prosodic and emotional features during audio synthesis, enhancing the expressiveness of prosodic and emotional features. A first semantic feature and a first phoneme are extracted from the first text. By focusing on the pronunciation in the text through phonemes, the accuracy of subsequent content prediction and timbre fusion is improved. A two-stage generation architecture is adopted. In the first stage, the semantic features of the first text, the first phoneme, and the first expression feature (prosodic and / or emotional features) are concatenated to predict and generate a first audio content representation. The generated audio content representation is a feature representation that integrates text semantics, text pronunciation (provided by phonemes), and expressed prosodic and emotional features. In the second stage, the first audio content representation generated in the first stage is fused with the first phoneme and the first timbre features to obtain the target timbre corresponding to the specified timbre. This can accurately transfer the timbre selected by the user to the content and pronunciation of the first text, while effectively preserving the rhythm, emotion and other expressive styles in the reference audio, making the emotional expression and / or rhythmic expression richer and more natural, improving the expressiveness of the audio, and ensuring that each pronunciation can be accurately fused with the timbre through phoneme fusion, thus achieving high-quality and highly natural speech synthesis.

[0096] Furthermore, by introducing joint modeling of the first text semantic features and the first phoneme, the pronunciation accuracy and semantic coherence of complex texts are improved, the dependence on large-scale labeled data is reduced, and the generalization ability and robustness of the audio synthesis model under different speakers and different contexts are enhanced.

[0097] In one embodiment, converting the reference audio into a first audio corresponding to a specified timbre includes:

[0098] Extract reference content features and reference prosodic features from each frame of the reference audio; obtain a specified timbre identifier and convert the specified timbre identifier into a timbre embedding representation; retrieve acoustic features that match the reference content features of each frame from the acoustic feature library corresponding to the specified timbre identifier; synthesize audio based on each acoustic feature, timbre embedding representation and each reference prosodic feature to obtain the first audio corresponding to the specified timbre.

[0099] The reference content features refer to the content information of each frame in the reference audio, such as the specific text or syllables in each frame. The reference prosodic features refer to the pitch, intonation, rhythm, and other information of each frame in the reference audio. The specified timbre identifier is a number for a specific timbre. The acoustic feature library is used to store acoustic features extracted from audio with specified timbres. The acoustic features stored in the acoustic feature library may include at least one of the following: specified timbre features, content features, second latent variables, or acoustic statistics.

[0100] The second latent variable represents the acoustic latent state of a frame in the latent space, and is a compressed representation that integrates the timbre, content, and prosodic features of the same frame. Acoustic statistics are the acoustic distribution statistics of the second latent variable in the latent space. Acoustic statistics can describe the distribution range of timbre, content, and prosodic features.

[0101] The reference audio can be divided into multiple frames, and reference content features and reference prosodic features can be extracted for each frame. A specified timbre identifier is obtained, converted into a timbre embedding representation, and an acoustic feature library corresponding to the specified timbre identifier is determined. Then, acoustic features that match the reference content features of each frame are retrieved from the acoustic feature library corresponding to the specified timbre identifier.

[0102] For example, the acoustic features include candidate content features and acoustic statistics. The candidate content features and acoustic statistics are stored together in the acoustic feature library. Candidate content features that match the reference content features of each frame can be retrieved from the acoustic feature library, and acoustic statistics associated with the candidate content features can be retrieved from the acoustic feature library to obtain the acoustic statistics that match the reference content features of each frame.

[0103] Audio is synthesized based on various acoustic features, timbre embedding representations, and various reference prosodic features to obtain the first audio corresponding to the specified timbre.

[0104] For example, the acoustic features of each frame are concatenated to obtain an acoustic feature sequence. The reference prosodic features of each frame are then concatenated to obtain a reference prosodic feature sequence. Audio synthesis is performed based on the acoustic feature sequence, the timbre embedding representation, and the reference prosodic feature sequence to obtain a first audio file corresponding to a specified timbre.

[0105] For example, the first audio corresponding to the specified timbre may refer to the first audio having the specified timbre.

[0106] In this embodiment, the reference audio is divided into multiple frames, and reference content features and reference prosodic features are extracted frame by frame. Combined with a timbre embedding representation derived from a specified timbre identifier, matching content features are retrieved from a pre-built acoustic feature library to obtain the corresponding specified timbre features. Finally, the three are fused to synthesize the first audio. This achieves decoupling and precise recombination of timbre and content, enabling the transfer of any target timbre to the original reference audio without altering personalized prosodic information such as speech rate, rhythm, and pauses. This efficiently generates a first audio that retains both the original speaking style and the specified timbre. Simultaneously, the frame-by-frame acoustic feature retrieval and matching significantly improves the precision and naturalness of timbre transfer, avoiding timbre leakage or prosodic distortion caused by overall conversion in traditional methods. Furthermore, the introduction of timbre embedding representation enhances the model's generalization ability to unseen timbres, allowing the system to flexibly adapt to various predefined or custom target timbres, reducing the amount of reference data and computational overhead required for timbre cloning.

[0107] In one embodiment, the acoustic features include acoustic statistics; audio synthesis is performed based on each acoustic feature, timbre embedding representation, and each reference prosodic feature to obtain a first audio corresponding to a specified timbre, including:

[0108] For each frame in the reference audio, based on the timbre embedding representation and the reference prosodic features of the frame, feature transformation is performed on the acoustic statistics of the frame to obtain the first latent variable of the frame; the first latent variables of each frame are concatenated, and the concatenated latent variables are decoded to obtain the first audio corresponding to the specified timbre.

[0109] Acoustic features include acoustic statistics, and the acoustic feature library stores candidate content features and acoustic statistics in association. For each frame in the reference audio, candidate content features that match the reference content features of the target frame can be retrieved from the acoustic feature library, and the acoustic statistics associated with the candidate content features can be retrieved from the acoustic feature library to obtain the acoustic statistics that match the reference content features of the target frame.

[0110] Based on the timbre embedding representation and the reference prosodic features of the target frame, feature transformation is performed on the acoustic statistics of the target frame to obtain the first latent variable of the target frame. After obtaining the first latent variable of each frame in the reference audio, the first latent variables of each frame can be concatenated, and the concatenated latent variables can be decoded to obtain the first audio corresponding to the specified timbre.

[0111] For example, the first latent variable of each frame is spliced ​​together, and the spliced ​​latent variable is converted into a time-domain waveform that can be played directly to obtain the first audio corresponding to the specified timbre.

[0112] For example, such as Figure 5The diagram shown illustrates the principle of an audio synthesis method.

[0113] The process involves acquiring reference audio, extracting features from each frame to obtain reference content features and reference prosodic features for each frame, obtaining a specified timbre identifier, and converting the timbre identifier into a timbre embedding representation. Then, in the acoustic feature library corresponding to the specified timbre identifier, acoustic features matching the reference content features of each frame are retrieved. For each frame of the reference audio, based on the timbre embedding representation and the reference prosodic features of the frame, feature transformation is performed on the acoustic statistics of the frame to obtain the first latent variable of the frame. Finally, the first latent variables of each frame are concatenated, and the concatenated latent variables are decoded to obtain the first audio corresponding to the specified timbre.

[0114] Obtain the first text, extract its semantic features, and convert it into a first phoneme. Extract the first timbre and first expression features from the first audio. Concatenate the semantic features, first phoneme, and first expression features to predict the audio content and obtain a first audio content representation. Based on the first audio content representation, first phoneme, and first timbre features, synthesize the audio to obtain the target audio with the specified timbre.

[0115] For example, the first audio corresponding to the specified timbre can be a first audio with the specified timbre.

[0116] For example, the first expressive feature includes at least one of a first prosodic feature or a first emotional feature.

[0117] In this embodiment, acoustic statistics are used as basic features. For each frame of the reference audio, feature transformation is performed on the acoustic statistics using timbre embedding representation and the reference prosodic features of that frame to obtain the first latent variable. This is then concatenated and decoded to generate the first audio. This significantly improves the flexibility and controllability of timbre transfer and prosodic preservation. Because acoustic statistics are used instead of deterministic features, the model can more naturally adapt to the interactive changes of different timbres and prosodices, avoiding overfitting or a mechanical feel that may result from deterministic mapping. Based on frame-level feature transformation and latent variable generation, the acoustic details of each frame (such as pitch, energy, and duration) can be accurately integrated with the target timbre attributes while preserving the prosodic pattern of the original reference audio. This results in the synthesis of a first audio that possesses the specified timbre while highly reproducing the original rhythm, emotional fluctuations, and intonation changes. Furthermore, using statistics for transformation reduces the model's dependence on large-scale paired data, enhances the system's robustness to noisy or non-standard pronunciation reference audio, and allows for continuous control of timbre and prosodices by adjusting the statistical distribution, providing richer operational space for subsequent personalized speech synthesis.

[0118] In one embodiment, the method further includes:

[0119] Acquire candidate audio with a specified timbre, and extract candidate content features for each frame of the candidate audio; convert the waveform of the candidate audio into a latent variable sequence, which includes the second latent variable of each frame of the candidate audio; align the candidate content features of each frame with the second latent variable of each frame, and filter out the candidate content features of each frame that meet the similarity conditions; perform acoustic distribution statistics based on the second latent variable aligned with the selected candidate content features of each frame to obtain the acoustic statistics that are common to the selected candidate content features of each frame; associate and store the candidate content features of each frame with the corresponding acoustic statistics in the acoustic feature library.

[0120] It can acquire candidate audio with a specified timbre, divide the candidate audio into multiple frames, and extract candidate content features from each frame. The waveform of the candidate audio is also divided into multiple frames, with each frame's waveform corresponding one-to-one with its candidate content features.

[0121] The waveforms of the candidate audio are converted into a sequence of latent variables, which includes a second latent variable for each frame of the candidate audio. The second latent variable represents the acoustic latent state of a frame in the latent space, and is a fused representation of the timbre, content, and prosodic features of the same frame.

[0122] For example, the second latent variable can also be a compressed representation of the fused representation of timbre features, content features, and prosodic features of the same frame. A compressed representation refers to a feature representation obtained by reducing the dimensionality of the fused representation. For instance, the fused representation could be a 512-dimensional feature representation; after dimensionality reduction, a 32-dimensional or 64-dimensional feature representation is obtained. This 32-dimensional or 64-dimensional feature representation is the second latent variable.

[0123] For example, converting the waveform of the candidate audio into a sequence of latent variables includes:

[0124] The waveform of the candidate audio is divided into multiple frames. A Fourier transform is performed on the waveform of each frame to obtain the spectral representation vector of each frame. The spectral representation vector of each frame is then encoded to obtain the latent variables of each frame.

[0125] Furthermore, the spectral representation vector of each frame is encoded. Specifically, the spectral representation vector can be encoded frame by frame along the time window by an encoder to obtain the latent variables of each frame.

[0126] The latent variables for each frame are obtained by encoding the spectral representation vector frame by frame along a time window using an encoder. This includes multiplying the window function with the spectral representation vector of each frame to obtain the latent variables. The window function can be a Hamming window, a Hanning window, etc.

[0127] For example, the candidate content features of each frame are aligned with the second latent variable of each frame. The similarity between the candidate content features of each frame is calculated, and the candidate content features of each frame that meet the similarity condition are selected.

[0128] For example, the similarity condition may specifically be that a similarity threshold needs to be reached, and the similarity between candidate content features of each frame can be calculated to filter out candidate content features of each frame whose similarity meets the similarity threshold.

[0129] Acoustic distribution statistics are performed based on the second latent variable aligned with the selected candidate content features of each frame to obtain the acoustic statistics that are common to the selected candidate content features of each frame.

[0130] For example, the mean of the second latent variable aligned with the selected candidate content features of each frame is taken, and this mean is the acoustic statistic that is commonly associated with the selected candidate content features of each frame.

[0131] For example, the variance of the second latent variable aligned with the selected candidate content features of each frame is calculated, and this variance is the acoustic statistic that is commonly associated with the selected candidate content features of each frame.

[0132] For example, if 100 sentences are spoken in the candidate audio of a specified timbre, and there are a total of 30 frames in which the word "i" is pronounced, then the mean or variance of the second latent variable of these 30 frames is calculated. The variance or mean is the acoustic statistic corresponding to each of these 30 frames.

[0133] Each frame's candidate content features and corresponding acoustic statistics can be associated and stored in an acoustic feature library, allowing the corresponding acoustic statistics to be quickly found through the candidate content features.

[0134] For example, candidate content features and corresponding acoustic statistics of the same frame can be stored in an acoustic feature library in a key-value pair manner. For a frame, the candidate content features of the frame can be used as the key, and the acoustic statistics of the frame can be used as the value corresponding to the key.

[0135] In this embodiment, candidate audio with a specified timbre is acquired, and candidate content features are extracted frame by frame. Simultaneously, the waveform is converted into a latent variable sequence, and the content features are aligned frame by frame with the latent variables. Frame pairs that meet similarity conditions are filtered out, and acoustic distributions are statistically analyzed based on the aligned latent variables to obtain acoustic statistics corresponding to specific content features. Finally, the content features and acoustic statistics are associated and stored in an acoustic feature library, constructing an efficient, robust, and generalizable timbre-content associated acoustic feature library. Through latent variable alignment and similarity filtering, abnormal frames caused by noise, pronunciation deviations, or prosodic changes in the candidate audio are effectively removed, ensuring that the stored acoustic statistics can stably represent the essential acoustic attributes of each content feature under the specified timbre. Furthermore, using acoustic statistics instead of deterministic features for storage compresses data dimensionality, reduces storage overhead, and preserves the natural variation range of timbre in content expression, enabling more realistic and diverse synthesized speech to be flexibly generated based on context during subsequent retrieval. Furthermore, this construction process does not require manual annotation or alignment supervision, and can automatically extract high-quality statistical features from a small number of candidate audios, significantly improving the construction efficiency and scalability of the acoustic feature library, and providing a reliable data foundation for retrieval-based timbre transfer and speech synthesis in the embodiments.

[0136] In one embodiment, such as Figure 6 The diagram illustrates a principle for converting reference audio into a first audio with a specified timbre. The process involves acquiring the reference audio and extracting a first feature sequence from it. This first feature sequence includes a reference content feature sequence and a reference prosodic feature sequence. The reference audio sequence includes reference content features for each frame of the reference audio, and the reference prosodic feature sequence includes reference prosodic features for each frame of the reference audio. Reference content features include information such as the specific words or syllables spoken, while reference prosodic features include prosodic information such as intonation, pitch, rhythm, and stress. The reference content features and pitch sequence are timbre-independent features.

[0137] Obtain the specified timbre identifier (Speaker ID, SID) and convert it into a high-dimensional vector (Embedding), i.e., the timbre embedding representation. This timbre embedding representation represents the coordinates of the specified timbre in the high-dimensional feature space. This timbre embedding representation is broadcast to every layer of the reconstructor, so that the decoder knows which specified timbre to use when processing each frame.

[0138] The process involves pre-establishing an acoustic feature library: acquiring candidate audio files with specified timbres; extracting candidate content features from each frame of the candidate audio files; converting the waveforms of the candidate audio files into a latent variable sequence, which includes a second latent variable for each frame of the candidate audio files; aligning the candidate content features of each frame with the second latent variable of each frame; filtering out candidate content features that meet similarity criteria; performing acoustic distribution statistics based on the second latent variable aligned with the filtered candidate content features of each frame; and obtaining the acoustic statistics commonly corresponding to the filtered candidate content features of each frame. Finally, storing the candidate content features of each frame and their corresponding acoustic statistics in the acoustic feature library.

[0139] In the acoustic feature library corresponding to the specified timbre identifier, feature retrieval is performed using the reference content feature to retrieve candidate content features that match the reference content feature in the acoustic feature library, and the acoustic statistics associated with the candidate content feature are retrieved to obtain the acoustic statistics that match the reference content feature.

[0140] Feature transformation (i.e., variational flow network): The acoustic statistics, the timbre embedding representation of the specified timbre identifier, and the pitch sequence are input into the variational flow (an inverse transformation network based on normalizing flow). This inverse transformation network enhances the expressive power of the acoustic statistics, so that the acoustic statistics incorporate the prosody in the specified timbre and the reference audio, forming the first latent variable.

[0141] Audio decoding and generation: The first latent variables are fed into the decoder, which converts these first latent variables into audio waveforms (Raw Waveform) that can be played directly, completing end-to-end audio synthesis and obtaining the first audio with the specified timbre.

[0142] Specifically, the decoder can be a generator containing a multi-receiver field fusion module. It fuses and transforms these first latent variables to obtain an audio waveform, that is, to obtain the first audio with the specified timbre. The first audio has the specified timbre and also incorporates the content features and melodic features of the reference audio, that is, the specified timbre is transferred to the reference audio.

[0143] For example, in the generator that includes a Multi-Receptive Field Fusion (MRF) module, the generator contains at least two parallel residual blocks, each of which uses a different convolution kernel size and dilation rate to form features at different scales. Finally, the features of these residual blocks are summed to effectively enhance the rich details of the synthesized audio.

[0144] In one embodiment, the first expressive feature includes at least one of a first prosodic feature or a first emotional feature.

[0145] For example, a first timbre feature is extracted from the first audio, and a first prosodic feature and / or a first sentiment feature is extracted from the first audio. The first text semantic feature and the first phoneme, as well as the first prosodic feature and / or the first sentiment feature, are concatenated, and the audio content is predicted from the concatenated result to obtain a representation of the first audio content.

[0146] For example, the first text semantic features, the first phoneme and the first prosodic features are concatenated and then used to predict the audio content to obtain the first audio content representation, which integrates the text semantics and phonemes of the first text and the prosodic features of the first audio.

[0147] For example, the first text semantic features, the first phoneme and the first sentiment features are concatenated and then used to predict the audio content to obtain the first audio content representation, which integrates the text semantics and phonemes of the first text and the sentiment of the first audio.

[0148] For example, the first text semantic features, the first phoneme, the first prosodic features and the first sentiment features are concatenated and then used to predict the audio content to obtain the first audio content representation, which integrates the text semantics and phonemes of the first text, as well as the prosodic features and sentiment of the first audio.

[0149] For example, the first expression feature sequence includes at least one of a first prosodic feature sequence or a first sentiment feature sequence. The first prosodic feature sequence is a prosodic feature sequence obtained by concatenating at least two first prosodic features. The first sentiment feature sequence is a prosodic feature sequence obtained by concatenating at least two first sentiment features.

[0150] For example, a first timbre feature sequence is extracted from the first audio, and a first prosodic feature sequence and / or a first emotional feature sequence are extracted from the first audio. The first text semantic feature sequence, the first phoneme sequence, and the first prosodic feature sequence and / or the first emotional feature sequence are concatenated, and audio content prediction is performed on the concatenated sequence to obtain a first audio content representation sequence.

[0151] For example, the first text semantic feature sequence, the first phoneme sequence, and the first prosodic feature sequence are concatenated and then used to predict the audio content to obtain the first audio content representation sequence, which integrates the text semantics and phonemes of the first text, as well as the prosodicity of the first audio.

[0152] For example, the first text semantic feature sequence, the first phoneme sequence, and the first sentiment feature sequence are concatenated and then used to predict the audio content to obtain the first audio content representation sequence, which integrates the text semantics and phonemes of the first text, as well as the sentiment of the first audio.

[0153] For example, the first text semantic feature sequence, the first phoneme sequence, the first prosodic feature sequence and the first sentiment feature sequence are concatenated and then used to predict the audio content to obtain the first audio content representation sequence, which integrates the text semantics and phonemes of the first text, as well as the prosodic and sentiment of the first audio.

[0154] In this embodiment, a first prosodic feature and / or a first emotional feature are added when extracting features from the first audio. The first prosodic feature and / or the first emotional feature are then concatenated with the first text semantic feature and the first phoneme sequence and subjected to autoregressive prediction to obtain the first audio content representation sequence. This allows the target audio to not only reproduce the specified timbre and the prosodic rhythm of the reference audio, but also accurately transfer the emotional state from the reference audio. This achieves joint modeling and controllable synthesis of the three attributes of timbre, prosodicity, and emotionality, significantly improving the emotional expressiveness and subtlety of the synthesized speech. Because the emotional feature participates in every time step of the autoregressive prediction, it enables the generation of intonation fluctuations, energy changes, and phonological features consistent with the emotional state, avoiding the problem of stiff emotional expression caused by the loose coupling of emotion and prosodicity in traditional methods.

[0155] Furthermore, this method enables the output of target audio with different emotional tones by replacing emotional features under the same text and timbre, providing richer, more natural, and more engaging speech synthesis capabilities for various application scenarios such as personalized voice broadcasting, virtual assistants, and audio content creation.

[0156] In one embodiment, the reference audio is an emotion reference audio, and the method further includes:

[0157] The process involves acquiring sentiment description information and candidate text, performing feature encoding on the candidate text to obtain a text embedding representation sequence, generating a supervised semantic token sequence based on the sentiment description information and the text embedding representation sequence, acquiring preset noise, and extracting a spectrogram from the preset noise based on the supervised semantic token sequence, performing waveform transformation on the extracted spectrogram to obtain a sentiment reference audio, which carries the sentiment expressed by the sentiment description information.

[0158] The emotion description information describes the emotion that needs to be expressed. This emotion description information is used to describe the emotion to be expressed in the generated emotion reference audio. This emotion description information can be an emotion prompt instruction; for example, an emotion prompt instruction could be: "Please generate an audio with a sad tone."

[0159] A text embedding representation sequence comprises at least two text embedding representations. A text embedding representation is an embedding vector at the text fragment level. For example, an embedding vector represents the features of a text fragment in a candidate text.

[0160] For example, a text embedding representation is a sentence-level embedding vector. For instance, a text fragment can be a sentence, and an embedding vector can be used to represent a sentence in the candidate text.

[0161] Supervised semantic token sequences refer to discrete representations of text with specified sentiments. The first spectrogram can be a Mel spectrogram.

[0162] Obtain sentiment description information, which specifies the reference audio to be generated based on the desired sentiment. For example, the sentiment description information could be descriptive text such as "happy" or "sad." Obtain candidate text, which provides the audio content from the reference audio to be generated.

[0163] Candidate text can be divided into at least two text segments, and feature encoding can be performed on at least two text segments to obtain a text feature embedding representation sequence, which includes at least two text feature embedding representations. Each of the at least two text feature embedding representations corresponds one-to-one with at least two text segments.

[0164] A supervised semantic token sequence is generated based on sentiment description information and text embedding representation sequence. Preset noise is acquired and used as a generation condition with the supervised semantic token sequence to denoise the preset noise, converting it into a first spectrogram. The first spectrogram is then subjected to waveform transformation to obtain a sentiment reference audio, which contains sentiment description information indicating the sentiment.

[0165] For example, converting a reference audio into a first audio corresponding to a specified timbre includes: converting an emotional reference audio into a first audio corresponding to a specified timbre, the first audio having the emotion expressed by the emotional description information;

[0166] The first expressive feature includes the first emotional feature. Extracting the first timbre feature and the first expressive feature from the first audio, including: extracting the first timbre feature and the first emotional feature from the emotional reference audio;

[0167] The first audio content representation is obtained by concatenating the first text semantic features, the first phoneme, and the first expression features and then performing audio content prediction.

[0168] By converting emotional reference audio into a first audio with a specified timbre, the original emotion can be accurately extracted and preserved as an independent emotional feature. The explicit introduction of emotional features makes the synthesized audio more vivid and impactful. In the audio content prediction stage, the first text semantic features, phonemes, and first emotional features are concatenated to ensure consistency in semantics and emotion of the generated content. Combined with the first timbre feature for synthesis, the target audio not only possesses the specified timbre but also accurately reproduces the emotional rhythm of the reference audio. Through independent control of emotional expression and timbre, the problem of altering timbre disrupting the naturalness of emotion, common in traditional methods, is avoided. Furthermore, by decoupling timbre, emotional expression, and text content, high-quality emotional transfer and timbre customization are achieved.

[0169] For example, the first expression feature further includes a first prosodic feature, and the extraction of the first timbre feature and the first expression feature from the first audio includes: extracting the first timbre feature, the first prosodic feature and the first emotional feature from the emotional reference audio;

[0170] The first audio content representation is obtained by concatenating the first text semantic features, the first phoneme, and the first expression features and then performing audio content prediction. This includes concatenating the first text semantic features, the first phoneme, the first prosodic features, and the first sentiment features and then performing audio content prediction to obtain the first audio content representation.

[0171] This embodiment significantly improves the expressiveness and controllability of speech synthesis by decoupling and recombining the timbre, prosody, and emotional features in the emotional reference audio. Specifically, a first timbre feature, a first prosodic feature, and a first emotional feature are extracted from the emotional reference audio, achieving independent modeling of timbre and expression (i.e., prosody and emotion). This allows the synthesized audio to accurately transfer the emotional style and prosodic rhythm of the reference audio while preserving the timbre of the specified speaker, rather than simply imitating the timbre. Furthermore, by jointly predicting the audio content using the first text semantic feature, the first phoneme, the first prosodic feature, and the first emotional feature, consistency in semantics, pronunciation, intonation, and emotion is ensured in the generated content. Based on the joint synthesis of the first audio content representation, the first phoneme, and the first timbre feature, the target audio can be output with high naturalness, rich emotion, and realistic timbre, significantly improving the expressiveness and controllability of speech synthesis.

[0172] For example, such as Figure 7 The diagram shown illustrates the principle of an audio synthesis method.

[0173] The process involves acquiring sentiment description information and candidate text, performing feature encoding on the candidate text to obtain a text embedding representation sequence, and generating a supervised semantic token sequence based on the sentiment description information and the text embedding representation sequence. A preset noise is then acquired, and a spectrogram is extracted from the preset noise based on the supervised semantic token sequence. The extracted spectrogram undergoes waveform transformation to obtain a sentiment reference audio, which carries the sentiment expressed by the sentiment description information.

[0174] Feature extraction is performed on each frame of the emotional reference audio to obtain the reference content features, reference prosodic features, and reference emotional features of each frame.

[0175] Obtain a specified timbre identifier and convert it into a timbre embedding representation. Retrieve acoustic features matching the reference content features for each frame from the acoustic feature library corresponding to the specified timbre identifier. For each frame in the reference audio, perform feature transformation on the acoustic statistics of the frame based on the timbre embedding representation, the reference prosodic features of the frame, and the reference sentiment features of the frame, to obtain the first latent variable of the frame. Concatenate the first latent variables of each frame and decode the concatenated latent variables to obtain the first audio corresponding to the specified timbre. This first audio also carries the sentiment expressed by the sentiment description information.

[0176] First, obtain the first text, extract its semantic features, and convert it into a first phoneme. Then, extract the first timbre feature and the first expression feature from the first audio, where the first expression feature includes a first sentiment feature. Concatenate the first text semantic features, the first phoneme, and the first expression feature to predict the audio content, obtaining a first audio content representation. Based on the first audio content representation, the first phoneme, and the first timbre feature, perform audio synthesis to obtain a target audio with a specified timbre. This target audio also carries the sentiment expressed by the sentiment description information.

[0177] For example, the first audio corresponding to the specified timbre can be a first audio with the specified timbre.

[0178] For example, the first expressive feature also includes a first prosodic feature.

[0179] In this embodiment, by converting emotional description information into a supervised semantic token sequence and further extracting spectrograms and waveforms from preset noise, reference audio with vivid emotional color and high consistency with the description can be synthesized. This enables accurate generation of highly expressive emotional reference audio from emotional description information. Furthermore, it eliminates the dependence on real emotional samples, effectively solving the problems of scarce emotional data and difficult annotation in emotional speech synthesis. At the same time, it ensures the continuous controllability and subtle expression of the emotional dimension, laying a high-quality audio foundation for subsequent specified emotion transfer.

[0180] Furthermore, after acquiring the emotional reference audio, the content, prosody, and emotional features of each frame are extracted. These features are then combined with the user-specified timbre identifier to retrieve matching content features from the acoustic feature library. Subsequently, the acoustic statistics are transformed to obtain latent variables. Finally, the first audio with the target emotion and specified timbre is decoded. By decoupling and recombining timbre, prosody, emotion, and content features, emotional speech synthesis under the specified timbre is achieved, significantly improving the naturalness of emotional expression and the accuracy of timbre reproduction.

[0181] After obtaining the first audio, its timbre features and expressive features, including emotions, are further extracted. Combined with the semantic features and phoneme sequence of the first text, the final target audio is generated through two stages: content prediction and audio synthesis. This allows users to not only read any text with a specified timbre, but also to make the synthesized speech accurately carry the preset emotional color, taking into account timbre customization, emotional expressiveness, and content generalization ability.

[0182] For example, a spectrogram can be generated autoregressively using a Large Language Model (LLM). Specifically, the text embedding sequence and sentiment description information are input into the LLM, which then progressively generates supervised semantic tokens based on these information, resulting in a supervised semantic token sequence. Using this sequence as the generation condition, the spectrogram is extracted step-by-step from pre-defined noise by solving ordinary differential equations.

[0183] For example, the preset noise can be Gaussian noise. The converted spectrogram can be a Mel spectrogram, also known as a Mel spectrogram.

[0184] For example, the generation of emotion reference audio utilizes CosyVoice, a scalable multilingual zero-shot text-to-speech (TTS) synthesizer based on supervised semantic tokens. This primarily includes the extraction of supervised semantic tokens, the use of a large language model (LLM) for TTS, and optimal transport conditional stream matching.

[0185] Supervised semantic token extraction: Supervised semantic tokens are extracted using a multilingual speech recognition model (such as Whisper). Specifically, a vector quantization layer is inserted into the encoder of the speech recognition model to convert continuous speech signals into discrete token sequences. First, the input spectrogram undergoes position encoding and the first part of the encoder to obtain a context-aware feature representation H:

[0186]

[0187] This formula maps continuous feature representations H to the nearest discrete codebook vector, thereby compressing continuous acoustic features into a discrete token sequence. Here, X is the input spectrogram. It is the position code of X. It is the first part of the encoder.

[0188] Then, the vector quantizer converts the hidden representation of each frame. Mapping to the nearest codebook embedding Generate voice token :

[0189]

[0190] Here, VQ is the vector quantizer. C is the codebook, containing n learnable embedding vectors. . Represents the Euclidean distance (L2 norm), which measures the hidden representation. With codebook vector The closer the distance between two objects, the more similar they are. This represents the supervisory semantic token of the l-th frame, which represents the codebook index corresponding to that frame.

[0191] During training, the codebook embeddings are updated using an exponential moving average. As the attenuation coefficient:

[0192]

[0193] in, This represents the currently selected codebook embedding vector (i.e., the one that is embedded in the codebook). The closest one). This represents the attenuation coefficient.

[0194] Large Language Models (LLM) for TTS: The TTS task is formulated as an autoregressive speech token generation problem. LLM is used to learn the entire sequence of text encodings and speech tokens. Sequence construction includes start and end markers, speaker embedding vectors, text encodings, and speech tokens. A teacher forcing scheme is used during training, aiming to minimize the cross-entropy loss of the speech tokens.

[0195]

[0196] in, The l-th token predicted by LLM is The probability is given by L. L is the length of the voice token sequence (i.e., the total number of frames), and (L+1) is the total sequence length after adding the end marker. This formula is the loss calculation formula for an autoregressive language model. At each step, the model predicts the probability distribution of the next token and then performs cross-entropy with the actual token. The smaller the loss, the more accurately the LLM can predict "what the next token should be".

[0197] Optimal Transport Conditional Flow Matching (i.e., Inference-Generated Mel Speech Spectrum): The model uses the Optimal Transport Conditional Flow Matching model (OT-CFM) to learn the distribution of the Mel speech spectrogram and uses the generated speech tokens as conditional generation samples. This is achieved by constructing a model from the prior distribution... Distribution of Mel language spectrogram data This is achieved through a probability density path. The probability density path is formed by a time-dependent vector field. Defined as a flow generated by the following ordinary differential equation :

[0198]

[0199] Where t is the time variable, That is, t=0 is the starting point and t=1 is the ending point. The stream is a function that changes over time, representing the new position of X after transmission from t=0 to t. It represents the derivative of the flow with respect to time t, i.e., the rate of change of the flow. The time-dependent vector field is a learnable function that tells each location and each point in time what speed it should move at.

[0200] This indicates the position of the flow at time t=0, i.e., the starting point. This represents the prior distribution, which is usually taken as the standard multivariate Gaussian distribution N(0,I). This represents a Gaussian distribution with a mean of 0 and a covariance of I, i.e., noise with independent dimensions and a variance of I.

[0201] This indicates the position of the flow at time t=1, i.e., the endpoint. This represents the distribution of the target data, i.e., the empirical distribution of the generated Mel language spectrogram.

[0202] The vector field is learned by minimizing the following loss. This allows the predicted vector field to approximate the speed of the optimal transmission path (i.e., the model learning objective):

[0203]

[0204] in, This represents the expectation over time t, where t is typically uniformly sampled over [0,1]. Indicates the optimal transmission path. It is the weight function on the optimal transmission path. This represents the vector field (velocity) predicted by the model. During the generation process, a cosine scheduler and classifier free guidance are used to improve generation quality.

[0205] In this embodiment, by acquiring sentiment description information and candidate text, encoding the candidate text to obtain a text embedding representation, and then combining it with sentiment description information to generate a supervised semantic token sequence, a first spectrogram is extracted from preset noise, and finally, a sentiment reference audio with a specified sentiment is synthesized. This achieves zero-sample or few-sample controllable generation of text-to-sensory speech, enabling the generation of reference audio that conforms to the sentiment using sentiment description information without relying on the target sentiment reference audio, greatly expanding the flexibility and customizability of sentiment expression in speech synthesis. By guiding the conversion process from noise to spectrogram through the supervised semantic token sequence, the generated sentiment audio maintains the semantic accuracy of the text while possessing natural, coherent, and identifiable sentiment prosodic features (such as intonation, rhythm, and energy changes), avoiding the limitations of traditional sentiment synthesis methods that require a large amount of labeled sentiment data or specific speaker recordings.

[0206] In one embodiment, fusing a first audio content representation, a first phoneme, and a first timbre feature to obtain a target audio corresponding to a specified timbre includes:

[0207] Obtain the specified timbre identifier and convert it into a timbre embedding representation; fuse the first audio content representation, the first phoneme, the first timbre feature and the timbre embedding representation to obtain the target audio corresponding to the specified timbre.

[0208] For example, a specified timbre identifier is obtained and converted into a timbre embedding representation. Based on the first audio content representation, the first phoneme, the first timbre feature, and the timbre embedding representation, a fusion is performed, and the fused feature is waveform-converted to obtain the target audio corresponding to the specified timbre.

[0209] For example, audio synthesis is performed based on a first audio content representation sequence, a first phoneme sequence, a first timbre feature sequence, and a timbre embedding representation to obtain a target audio with a specified timbre. Specifically, a specified timbre identifier can be obtained, encoded, and used to obtain a timbre embedding representation. This timbre embedding representation can be a timbre embedding vector. The first audio content representation sequence, the first phoneme sequence, the first timbre feature sequence, and the timbre embedding representation are then fused, and the fused features are waveform-transformed to obtain the target audio with the specified timbre.

[0210] In this embodiment, when synthesizing audio based on the first audio content representation, the first phoneme, and the first timbre feature, a specified timbre identifier is additionally acquired and converted into a timbre embedding representation. This identifier, along with the aforementioned features, participates in audio synthesis. By introducing explicit timbre identifier embedding, the precise control and directional transfer capability of the target timbre are enhanced. This allows the synthesis system to not only rely on the implicit first timbre feature extracted from the prompt audio but also to use timbre embedding for explicit calibration and guidance. This effectively avoids timbre deviation caused by fluctuations in the quality of the reference audio or extraction errors, improving the robustness and consistency of timbre cloning. Furthermore, the fusion of the timbre embedding representation and the first timbre feature enables the model to flexibly handle scenarios with the same specified timbre and different reference audio. It preserves the personalized prosody and emotional details in the reference audio while ensuring that the timbre of the final target audio strictly aligns with the preset specified timbre identifier. This provides a more reliable timbre control mechanism for multi-speaker, multi-style speech synthesis systems and reduces the dependence on high-quality reference audio, enhancing the stability and adaptability of the solution in actual deployment.

[0211] In one embodiment, fusing a first audio content representation, a first phoneme, and a first timbre feature to obtain a target audio corresponding to a specified timbre includes:

[0212] The first audio content representation sequence, the first phoneme sequence, and the first spectrogram are fused to obtain the second spectrogram. The first audio content representation sequence includes the first audio content representation, the first phoneme sequence includes the first phoneme, the first timbre feature is a feature in the first timbre feature sequence, and the first timbre feature sequence is represented by the first spectrogram. The second spectrogram is waveform-transformed to obtain the target audio corresponding to the specified timbre.

[0213] For example, a specified timbre identifier can be obtained and converted into a timbre embedding representation. A second spectrogram is obtained by fusing the first audio content representation sequence, the first phoneme sequence, the first spectrogram, and the timbre embedding representation. The second spectrogram is then subjected to waveform transformation to obtain the target audio corresponding to the specified timbre.

[0214] In this embodiment, during audio synthesis, the first audio content representation sequence, the first phoneme sequence, and the first spectrogram are directly fused to obtain the second spectrogram. This second spectrogram is then obtained through waveform conversion to obtain the target audio. By introducing a fine-grained cue spectrogram as a timbre carrier, the timbre transfer of the target audio becomes more direct, accurate, and detailed. This is because the spectrogram contains acoustic details closely related to timbre in the original reference audio, such as the spectral envelope, formant structure, and harmonic distribution. Compared to abstract embedding vectors or statistical features, it can more completely preserve the personalized attributes of the specified timbre, such as timbre quality, warmth, and brightness. During the fusion process, the joint modeling of the first spectrogram with the content representation sequence and phoneme sequence achieves deep alignment of timbre information and speech content in the frequency domain. This avoids spectral breaks or unnatural transitions caused by separation processing in traditional methods, resulting in a significant improvement in both timbre similarity and naturalness of the synthesized target audio. Furthermore, by directly utilizing spectrograms for fusion and waveform conversion, the intermediate steps of timbre extraction and redrawing are simplified, cumulative errors are reduced, and the system can generate high-fidelity target audio with specified timbre more efficiently.

[0215] For example, such as Figure 8 The diagram shown is a schematic of an audio synthesis method in one embodiment.

[0216] The process involves acquiring sentiment description information and candidate text, performing feature encoding on the candidate text to obtain a text embedding representation sequence, and generating a supervised semantic token sequence based on the sentiment description information and the text embedding representation sequence. A preset noise is then acquired, and a spectrogram is extracted from the preset noise based on the supervised semantic token sequence. The extracted spectrogram undergoes waveform transformation to obtain a sentiment reference audio, which carries the sentiment expressed by the sentiment description information.

[0217] Feature extraction is performed on each frame of the emotional reference audio to obtain the reference content features, reference prosodic features, and reference emotional features of each frame.

[0218] Obtain a specified timbre identifier and convert it into a timbre embedding representation. Retrieve acoustic features matching the reference content features for each frame from the acoustic feature library corresponding to the specified timbre identifier. For each frame in the reference audio, perform feature transformation on the acoustic statistics of the frame based on the timbre embedding representation, the reference prosodic features of the frame, and the reference sentiment features of the frame, to obtain the first latent variable of the frame. Concatenate the first latent variables of each frame and decode the concatenated latent variables to obtain the first audio corresponding to the specified timbre. This first audio also carries the sentiment expressed by the sentiment description information.

[0219] First text is obtained, and its semantic feature sequence is extracted and converted into a first phoneme sequence. First timbre feature sequence and first expression feature sequence, including a first emotion feature sequence, are extracted from the first audio. The semantic feature sequence, phoneme sequence, and expression feature sequence are concatenated and used for audio content prediction to obtain a first audio content representation sequence. The first audio content representation sequence, phoneme sequence, spectrogram, and timbre embedding representation are fused to obtain a second spectrogram. The second spectrogram is then waveform-transformed to obtain the target audio with the specified timbre. This target audio also carries the emotion expressed by the emotion description information.

[0220] For example, the first audio corresponding to the specified timbre can be a first audio with the specified timbre.

[0221] For example, the first expression feature sequence also includes a first prosodic feature sequence.

[0222] In this embodiment, by converting emotional description information into a supervised semantic token sequence and further extracting spectrograms and waveforms from preset noise, reference audio with vivid emotional color and high consistency with the description can be synthesized. This enables accurate generation of highly expressive emotional reference audio from emotional description information. Furthermore, it eliminates the dependence on real emotional samples, effectively solving the problems of scarce emotional data and difficult annotation in emotional speech synthesis. At the same time, it ensures the continuous controllability and subtle expression of the emotional dimension, laying a high-quality audio foundation for subsequent specified emotion transfer.

[0223] Furthermore, after acquiring the reference audio, the content, prosody, and emotion features of each frame are extracted, and the matching content features are retrieved from the acoustic feature library in combination with the user-specified timbre identifier. Then, the acoustic statistics are transformed to obtain latent variables, and finally the first audio with the target emotion and specified timbre is decoded. By decoupling and recombining timbre, prosody, emotion, and content features, emotional speech synthesis under the specified timbre is realized, which significantly improves the naturalness of emotional expression and the accuracy of timbre reproduction.

[0224] After obtaining the first audio, its timbre features and expressive features, including emotional ones, are further extracted. Combined with the semantic features and phoneme sequence of the first text, and the joint modeling of the first spectrogram with the content representation sequence and phoneme sequence, deep alignment of timbre information and speech content in the frequency domain is achieved. This effectively enhances the timbre similarity and naturalness of the synthesized target audio. Utilizing spectrograms for fusion and waveform transformation simplifies the intermediate steps of timbre extraction and redrawing, reduces accumulated errors, and more efficiently generates high-fidelity target audio with specified timbre, balancing timbre customization, emotional expressiveness, and content generalization ability.

[0225] In one embodiment, such as Figure 9 As shown, a schematic diagram of audio synthesis is provided, including feature extraction, core two-stage modeling, and vocoder output.

[0226] Feature extraction layer:

[0227] First text: The first text is input into the encoder (BERT) and the phoneme converter respectively. BERT extracts the semantic feature sequence of the first text (extracting deep semantics, context, sentiment, and stress information). The phoneme converter converts the first text into a first phoneme sequence, thus converting the text into phonetic units and providing the specific pronunciation of each word.

[0228] The first audio audio is input into a feature extractor (CNHuBERT) to extract a first prosodic feature sequence (e.g., a first semantic token sequence) to extract pronunciation style features such as rhythm, prosody, and intonation. The first audio audio is then input into a spectrogram extractor to obtain a first spectrogram, which is used to capture the underlying physical timbre features. The first cue audio is input into a speaker verification (SV) model to obtain a timbre embedding representation. This timbre embedding representation is used to lock in the speaker's identity features and enhance the stability of timbre cloning.

[0229] The core two-stage modeling includes processing in Stage 1 and Stage 2, where:

[0230] Phase 1: Input the first text semantic feature sequence, the first phoneme sequence, and the first prosodic feature sequence into the autoregressive GPT. Use the autoregressive GPT model to predict the semantic token sequence frame by frame (this process is similar to the model first "reading" the text silently in its mind while imitating the speaking rhythm of the Prompt audio), to obtain the first audio content representation sequence. This first audio content representation sequence is the predicted semantic token sequence (this is a series of abstract "pronunciation instructions" that contain prosodic information such as intonation, rhythm, and stress of the first text, but does not contain specific timbre).

[0231] Phase 2: The predicted semantic token sequence, first phoneme sequence, first spectrogram, and timbre embedding representation input transformation model (e.g., the SoVITS model) are fused together (this model can determine "what sound to pronounce" based on phonemes and semantic tokens, and determine "what timbre to use" based on the reference cue spectrogram and timbre embedding representation) to obtain the second spectrogram. The second spectrogram is a two-dimensional time-frequency representation that contains all the acoustic details of the sound, but humans cannot hear it directly. Therefore, a vocoder (e.g., BigVGAN) is needed to map the second spectrogram into a one-dimensional audio waveform to obtain the target audio with the specified timbre.

[0232] For example, the sampling rate of a one-dimensional audio waveform can also be increased (e.g., from 24kHz to 48kHz) by using an audio super-resolution network (e.g., AP-BWE), increasing high-frequency details, eliminating dullness, and obtaining high-quality target audio.

[0233] In one embodiment, such as Figure 10 As shown, a model training method is provided. This method can be executed by the server or terminal alone, or by both the server and terminal. This method can be applied to... Figure 1 Taking a terminal as an example, this audio synthesis method may include the following steps:

[0234] Step S1002: Obtain sample audio and convert it into a second audio with a specified timbre through the conversion network in the initial audio synthesis model. The timbre of the sample audio is different from the specified timbre, which is selected from a variety of candidate timbres.

[0235] For example, converting sample audio into a second audio with a specified timbre using a transformation network in an initial audio synthesis model includes:

[0236] Using the transformation network in the initial audio synthesis model, reference content features and reference prosodic features are extracted from each frame of the sample audio to obtain a specified timbre identifier. The specified timbre identifier is converted into a timbre embedding representation. In the acoustic feature library corresponding to the specified timbre identifier, acoustic features that match the reference content features of each frame are retrieved. Based on each acoustic feature, timbre embedding representation and each reference prosodic feature, audio synthesis is performed to obtain a second audio with the specified timbre.

[0237] For example, the second audio corresponding to the specified timbre can be a second audio with the specified timbre.

[0238] For example, acoustic features include acoustic statistics; audio synthesis is performed based on each acoustic feature, timbre embedding representation, and each reference prosodic feature to obtain a second audio with a specified timbre, including:

[0239] For each frame in the sample audio, based on the timbre embedding representation and the reference prosodic features of the frame, feature transformation is performed on the acoustic statistics of the frame to obtain the third latent variable of the frame. The third latent variables of each frame are concatenated, and the concatenated latent variables are decoded to obtain the second audio corresponding to the specified timbre.

[0240] For example, the method further includes:

[0241] Candidate audio with a specified timbre is acquired. Candidate content features are extracted from each frame of the candidate audio. The waveform of the candidate audio is converted into a latent variable sequence, which includes the second latent variable of each frame of the candidate audio. The candidate content features of each frame are aligned with the second latent variable of each frame. Candidate content features of each frame that meet similar conditions are selected. Acoustic distribution statistics are performed based on the second latent variable aligned with the selected candidate content features of each frame to obtain the acoustic statistics that are common to the selected candidate content features of each frame. The candidate content features of each frame and the corresponding acoustic statistics are associated and stored in the acoustic feature library.

[0242] For example, the sample audio is an emotion reference audio, and the method further includes:

[0243] The process involves acquiring sentiment description information and candidate text, performing feature encoding on the candidate text to obtain a text embedding representation sequence, generating a supervised semantic token sequence based on the sentiment description information and the text embedding representation sequence, acquiring preset noise, extracting a spectrogram from the preset noise based on the supervised semantic token sequence, performing waveform transformation on the extracted spectrogram, and obtaining sentiment reference audio. The sentiment reference audio contains sentiment indicated by sentiment description information.

[0244] Step S1004: Obtain the second text, extract the semantic features of the second text through the feature extraction network in the initial audio synthesis model, convert the second text into a second phoneme, extract the second timbre features and the second expression features from the second audio. The second text is used to describe the content in the audio to be synthesized, and the second expression features include at least one of the second prosodic features or the second emotional features.

[0245] For example, the second text semantic feature sequence can be extracted from the second text through the feature extraction network in the initial audio synthesis model, and the second text can be converted into a second phoneme sequence. The second timbre feature sequence and the second expression feature sequence can be extracted from the second audio.

[0246] The second text semantic feature sequence is a sequence consisting of at least two second text semantic features, each of which is used to express the semantic features of a portion of the content in the second text. For example, each text semantic feature is used to express the semantic features of a segment, a sentence, a character, or a word in the second text.

[0247] A second phoneme sequence is a sequence consisting of at least two second phonemes. For example, the second text can be divided to obtain at least two segments, each segment potentially including at least one sentence. Semantic features of the second text are extracted from each segment, and these features are then concatenated sequentially to obtain a sequence of second text semantic features.

[0248] For example, the second text can be segmented by character or word using the transformation network in the initial audio synthesis model, and features can be extracted from each character or word to obtain the semantic features of the second text. The semantic features of the second text are then concatenated in sequence to obtain the semantic feature sequence of the second text.

[0249] For example, the pronunciation of each character in the second text can be determined, the phoneme corresponding to the pronunciation of each character (i.e., the second phoneme) can be determined, and the second phonemes can be concatenated in sequence to obtain a second phoneme sequence.

[0250] For example, the second audio can be divided into at least two frames using a transformation network in the initial audio synthesis model. For each frame, a second timbre feature and a second expression feature are extracted. The second timbre features of each frame are concatenated in sequence to form a second timbre feature sequence, and the second expression features of each frame are concatenated in sequence to form a second expression feature sequence.

[0251] For example, the second expression feature sequence includes a second prosodic feature sequence and / or a second emotional feature sequence.

[0252] Step S1006: The second audio content representation is obtained by concatenating the second text semantic features, the second phoneme, and the second expression features through the semantic prediction network in the initial audio synthesis model.

[0253] For example, the second audio content representation sequence can be obtained by concatenating the second text semantic feature sequence, the second phoneme sequence, and the second expression feature sequence through the semantic prediction network in the initial audio synthesis model.

[0254] The second text semantic feature sequence, the second phoneme sequence, and the second expression feature sequence can be concatenated by the semantic prediction network in the initial audio synthesis model. The concatenated feature sequence is then used to predict the audio content to obtain the second audio content representation sequence.

[0255] For example, the semantic prediction network may include an autoregressive prediction network, which concatenates the second text semantic feature sequence, the second phoneme sequence, and the second expression feature sequence, and performs autoregressive prediction on the concatenated feature sequence to obtain the second audio content representation sequence.

[0256] Step S1008: The second audio content representation, the second phoneme, and the second timbre feature are fused through the synthesis network in the initial audio synthesis model to obtain the third audio corresponding to the specified timbre.

[0257] For example, the second audio content representation sequence, the second phoneme sequence, and the second timbre feature sequence are fused together through the synthesis network in the initial audio synthesis model to obtain a third audio with a specified timbre.

[0258] For example, the method further includes: obtaining a specified timbre identifier, and converting the specified timbre identifier into a timbre embedding representation through a feature extraction network in the initial audio synthesis model;

[0259] By fusing the second audio content representation sequence, the second phoneme sequence, and the second timbre feature sequence through the synthesis network in the initial audio synthesis model, a third audio with a specified timbre is obtained, including:

[0260] By using the synthesis network in the initial audio synthesis model, the second audio content representation sequence, the second phoneme sequence, the second timbre feature sequence, and the timbre embedding representation are fused to obtain a third audio with a specified timbre.

[0261] For example, the second timbre feature sequence includes a training spectrogram. Through the synthesis network in the initial audio synthesis model, the second audio content representation sequence, the second phoneme sequence, the second timbre feature sequence, and the timbre embedding representation are fused to obtain a third audio with a specified timbre, including:

[0262] The synthesis network in the initial audio synthesis model is used to fuse the second audio content representation sequence, the second phoneme sequence, the training spectrogram, and the timbre embedding representation. The obtained spectrogram is then waveform-transformed to obtain a third audio with a specified timbre.

[0263] Step S1010: Determine the audio conversion loss based on the second audio and the first audio tag, and determine the audio synthesis loss based on the third audio and the second audio tag.

[0264] Audio conversion loss is the loss incurred by the conversion network in converting sample audio into a second audio. The first audio label can be real audio with a specified timbre, and the audio conversion loss characterizes the difference between the second audio generated from the conversion network and the real audio.

[0265] The difference between the second audio and the first audio tag can be determined; this difference is the audio conversion loss. The difference between the third audio and the second audio tag can be determined; this difference is the audio synthesis loss.

[0266] Step S1012: Train the initial audio synthesis model based on audio conversion loss and audio synthesis loss to obtain the audio synthesis model.

[0267] After determining the audio conversion loss and audio synthesis loss, the model parameters of the initial audio synthesis model can be adjusted based on these losses until the training termination condition is met, thus obtaining the final audio synthesis model. Alternatively, the total loss can be determined first based on the audio conversion and synthesis losses. The model parameters of the initial audio synthesis model can then be adjusted based on this total loss. Afterward, it can be determined whether the training termination condition is met. If the training termination condition is met, this initial audio synthesis model can be designated as the final audio synthesis model. If the training termination condition is not met, training can continue until it is satisfied.

[0268] The processing logic of the model training method is as follows: Figure 11 As shown.

[0269] In this embodiment, an initial audio synthesis model is constructed, comprising a conversion network, a feature extraction network, a semantic prediction network, and a synthesis network. The training data stream consists of a sample audio converted to generate a second audio with a specified timbre, a second text extracting semantic and phoneme sequences, and a prompt audio extracting timbre and expressive feature sequences. The audio conversion loss is calculated based on the prompt audio and the label, and the audio synthesis loss is calculated based on the third audio and the label. The entire model is jointly optimized, achieving end-to-end multi-task joint training. This enables the model to learn both "timbre conversion" and "speech synthesis" capabilities simultaneously. Furthermore, by sharing feature representations (such as timbre and expressive features), knowledge transfer between tasks is promoted, thereby improving the generalization ability and robustness of the final audio synthesis model in cloning a specified timbre and preserving expressive features. Furthermore, an audio conversion loss is introduced to constrain the conversion network, ensuring that the prompt audio retains the expressive features (such as rhythm and emotion) of the original reference audio after timbre transfer, avoiding information loss or distortion. An audio synthesis loss is introduced to constrain the overall generation quality, enabling the model to directly synthesize target audio with both specified timbre and rich expressiveness from text. This reduces the dependence on paired "text-speech" data, lowers training costs, and enhances the model's performance and timbre similarity in scenarios with few or zero samples.

[0270] In one embodiment, such as Figure 12 As shown, a processing logic diagram of a model training method is provided.

[0271] Preprocessing and feature extraction in the original training set: Obtain audio with a specified timbre, and perform preprocessing such as vocal accompaniment separation, sound noise reduction / enhancement, and audio segmentation to obtain audio segments.

[0272] Fine-tuning the initial audio synthesis model involves fine-tuning the transformation network in the initial audio synthesis model to obtain the VC model.

[0273] The conversion network consists of an encoder and a decoder. The encoder extracts speech features, and the decoder reconstructs the speech audio. The encoder uses a pre-trained model, ContentVec, and the decoder comprises a backbone (a conditional variational autoencoder VITS-e2e) and a vocoder (HifiGAN / NSF) architecture. The purpose of the training and fine-tuning phase is to learn the specified timbre: this phase aims to make the model remember the specified timbre to be converted. The main purpose of this phase is to "wash" the timbre from the original training set in the reconstructor to the target timbre, complete the fine-tuning of the reconstructor, and establish an acoustic feature library after fine-tuning.

[0274] The TTS model can be controlled via pre-trained instructions to generate cue audio based on sentiment description information and candidate text. The fine-tuned VC model then converts the sentiment audio into cue audio with a specified timbre, or vice versa, converts external reference audio into cue audio with a specified timbre.

[0275] Using a trained synthesis network, target audio with specified timbre, specified text, and specified emotion is synthesized based on text and audio prompts.

[0276] In one embodiment, the conversion network of the initial audio synthesis model has an encoder, a first generator, and a first discriminator. The conversion network in the initial audio synthesis model converts sample audio into a second audio corresponding to a specified timbre, including:

[0277] The encoder in the conversion network converts the sample audio into a third spectrogram with a specified timbre; the first generator in the conversion network converts the third spectrogram into a first waveform; the first waveform is the waveform of the second audio.

[0278] The audio conversion loss includes a first reconstruction loss and a first adversarial loss. The first audio label includes the true waveform of the third spectrogram. The audio conversion loss is determined based on the second audio and the first audio label, including:

[0279] The first discriminator in the transformation network is used to discriminate the first waveform and the real waveform of the third spectrogram, respectively, to obtain the first discrimination result of the first waveform and the second discrimination result of the real waveform; the first waveform is converted into the fourth spectrogram to determine the first reconstruction loss between the third spectrogram and the fourth spectrogram; based on the first discrimination result and the second discrimination result, the first adversarial loss is determined.

[0280] The first reconstruction loss ensures that the generated spectrogram is consistent with the original audio in content. The first adversarial loss enhances the discriminator's ability to better distinguish between real and synthesized audio.

[0281] Among them, the first feature matching loss can stabilize adversarial training and learn richer features from the intermediate layer of the discriminator.

[0282] like Figure 13 The diagram shows the architecture of the transformation network for the initial audio synthesis model. This network includes an encoder, a first generator, and a first discriminator. Specifically, the encoder can be a conditional variational autoencoder (VITS-e2e). The first generator can be a vocoder (HiFi-GAN) and includes the following main components:

[0283] The core task of the generator is to upsample the input Mel-language spectrogram and convert it into a continuous audio waveform.

[0284] Multi-Receptive Field Fusion (MRF) Module: To capture audio feature patterns of varying lengths in parallel, HiFi-GAN incorporates a Multi-Receptive Field Fusion (MRF) module within its generator. This module comprises multiple parallel residual blocks, each employing different kernel sizes and dilation rates to create diverse receptive fields. The MRF ultimately sums the outputs of these residual blocks, effectively enhancing the rich details of the synthesized audio.

[0285] The first discriminator can be a hybrid discriminator. To comprehensively and multi-dimensionally evaluate the quality of the generated audio, a dual discriminator system containing two different architectures can be used:

[0286] Multi-Period Discriminator (MPD): Focused on capturing the periodic features of audio. It consists of multiple sub-discriminators, each processing only discrete sample points with a specific period (interval) in the audio waveform. This design allows the model to capture high-frequency periodic patterns such as the fundamental frequency and harmonics of speech very well, greatly improving the realism of the sound.

[0287] Multi-Scale Discriminator (MSD): Focuses on capturing continuous patterns and long-term dependencies in audio (adapted from MelGAN). It consists of three sub-discriminators that evaluate audio at different resolution scales: the original audio waveform, audio downsampled by 2x (x2 average pooling), and audio downsampled by 4x (x4 average pooling).

[0288] The transformation network consists of a combination of the following three loss functions, used to jointly optimize the first generator and the first discriminator:

[0289] Adversarial Loss (GAN Loss / Real-Fake Loss): Used for basic adversarial training. The discriminator strives to distinguish between real and generated waveforms, while the generator attempts to produce realistic audio that can fool the MPD and MSD discriminators. The adversarial loss can be calculated using the following formula:

[0290]

[0291]

[0292] in, This represents the adversarial loss of the discriminator. Let (x, s) represent the expected value. Let (x, s) represent the real audio x sampled from the training data and the corresponding spectrogram s. Let D(x) represent the discriminator's judgment result on the real audio x. Let D(G(s)) represent the discriminator's judgment result on the fake audio G(s) generated by the generator. This represents the loss term of the actual audio. This represents the loss term for fake audio. Using this formula, the discriminator is trained to maximize the difference between real and fake audio outputs, i.e., to give real audio a high score (e.g., close to 1) and fake audio a low score (e.g., close to 0).

[0293] This represents the adversarial loss of the generator. (G;D) indicates that this is the loss with respect to the generator G (where the parameters of the discriminator D are fixed). This represents the expectation of the spectrogram of the conditional input. Here, there is no real audio x because the generator only generates fake audio. This indicates the discriminator's judgment result on the generated audio. The "deception" loss is represented by D(G(s)), which the generator aims to make the discriminator misclassify fake audio as real. This means D(G(s)) should be close to 1; when the discriminator outputs 1, this term is 0. Through this formula, the generator is trained to counter the discriminator, generating increasingly realistic audio that makes it impossible for the discriminator to distinguish between real and fake audio.

[0294] Mel-spectrogram Loss (L1 Loss), also known as reconstruction loss, calculates the L1 distance between the generated audio waveform and the corresponding Mel-spectrogram of the ground truth audio waveform. This effectively accelerates the early convergence speed of the generator and ensures high fidelity of the synthesized audio in the macroscopic frequency domain. The Mel-spectrogram Loss can be calculated using the following formula:

[0295]

[0296] in, This indicates Mel-spectral loss, which only applies to the generator G. A Melanographic spectrogram representing the actual audio value x. The Mel-language spectrogram represents the audio G(s) generated by the generator. This represents the L1 norm. This formula allows the generated Mel spectrum to approximate the real audio in the frequency domain, accelerating convergence and ensuring basic frequency domain fidelity.

[0297] Feature Matching Loss: Measures the L1 difference between real and generated samples in the feature maps extracted by each layer of the discriminator. This is an additional constraint that guides the generator not only to sound realistic in the final output, but also to maintain consistency with real audio in terms of deep feature representation.

[0298] Feature matching loss can be calculated using the following formula:

[0299]

[0300] in, represents the feature matching loss used in the generator G. T represents the number of layers (or feature extraction layers) in the discriminator D. Each sub-discriminator has multiple convolutional layers. The feature map representing the output of the i-th layer of the discriminator D refers to the feature map of the real audio x. The feature map representing the output of the i-th layer of the discriminator D refers to the feature map of the audio G(s) generated by the generator. This represents the total number of elements in the feature map of the i-th layer, used for normalization so that the losses of different layers can be added together. This represents the feature map output of the i-th layer of the discriminator D. and The difference between them is taken as L1 distance.

[0301] This means that after taking the L1 distance for the feature differences of the i-th layer, the average feature difference of the i-th layer is obtained by averaging the differences by the number of elements. Then, the summation is performed over all layers to obtain the total feature matching loss.

[0302] In this embodiment, an encoder, a first generator, and a first discriminator are introduced into the conversion network of the initial audio synthesis model. The encoder converts the sample audio into a third spectrogram corresponding to the specified timbre, and the first generator converts it into a first waveform as the waveform of the second audio. The first reconstruction loss between the first waveform and the real waveform (by comparing the first waveform back to the fourth spectrogram with the third spectrogram) and the adversarial loss based on the first discriminator are calculated respectively. This significantly improves the fidelity and perceived naturalness of the timbre conversion. The reconstruction loss forces the converted waveform to accurately align with the acoustic details of the target timbre at the spectral level, avoiding information loss or spectral distortion. The first adversarial loss, through game-theoretic training of the discriminator, makes the generated first waveform statistically approximate the waveform of the real target timbre, thereby eliminating common artifacts such as mechanical sound and metallic sound, and synthesizing a prompt audio with high naturalness and high timbre similarity. By jointly optimizing the first reconstruction loss and the first adversarial loss to form dual constraints of spectral consistency and perceptual realism, the conversion network can learn the decoupling and transfer from the source timbre to the target timbre without relying on paired training data (same content, different timbres). This provides high-quality, timbre-pure cue audio for subsequent audio synthesis tasks, while also enhancing the model's robustness to noise or non-ideal reference audio.

[0303] In one embodiment, the audio conversion loss further includes a first feature matching loss, and the method further includes:

[0304] By using the first discriminator in the transformation network, multi-scale feature maps of the first waveform and the real waveform are extracted respectively; based on the multi-scale feature maps of the first waveform and the real waveform, the first feature matching loss is determined.

[0305] For example, the transformation network has an encoder, a first generator, and a first discriminator. Figure 14 As shown, a schematic diagram of the processing principle in a conversion network is provided.

[0306] The sample audio is acquired, and the encoder in the conversion network converts the sample audio into a third spectrogram corresponding to the specified timbre. The first generator in the conversion network then converts the third spectrogram into a first waveform, which is the waveform of the second audio.

[0307] By using the first discriminator in the transformation network, the first waveform and the real waveform of the third spectrogram are discriminated respectively, and the first discrimination result of the first waveform and the second discrimination result of the real waveform are obtained. Based on the first discrimination result and the second discrimination result, the first adversarial loss is determined.

[0308] For example, the difference between the second discrimination result and 1 is calculated and the square of the difference is calculated (i.e., the first squared value). The square of the first discrimination result is calculated (i.e., the second squared value). The first squared value and the second squared value are summed to obtain the first adversarial loss.

[0309] For example, the first adversarial loss includes a first discriminator loss and a first generator loss, which can be implemented as described above. Figure 13 The formula for calculating the adversarial loss of the discriminator is used to calculate the first discriminator loss. That is, the difference between the second discrimination result and 1 is calculated and the square of the difference is calculated (i.e., the first squared value), the square of the first discrimination result is calculated (i.e., the second squared value), and the sum of the first squared value and the second squared value is used as the first discriminator loss.

[0310] According to the above Figure 13 The formula for calculating the adversarial loss of the generator is used to calculate the first generator loss. That is, the difference between the first discrimination result and 1 is calculated, and the square of this difference is calculated to obtain the first generator loss.

[0311] By transforming the first generator in the network, the first waveform is converted into a fourth spectrogram, and the first reconstruction loss between the third and fourth spectrograms is determined.

[0312] For example, the spectrogram difference between the third and fourth spectrograms is calculated, and this difference is the first reconstruction loss.

[0313] Furthermore, the difference between the third and fourth spectrograms can be expressed using the L1 norm (i.e., the L1 distance), and this L1 distance is the first reconstruction loss.

[0314] The encoder in the transformation network extracts multi-scale feature maps of the first waveform and the real waveform, respectively. The first discriminator in the transformation network then determines the first feature matching loss based on the multi-scale feature maps of the first waveform and the real waveform.

[0315] For example, the feature differences between the multi-scale feature maps of the real waveform and the first waveform at the corresponding scale are calculated, thereby obtaining the feature differences between the real waveform and the first waveform at each scale of the multi-scale model. The average of the feature differences at each scale is then calculated to obtain the first feature matching loss.

[0316] Furthermore, the difference between the multi-scale feature map of the real waveform and the multi-scale feature map of the first waveform at the corresponding scale can be calculated as the L1 distance, which is the feature difference at the same scale.

[0317] For example, the first discriminator may include multiple sub-discriminators, each of which has multiple convolutional layers, and each convolutional layer outputs a feature map of one scale, as described above. Figure 13The formula for calculating the feature matching loss is used to calculate the first feature matching loss. For example, for each sub-discriminator, the L1 distance is taken for the feature differences of the i-th layer (i.e., the convolutional layer) of the target discriminator, and then the average is obtained by averaging the number of elements to get the average feature difference of that layer. Then, the summation is applied to all layers to obtain the total feature matching loss of a sub-discriminator. The average of the total feature matching losses of all sub-discriminators is then taken to obtain the first feature matching loss.

[0318] For example, it can be done according to the above. Figure 13 The formula for calculating Mel spectrogram loss is used to calculate the first reconstruction loss.

[0319] In this embodiment, a first feature matching loss is further introduced into the audio conversion loss. The first discriminator in the conversion network extracts multi-scale feature maps of the first waveform (generated prompt audio waveform) and the real waveform, and calculates the feature matching loss based on the difference between the two. This significantly enhances the similarity between the generated waveform and the real waveform at the intermediate semantic level, rather than being limited to the adversarial nature of the final discrimination result. The multi-scale feature maps can capture acoustic features of different granularities, from local details (such as short-term spectral fluctuations) to global structures (such as long-term envelopes). The feature matching loss forces the internal feature representation of the generated waveform to be consistent with the real waveform, thereby effectively guiding the generator to learn a mapping that is closer to the real data distribution and reducing the pattern collapse or artifact problems that may occur due to relying solely on adversarial loss. At the same time, this loss provides dense supervision signals for the intermediate layers of the discriminator, accelerates model convergence and improves training stability, so that the prompt audio generated by the conversion network is further improved in terms of timbre reproduction, waveform continuity and naturalness of sound, and finally synthesizes high-quality audio that is more difficult to distinguish from the real one.

[0320] In one embodiment, the second timbre feature sequence includes a training spectrogram, and the synthesis network of the initial audio synthesis model has a second generator and a second discriminator; fusing the second audio content representation sequence, the second phoneme sequence, and the second timbre feature sequence to obtain a third audio with a specified timbre includes:

[0321] The second generator in the synthesis network fuses the second audio content representation sequence, the second phoneme sequence, and the training spectrogram to obtain the fifth spectrogram, which is the spectrogram of the third audio with a specified timbre. The second audio content representation sequence includes the second audio content representation, the second phoneme sequence includes the second phonemes, and the second timbre feature is the feature in the second timbre feature sequence, which is represented by the training spectrogram. The second generator in the synthesis network converts the fourth spectrogram into a second waveform, which is the waveform of the third audio.

[0322] The audio synthesis loss includes a second reconstruction loss, and the second audio label includes the second true waveform of the fifth spectrogram. The audio synthesis loss is determined based on the third audio and the second audio label, including:

[0323] The second discriminator in the synthetic network is used to discriminate the real waveforms of the second waveform and the fifth spectrogram, respectively, to obtain the third discrimination result of the second waveform and the fourth discrimination result of the real waveform of the fifth spectrogram; based on the third discrimination result and the fourth discrimination result, the second adversarial loss is determined.

[0324] For example, the difference between the fourth discrimination result and 1 is calculated and the square of the difference is calculated (i.e., the third squared value). The square of the third discrimination result is calculated (i.e., the fourth squared value). The third squared value and the fourth squared value are summed to obtain the second adversarial loss.

[0325] For example, the second adversarial loss includes a second discriminator loss and a second generator loss, which can be implemented as described above. Figure 13 The formula for calculating the adversarial loss of the discriminator is used to calculate the second discriminator loss. That is, the difference between the fourth discrimination result and 1 is calculated and the square of the difference is calculated (i.e., the third squared value). The square of the third discrimination result is calculated (i.e., the fourth squared value). The third squared value and the fourth squared value are summed to obtain the second discriminator loss.

[0326] According to the above Figure 13 The formula for calculating the adversarial loss of the generator is used to calculate the second generator loss. That is, the difference between the third discrimination result and 1 is calculated, and the square of this difference is calculated to obtain the second generator loss.

[0327] For example, the audio synthesis loss also includes a second reconstruction loss, and the method further includes:

[0328] The second waveform is converted into a sixth spectrogram by the second generator in the synthesis network, and the second reconstruction loss between the sixth spectrogram and the fifth spectrogram is determined.

[0329] For example, the spectrogram difference between the fifth and sixth spectrograms is calculated, and this difference is the second reconstruction loss.

[0330] Furthermore, the difference between the fifth and sixth spectrograms can be expressed using the L1 norm (i.e., the L1 distance), and this L1 distance is the second reconstruction loss.

[0331] For example, it can be done according to the above. Figure 13 The formula for calculating Mel spectrogram loss is used to calculate the second reconstruction loss.

[0332] In this embodiment, a second generator and a second discriminator are set in the synthesis network. The second generator fuses the second audio content representation sequence, the second phoneme sequence, and the training spectrogram to obtain the fifth spectrogram, which is then converted into the second waveform as the waveform of the third audio. The second reconstruction loss between this waveform and the real waveform is calculated (by converting the second waveform back to the sixth spectrogram and comparing it with the fifth spectrogram), and the second adversarial loss based on the second discriminator is calculated. This achieves end-to-end spectrum-waveform joint optimization in the synthesis stage. The reconstruction loss forces the generated third audio to have accurate acoustic details in the spectral domain and the target timbre. Alignment ensures that the fusion of content (semantics, phonemes) and timbre features is not distorted; adversarial loss, through game-theoretic training of the discriminator, makes the statistical distribution of the synthesized waveform approximate the natural waveform of the real target timbre, significantly improving the realism of the sound and the similarity of the timbre; the two work together to effectively suppress synthesis artifacts such as discontinuous spectral splicing, metallic sound, and mechanical feel, so that the third audio retains rich expressive features (such as rhythm and emotion) while possessing the specified timbre, providing high-quality self-supervised signals for model training, thereby improving the generalization ability and synthesis naturalness of the final audio synthesis model in scenarios with few or zero samples.

[0333] In one embodiment, the audio synthesis loss further includes a second feature matching loss, and the method further includes:

[0334] The second discriminator in the synthetic network extracts the multi-scale feature maps of the second waveform and the real waveform, respectively; based on the multi-scale feature maps of the second waveform and the real waveform, the second feature matching loss is determined.

[0335] For example, the multi-scale feature maps of the second waveform and the real waveform can be extracted separately by the second discriminator in the synthetic network. The feature differences between the multi-scale feature maps of the real waveform and the second waveform at the corresponding scale are calculated, thus obtaining the feature differences between the second waveform and the real waveform at each scale of the multi-scale network. The average of the feature differences at each scale is then calculated to obtain the second feature matching loss.

[0336] For example, the difference between the multi-scale feature map of the real waveform and the multi-scale feature map of the second waveform at the same scale can be calculated by taking the L1 distance. This L1 distance is the feature difference at the same scale.

[0337] For example, the second discriminator may include multiple sub-discriminators, each with multiple convolutional layers, and each convolutional layer outputs a feature map of one scale, as described above. Figure 13The formula for calculating the feature matching loss is used to calculate the second feature matching loss. For example, for each sub-discriminator, the L1 distance is taken for the feature differences of the i-th layer (i.e., the convolutional layer) of the target discriminator, and then the average is obtained by averaging the number of elements to get the average feature difference of that layer. Then, the sum is applied to all layers to obtain the total feature matching loss of a sub-discriminator. The average of the total feature matching losses of all sub-discriminators is then used to obtain the second feature matching loss.

[0338] In this embodiment, a second feature matching loss is further introduced into the audio synthesis loss. The second discriminator in the synthesis network extracts multi-scale feature maps of the second waveform (the generated cue audio waveform) and the real waveform, respectively. The feature matching loss is calculated based on the difference between the two, significantly enhancing the similarity between the generated and real waveforms at the intermediate semantic level, rather than being limited to adversarial games in the final discrimination result. The multi-scale feature maps capture acoustic features of different granularities, from local details (such as short-term spectral fluctuations) to global structures (such as long-term envelopes). The feature matching loss forces the internal feature representation of the generated waveform to be consistent with the real waveform, effectively guiding the generator to learn a mapping closer to the real data distribution and reducing the pattern collapse or artifact problems that may arise from relying solely on adversarial losses. Furthermore, this loss provides dense supervision signals for the intermediate layers of the discriminator, accelerating model convergence and improving training stability. This results in further improvements in the timbre reproduction, waveform continuity, and naturalness of the generated cue audio, ultimately synthesizing high-quality audio that is more difficult to distinguish from real audio.

[0339] In one embodiment, a model training method is provided, applied to a computer device. The initial audio synthesis model includes a conversion network, a feature extraction network, a semantic prediction network, and a synthesis network. The conversion network has an encoder, a first generator, and a first discriminator, and the synthesis network has a second generator and a second discriminator. The method includes:

[0340] The sample audio is acquired, and the encoder in the conversion network converts the sample audio into a third spectrogram corresponding to the specified timbre. The first generator in the conversion network then converts the third spectrogram into a first waveform, which is the waveform of the second audio.

[0341] By using the first discriminator in the transformation network, the first waveform and the real waveform of the third spectrogram are discriminated respectively, and the first discrimination result of the first waveform and the second discrimination result of the real waveform are obtained. Based on the first discrimination result and the second discrimination result, the first adversarial loss is determined.

[0342] By transforming the first generator in the network, the first waveform is converted into a fourth spectrogram, and the first reconstruction loss between the third and fourth spectrograms is determined.

[0343] The encoder in the transformation network extracts multi-scale feature maps of the first waveform and the real waveform, respectively. The first discriminator in the transformation network then determines the first feature matching loss based on the multi-scale feature maps of the first waveform and the real waveform.

[0344] The second text is obtained, and the semantic feature sequence of the second text is extracted through the feature extraction network in the initial audio synthesis model. The second text is then converted into a second phoneme sequence. The second timbre feature sequence and the second expression feature sequence are extracted from the second audio. The second timbre feature sequence is represented by the training spectrogram.

[0345] By using the semantic prediction network in the initial audio synthesis model, the second text semantic feature sequence, the second phoneme sequence, and the second expression feature sequence are concatenated to predict the audio content, thereby obtaining the second audio content representation sequence.

[0346] The second generator in the synthesis network fuses the second audio content representation sequence, the second phoneme sequence, and the training spectrogram to obtain the fifth spectrogram, which is the spectrogram of the third audio corresponding to the specified timbre.

[0347] The second generator in the synthesis network converts the fourth spectrogram into a second waveform, which is the waveform of the third audio. The second discriminator in the synthesis network then distinguishes between the second waveform and the actual waveform of the fifth spectrogram, obtaining a third discrimination result for the second waveform and a fourth discrimination result for the actual waveform of the fifth spectrogram.

[0348] The second waveform is converted into the sixth spectrogram, and the second reconstruction loss between the sixth and fifth spectrograms is determined.

[0349] Based on the third and fourth discrimination results, the second adversarial loss is determined.

[0350] The second discriminator in the synthetic network extracts the multi-scale feature maps of the second waveform and the real waveform, respectively. Based on the multi-scale feature maps of the second waveform and the real waveform, the second feature matching loss is determined.

[0351] The total loss is calculated based on the first feature matching loss, the second feature matching loss, the first adversarial loss, the second adversarial loss, the first reconstruction loss, and the second reconstruction loss.

[0352] The model parameters of the initial audio synthesis model are adjusted based on the total loss, and the audio synthesis model is obtained when the training termination condition is met.

[0353] In this embodiment, the encoder in the conversion network converts the sample audio into a third spectrogram with a specified timbre. A first generator then generates a first waveform, and a first reconstruction loss is calculated between the third and fourth spectrograms, enabling the model to accurately preserve the acoustic details of the original audio. Simultaneously, the first discriminator, through adversarial and feature matching losses, forces the generated first waveform to be indistinguishable from the real waveform in terms of overall distribution and multi-scale features. This joint constraint of "reconstruction + adversarial + feature matching" effectively avoids content distortion or timbre drift during the timbre conversion process, significantly improving the timbre similarity and waveform quality of the synthesized audio.

[0354] Furthermore, a semantic prediction network concatenates text semantic features, phoneme sequences, and expression feature sequences to predict the audio content representation sequence, enabling the model to fully understand the semantic and prosodic information of the text. The second generator, based on this content representation, phoneme sequence, and training spectrogram, fuses to obtain a fifth spectrogram, which is then further converted into a second waveform. Through joint optimization of the second reconstruction loss, adversarial loss, and feature matching loss, the generated audio not only meets the specified timbre requirements but also exhibits natural, fluent speech with rich expression, avoiding the problems of mechanical synthesis or lack of emotion found in traditional methods.

[0355] Furthermore, a multi-level, multi-task joint training strategy was adopted, significantly improving the model's generalization ability and training stability. On one hand, the conversion network and the synthesis network each have independent generators and discriminators, each undertaking the tasks of timbre conversion and speech synthesis, respectively. They form an information bridge by sharing intermediate features such as audio content representation, avoiding the shortcomings of a single network in handling multiple objectives. On the other hand, the scheme simultaneously uses reconstruction loss, adversarial loss, and feature matching loss. Among them, the feature matching loss constrains the generator's output at the multi-scale feature map level, effectively mitigating the problems of pattern collapse and training oscillation in adversarial training. This multi-loss collaborative mechanism enables the model to converge more smoothly during training and has a stronger adaptability to inputs with different timbres and different texts.

[0356] Finally, without requiring paired training data, an end-to-end capability was achieved to extract timbre from arbitrary audio samples and synthesize corresponding text speech. By separating and recombining timbre features (represented through training spectrograms), expressive features, and text semantic features, the model can transfer the timbre of the source audio to the read-aloud speech of the target text without requiring the same speaker to read all the labeled data of the text. Simultaneously, the first reconstruction loss ensures that the acoustic characteristics of the original audio are not destroyed during timbre conversion, while the second reconstruction loss guarantees the alignment of the synthesized speech with the expected content. This unpaired training method significantly reduces the dependence on high-quality labeled data, making the model highly adaptable and deployable in practical application scenarios such as personalized speech synthesis and few-sample timbre cloning.

[0357] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0358] Based on the same inventive concept, this application also provides an audio synthesis apparatus for implementing the audio synthesis method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations of one or at least two audio synthesis apparatus embodiments provided below can be found in the limitations of the audio synthesis method described above, and will not be repeated here.

[0359] In one embodiment, such as Figure 15 As shown, an audio synthesis apparatus 1500 is provided, comprising:

[0360] The first conversion module 1502 is used to acquire reference audio and convert the reference audio into a first audio with a corresponding specified timbre. The timbre of the reference audio is different from the specified timbre, which is selected from a variety of candidate timbres.

[0361] The first feature extraction module 1504 is used to acquire the first text, extract the semantic features of the first text, convert the first text into a first phoneme, extract the first timbre features and the first expression features from the first audio. The first text is used to describe the content in the audio to be synthesized, and the first expression features include at least one of the first prosodic features or the first emotional features.

[0362] The first prediction module 1506 is used to concatenate the first text semantic features, the first phoneme and the first expression features to predict the audio content and obtain the first audio content representation.

[0363] The first synthesis module 1508 is used to fuse the first audio content representation, the first phoneme and the first timbre feature to obtain the target audio corresponding to the specified timbre.

[0364] In one embodiment, the first conversion module 1502 is used to extract reference content features and reference prosodic features from each frame of the reference audio, obtain a specified timbre identifier, convert the specified timbre identifier into a timbre embedding representation, retrieve acoustic features that match the reference content features of each frame in the acoustic feature library corresponding to the specified timbre identifier, and perform audio synthesis based on each acoustic feature, timbre embedding representation and each reference prosodic feature to obtain the first audio corresponding to the specified timbre.

[0365] In one embodiment, the acoustic features include acoustic statistics; the first conversion module 1502 is used to perform feature conversion on the acoustic statistics of the target frame for each frame in the reference audio based on the timbre embedding representation and the reference prosodic features of the target frame, to obtain the first latent variable of the target frame, to concatenate the first latent variables of each frame, and to decode the concatenated latent variables to obtain the first audio corresponding to the specified timbre.

[0366] In one embodiment, the device further includes:

[0367] The feature library construction module is used to acquire candidate audio with a specified timbre, extract candidate content features for each frame of the candidate audio; convert the waveform of the candidate audio into a latent variable sequence, which includes the second latent variable of each frame of the candidate audio; align the candidate content features of each frame with the second latent variable of each frame, filter out the candidate content features of each frame that meet similar conditions, perform acoustic distribution statistics based on the second latent variable aligned with the selected candidate content features of each frame, and obtain the acoustic statistics that are common to the selected candidate content features of each frame; and associate and store the candidate content features of each frame with the corresponding acoustic statistics in the acoustic feature library.

[0368] In one embodiment, the first expressive feature includes at least one of a first prosodic feature or a first emotional feature.

[0369] In one embodiment, the reference audio is an emotion reference audio. The first conversion module 1502 is used to acquire emotion description information and candidate text, perform feature encoding on the candidate text to obtain a text embedding representation sequence; generate a supervised semantic token sequence based on the emotion description information and the text embedding representation sequence; acquire preset noise, and convert a spectrogram from the preset noise based on the supervised semantic token sequence; perform waveform conversion on the converted spectrogram to obtain the emotion reference audio, which has the emotion indicated by the emotion description information.

[0370] In one embodiment, the first synthesis module 1508 is used to obtain a specified timbre identifier, convert the specified timbre identifier into a timbre embedding representation, and fuse the first audio content representation, the first phoneme, the first timbre feature and the timbre embedding representation to obtain the target audio corresponding to the specified timbre.

[0371] In one embodiment, the first synthesis module 1508 is used to fuse the first audio content representation sequence, the first phoneme sequence, and the first spectrogram to obtain a second spectrogram. The first audio content representation sequence includes the first audio content representation, the first phoneme sequence includes the first phoneme, the first timbre feature is a feature in the first timbre feature sequence, and the first timbre feature sequence is represented by the first spectrogram. The second spectrogram is then subjected to waveform conversion to obtain the target audio corresponding to the specified timbre.

[0372] Based on the same inventive concept, this application also provides a model training apparatus for implementing the model training method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations of one or at least two model training apparatus embodiments provided below can be found in the limitations of the model training method described above, and will not be repeated here.

[0373] In one embodiment, such as Figure 16 As shown, a model training device 1600 is provided, comprising:

[0374] The second conversion module 1602 is used to acquire sample audio and convert it into a second audio with a specified timbre through the conversion network in the initial audio synthesis model. The timbre of the sample audio is different from the specified timbre, which is selected from a variety of candidate timbres.

[0375] The second feature extraction module 1604 is used to obtain the second text, extract the semantic features of the second text through the feature extraction network in the initial audio synthesis model, convert the second text into a second phoneme, extract the second timbre features and the second expression features from the second audio, the second text is used to describe the content in the audio to be synthesized, and the second expression features include at least one of the second prosodic features or the second emotional features.

[0376] The second prediction module 1606 is used to predict the audio content by concatenating the second text semantic features, the second phoneme, and the second expression features through the semantic prediction network in the initial audio synthesis model, thereby obtaining the second audio content representation.

[0377] The second synthesis module 1608 is used to fuse the second audio content representation, the second phoneme, and the second timbre features through the synthesis network in the initial audio synthesis model to obtain a third audio with a specified timbre.

[0378] The loss determination module 1610 is used to determine the audio conversion loss based on the second audio and the first audio label, and to determine the audio synthesis loss based on the third audio and the second audio label.

[0379] Training module 1612 is used to train the initial audio synthesis model based on audio conversion loss and audio synthesis loss to obtain the audio synthesis model.

[0380] In one embodiment, the conversion network of the initial audio synthesis model has an encoder, a first generator, and a first discriminator. The second conversion module 1602 is used to convert the sample audio into a third spectrogram corresponding to a specified timbre through the encoder in the conversion network; and to convert the third spectrogram into a first waveform through the first generator in the conversion network; the first waveform is the waveform of the second audio.

[0381] The audio conversion loss includes a first reconstruction loss and a first adversarial loss. The first audio label includes the real waveform of the third spectrogram. The loss determination module 1610 is used to discriminate the first waveform and the real waveform of the third spectrogram respectively through the first discriminator in the conversion network to obtain a first discrimination result of the first waveform and a second discrimination result of the real waveform. The first waveform is converted into a fourth spectrogram to determine the first reconstruction loss between the third spectrogram and the fourth spectrogram. Based on the first discrimination result and the second discrimination result, the first adversarial loss is determined.

[0382] In one embodiment, the audio conversion loss further includes a first feature matching loss. The loss determination module 1610 is used to extract the multi-scale feature map of the first waveform and the multi-scale feature map of the real waveform through the first discriminator in the conversion network; and to determine the first feature matching loss based on the multi-scale feature map of the first waveform and the multi-scale feature map of the real waveform.

[0383] In one embodiment, the second timbre feature sequence includes a training spectrogram, and the synthesis network of the initial audio synthesis model has a second generator and a second discriminator;

[0384] The second synthesis module 1608 is used to obtain a fifth spectrogram by fusing a second audio content representation sequence, a second phoneme sequence, and a training spectrogram through a second generator in the synthesis network. The fifth spectrogram is the spectrogram of the third audio corresponding to a specified timbre. The second audio content representation sequence includes the second audio content representation, the second phoneme sequence includes the second phonemes, and the second timbre feature is a feature in the second timbre feature sequence, which is represented by the training spectrogram. The second generator in the synthesis network converts the fourth spectrogram into a second waveform, which is the waveform of the third audio.

[0385] The audio synthesis loss includes a second reconstruction loss and a second adversarial loss. The second audio tag includes the second real waveform of the fifth spectrogram. The loss determination module 1610 is used to discriminate the second waveform and the real waveform of the fifth spectrogram respectively through the second discriminator in the synthesis network to obtain the third discrimination result of the second waveform and the fourth discrimination result of the real waveform of the fifth spectrogram. The second waveform is converted into a sixth spectrogram to determine the second reconstruction loss between the sixth spectrogram and the fifth spectrogram. Based on the third discrimination result and the fourth discrimination result, the second adversarial loss is determined.

[0386] Each module in the aforementioned devices can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0387] In one embodiment, a computer device is provided, which may be a terminal or a server. Taking a server as an example, its internal structure diagram may be as follows: Figure 17 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores audio synthesis data and model training data. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the aforementioned audio synthesis methods and model training methods.

[0388] Those skilled in the art will understand that Figure 17 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0389] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0390] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0391] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0392] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0393] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, etc., and are not limited to these.

[0394] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0395] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An audio synthesis method, characterized in that, The method includes: Obtain a reference audio, convert the reference audio into a first audio with a specified timbre, wherein the timbre of the reference audio is different from the specified timbre, and the specified timbre is a timbre selected from a variety of candidate timbres; Obtain the first text, extract the semantic features of the first text, and convert the first text into the first phoneme. The first text is used to describe the content in the audio to be synthesized. Extract a first timbre feature and a first expression feature from the first audio, wherein the first expression feature includes at least one of a first prosodic feature or a first emotional feature; The first text semantic features, the first phoneme and the first expression features are concatenated and then audio content prediction is performed to obtain the first audio content representation; The first audio content representation, the first phoneme, and the first timbre feature are fused together to obtain the target audio corresponding to the specified timbre.

2. The method according to claim 1, characterized in that, The step of converting the reference audio into a first audio with a corresponding specified timbre includes: Extract reference content features and reference prosodic features from each frame of the reference audio; Obtain the specified timbre identifier and convert the specified timbre identifier into a timbre embedding representation; In the acoustic feature library corresponding to the specified timbre identifier, retrieve the acoustic features that match the reference content features of each frame; Audio synthesis is performed based on the acoustic features, the timbre embedding representation, and the reference prosodic features to obtain a first audio with a specified timbre.

3. The method according to claim 2, characterized in that, The acoustic features include acoustic statistics; the audio synthesis based on each of the acoustic features, the timbre embedding representation, and each of the reference prosodic features to obtain a first audio corresponding to a specified timbre includes: For each frame in the reference audio, based on the timbre embedding representation and the reference prosodic features of the frame, feature transformation is performed on the acoustic statistics of the frame to obtain the first latent variable of the frame. The first latent variables of each frame are concatenated, and the concatenated latent variables are decoded to obtain the first audio corresponding to the specified timbre.

4. The method according to claim 2 or 3, characterized in that, The method further includes: Acquire candidate audio with a specified timbre, and extract candidate content features for each frame of the candidate audio; The waveform of the candidate audio is converted into a sequence of latent variables, the sequence of latent variables including the second latent variable of each frame of the candidate audio; Align the candidate content features of each frame with the second latent variable of each frame, filter out the candidate content features of each frame that meet the similarity conditions, perform acoustic distribution statistics based on the second latent variable aligned with the selected candidate content features of each frame, and obtain the acoustic statistics that are commonly corresponding to the selected candidate content features of each frame. The candidate content features and corresponding acoustic statistics of each frame are associated and stored in the acoustic feature library.

5. The method according to claim 1, characterized in that, The reference audio is an emotional reference audio, and the method further includes: Obtain sentiment description information and candidate text, perform feature encoding on the candidate text, and obtain a text embedding representation sequence; A supervised semantic token sequence is generated based on the sentiment description information and the text embedding representation sequence; Obtain preset noise, and extract a spectrogram from the preset noise based on the supervised semantic token sequence; The converted spectrogram is subjected to waveform conversion to obtain an emotional reference audio, which has the emotion expressed by the emotional description information.

6. The method according to claim 1, characterized in that, The step of fusing the first audio content representation, the first phoneme, and the first timbre feature to obtain the target audio corresponding to the specified timbre includes: Obtain the specified timbre identifier and convert the specified timbre identifier into a timbre embedding representation; The first audio content representation, the first phoneme, the first timbre feature, and the timbre embedding representation are fused to obtain the target audio corresponding to the specified timbre.

7. The method according to claim 1, characterized in that, The step of fusing the first audio content representation, the first phoneme, and the first timbre feature to obtain the target audio corresponding to the specified timbre includes: The first audio content representation sequence, the first phoneme sequence, and the first spectrogram are fused to obtain the second spectrogram; the first audio content representation sequence includes the first audio content representation, the first phoneme sequence includes the first phoneme, the first timbre feature is a feature in the first timbre feature sequence, and the first timbre feature sequence is represented by the first spectrogram. The second spectrogram is waveform-converted to obtain the target audio corresponding to the specified timbre.

8. A model training method, characterized in that, The method includes: A sample audio is acquired and converted into a second audio with a specified timbre through a conversion network in an initial audio synthesis model. The timbre of the sample audio is different from the specified timbre, which is selected from a variety of candidate timbres. The second text is obtained, and its semantic features are extracted through the feature extraction network in the initial audio synthesis model. The second text is then converted into a second phoneme. The second timbre features and second expression features are extracted from the second audio. The second text is used to describe the content in the audio to be synthesized. The second expression features include at least one of the second prosodic features or the second emotional features. The second audio content representation is obtained by concatenating the second text semantic features, the second phoneme and the second expression features through the semantic prediction network in the initial audio synthesis model and then predicting the audio content. By fusing the second audio content representation, the second phoneme, and the second timbre feature through the synthesis network in the initial audio synthesis model, a third audio corresponding to the specified timbre is obtained; The audio conversion loss is determined based on the second audio and the first audio tag, and the audio synthesis loss is determined based on the third audio and the second audio tag; The initial audio synthesis model is trained based on the audio conversion loss and the audio synthesis loss to obtain the audio synthesis model.

9. The method according to claim 8, characterized in that, The initial audio synthesis model's conversion network includes an encoder, a first generator, and a first discriminator. The step of converting the sample audio into a second audio with a specified timbre using the conversion network in the initial audio synthesis model includes: The encoder in the conversion network converts the sample audio into a third spectrogram corresponding to the specified timbre. The third spectrogram is converted into a first waveform by the first generator in the conversion network; the first waveform is the waveform of the second audio. The audio conversion loss includes a first reconstruction loss and a first adversarial loss, the first audio tag includes the true waveform of the third spectrogram, and the step of determining the audio conversion loss based on the second audio and the first audio tag includes: The first discriminator in the conversion network is used to distinguish between the first waveform and the real waveform of the third spectrogram, respectively, to obtain a first discrimination result of the first waveform and a second discrimination result of the real waveform; The first waveform is converted into a fourth spectrogram, and the first reconstruction loss between the third spectrogram and the fourth spectrogram is determined. Based on the first discrimination result and the second discrimination result, the first adversarial loss is determined.

10. The method according to claim 9, characterized in that, The audio conversion loss also includes a first feature matching loss, and the method further includes: The first discriminator in the transformation network extracts the multi-scale feature map of the first waveform and the multi-scale feature map of the real waveform, respectively. Based on the multi-scale feature map of the first waveform and the multi-scale feature map of the real waveform, the first feature matching loss is determined.

11. The method according to any one of claims 8 to 10, characterized in that, The synthesis network of the initial audio synthesis model has a second generator and a second discriminator; The step of fusing the second audio content representation, the second phoneme, and the second timbre feature to obtain a third audio corresponding to the specified timbre includes: The second generator in the synthesis network fuses the second audio content representation sequence, the second phoneme sequence, and the training spectrogram to obtain a fifth spectrogram. The fifth spectrogram is a spectrogram of the third audio corresponding to the specified timbre. The second audio content representation sequence includes the second audio content representation, the second phoneme sequence includes the second phoneme, and the second timbre feature is a feature in the second timbre feature sequence. The second timbre feature sequence is represented by the training spectrogram. The fourth spectrogram is converted into a second waveform by the second generator in the synthesis network, and the second waveform is the waveform of the third audio. The audio synthesis loss includes a second adversarial loss, the second audio tag includes the second true waveform of the fifth spectrogram, and the determination of the audio synthesis loss based on the third audio and the second audio tag includes: The second discriminator in the synthesis network is used to distinguish between the second waveform and the real waveform of the fifth spectrogram, respectively, to obtain the third discrimination result of the second waveform and the fourth discrimination result of the real waveform of the fifth spectrogram. Based on the third and fourth discrimination results, the second adversarial loss is determined.

12. The method according to claim 11, characterized in that, The audio synthesis loss also includes a second reconstruction loss, and the method further includes: The second waveform is converted into a sixth spectrogram by the second generator in the synthesis network; Determine the second reconstruction loss between the sixth spectrogram and the fifth spectrogram.

13. An audio synthesis device, characterized in that, The device includes: The first conversion module is used to acquire reference audio and convert the reference audio into a first audio with a corresponding specified timbre. The timbre of the reference audio is different from the specified timbre, and the specified timbre is a timbre selected from a variety of candidate timbres. The first feature extraction module is used to acquire first text, extract semantic features from the first text, and convert the first text into first phonemes. The first text is used to describe the content in the audio to be synthesized. The module also extracts first timbre features and first expression features from the first audio. The first expression features include at least one of first prosodic features or first emotional features. The first prediction module is used to concatenate the first text semantic features, the first phoneme and the first expression features to predict the audio content and obtain the first audio content representation. The first synthesis module is used to fuse the first audio content representation, the first phoneme, and the first timbre feature to obtain the target audio corresponding to the specified timbre.

14. A model training device, characterized in that, The device includes: The second conversion module is used to acquire sample audio and convert it into a second audio with a specified timbre through the conversion network in the initial audio synthesis model. The timbre of the sample audio is different from the specified timbre, which is selected from a variety of candidate timbres. The second feature extraction module is used to acquire the second text, extract the semantic features of the second text through the feature extraction network in the initial audio synthesis model, convert the second text into a second phoneme, extract the second timbre features and the second expression features from the second audio, the second text is used to describe the content in the audio to be synthesized, and the second expression features include at least one of the second prosodic features or the second emotional features. The second prediction module is used to concatenate the second text semantic features, the second phoneme, and the second expression features through the semantic prediction network in the initial audio synthesis model to predict the audio content and obtain the second audio content representation. The second synthesis module is used to fuse the second audio content representation, the second phoneme, and the second timbre feature through the synthesis network in the initial audio synthesis model to obtain a third audio corresponding to the specified timbre; The loss determination module is used to determine audio conversion loss based on the second audio and the first audio tag, and to determine audio synthesis loss based on the third audio and the second audio tag; The training module is used to train the initial audio synthesis model based on the audio conversion loss and the audio synthesis loss to obtain the audio synthesis model.

15. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 12.

16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.

17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 12.