Rhythm migration method and device, electronic equipment and storage medium
By acquiring decoupled prosodic and timbre features, the problems of insufficient prosodic expression and timbre leakage in existing speech synthesis are solved, achieving high-quality cross-speaker prosodic transfer and generating natural and vivid audio.
Patent Information
- Application Number
- CN202511372197.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2026-01-02
AI Technical Summary
Existing speech synthesis technology struggles to accurately express complex prosodic details during prosodic transfer, resulting in a lack of expressiveness in synthesized speech. Furthermore, severe timbre leakage occurs during cross-speaker prosodic transfer, affecting the purity of the target speaker's timbre.
By acquiring and utilizing decoupled prosodic and timbre features, the prosody of the source speech and the timbre of the target speaker are characterized respectively. A target speech vector sequence is generated using multimodal input, and the target audio is synthesized through an acoustic decoder to ensure timbre purity.
It improves the quality of cross-speaker prosodic transfer, maintains the purity of the target speaker's timbre, and enhances the naturalness and expressiveness of synthesized speech.
Smart Images

Figure CN121260146A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech processing technology, and in particular to a prosody transfer method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the development of deep learning technology, modern speech synthesis no longer only pursues clarity and naturalness of pronunciation, but also pays more attention to the expressiveness of speech. Prosody transfer is one of the key technologies to achieve this goal. It aims to apply the speaking style (such as intonation, rhythm and emotion) of a source speech to the speech to be synthesized by the target speaker in order to generate more expressive audio.
[0003] Currently, common prosodic transfer schemes explicitly decompose speech signals into prosodic features representing speaking style and timbre features representing speaker identity. This scheme extracts the prosody of the source speech and the timbre of the target speaker, merges them, and then uses a vocoder to synthesize the final speech waveform. However, this approach struggles to accurately capture and express complex prosodic details, resulting in synthesized speech that lacks expressiveness and sounds bland. Furthermore, when performing cross-speaker prosodic transfer, the synthesized speech often contains elements of the source speaker's timbre, affecting the purity of the target speaker's timbre. Summary of the Invention
[0004] This invention provides a prosody transfer method, apparatus, electronic device, and storage medium to address the shortcomings of existing speech synthesis technologies, such as limited prosodic expression capabilities and severe timbre leakage during cross-speaker prosodic transfer.
[0005] This invention provides a prosodic transfer method, comprising: Obtain decoupled prosodic features based on source prosodic speech and decoupled timbre features based on target speaker speech, wherein the decoupled prosodic features characterize the prosody of the source prosodic speech and the decoupled timbre features characterize the timbre of the target speaker speech; Based on the text features of the target text, the speech features of the target speaker, the decoupled prosodic features, and the decoupled timbre features, a target speech vector sequence is generated; Based on the target speech vector sequence, the target audio is synthesized.
[0006] According to a prosodic transfer method provided by the present invention, the step of obtaining decoupled prosodic features based on source prosodic speech and decoupled timbre features based on target speaker speech includes: Based on the speech vector encoder, the source prosodic speech is encoded into source speech features, and the source speech features are mapped to the prosodic feature space through the prosodic mapper to obtain prosodic mapping features. Based on the speech vector encoder, the speech of the target speaker is encoded into target speech features, and the target speech features are mapped to the timbre feature space through the timbre mapper to obtain timbre mapping features; Based on the prosodic timbre decoupling module, the prosodic mapping feature and the timbre mapping feature are decoupled respectively to obtain the decoupled prosodic feature and the decoupled timbre feature.
[0007] According to a prosody transfer method provided by the present invention, the step of mapping the source speech features to a prosody feature space through a prosody mapper to obtain prosody mapping features includes: Based on a text encoder, the source speech text is encoded to obtain the text features of the source speech text, and the source speech text corresponds to the source prosodic speech in terms of content. Based on the prosodic encoder, the source speech features and the text features of the source speech text are encoded to obtain the initial prosodic features; Based on the prosodic mapper, the initial prosodic features are mapped to the prosodic feature space to obtain the prosodic mapping features.
[0008] According to a prosodic transfer method provided by the present invention, the prosodic timbre decoupling module is trained using a contrastive learning strategy, and the training steps of the prosodic timbre decoupling module include: Construct positive and negative sample pairs for prosodic features and positive and negative sample pairs for timbre features; The prosodic feature positive and negative sample pairs and the timbre feature positive and negative sample pairs are input into the initial decoupling module to obtain the decoupling result output by the initial decoupling module. Based on the decoupling result, the prosodic contrast loss, timbre contrast loss and cross-contrast loss are calculated. The cross-contrast loss is used to suppress the correlation between the decoupled prosodic features and the decoupled timbre features in the decoupling result. Based on the prosodic contrast loss, the timbre contrast loss, and the cross-contrast loss, a combined loss is determined, and the parameters of the initial decoupling module are updated and iterated according to the combined loss to obtain the prosodic timbre decoupling module.
[0009] According to a prosodic transfer method provided by the present invention, the construction of positive and negative sample pairs of prosodic features and positive and negative sample pairs of timbre features includes: According to a prosody transfer method provided by the present invention, the training steps of the speech vector encoder include: By using automated annotation methods, large-scale raw audio data is processed to obtain a pre-trained corpus that includes prosodic information and speaker information; Based on the pre-trained corpus, the speech vector encoder is trained to convert continuous speech signals into discrete speech features.
[0010] According to a prosodic transfer method provided by the present invention, the method involves processing large-scale raw audio data using an automated annotation method to obtain a pre-trained corpus including prosodic information and speaker information, comprising: The original audio data is subjected to speech recognition to generate audio-text data pairs; Align the audio and text data pairs to obtain phoneme-level aligned data; Prosodic features and speaker features are extracted from the phoneme-level aligned data to form the pre-trained corpus.
[0011] According to a prosody transfer method provided by the present invention, generating a target speech vector sequence based on text features of the target text, speech features of the target speaker's speech, decoupled prosody features, and decoupled timbre features includes: Based on the text features of the target text, the speech features of the target speaker, the decoupled prosodic features, and the decoupled timbre features, a target speech vector sequence is generated through an autoregressive generation model. The autoregressive generation model is used to predict the speech vector of the next frame in an autoregressive manner. The autoregressive generation model is fine-tuned based on a labeled speech dataset.
[0012] According to a prosody transfer method provided by the present invention, the step of synthesizing target audio based on the target speech vector sequence includes: Based on an acoustic decoder, the target speech vector sequence is converted into an intermediate acoustic representation, wherein the acoustic decoder converts the target speech vector sequence based on the decoupled prosodic features and the decoupled timbre features; The intermediate acoustic representation is synthesized into the target audio based on a neural vocoder.
[0013] The present invention also provides a prosody transfer device, comprising: The acquisition unit is used to acquire decoupled prosodic features based on source prosodic speech and decoupled timbre features based on target speaker speech, wherein the decoupled prosodic features characterize the prosody of the source prosodic speech and the decoupled timbre features characterize the timbre of the target speaker speech. The generation unit is used to generate a target speech vector sequence based on the text features of the target text, the speech features of the target speaker's speech, the decoupled prosodic features, and the decoupled timbre features; A synthesis unit is used to synthesize target audio based on the target speech vector sequence.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the prosody transfer method as described above.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the prosody transfer method as described above.
[0016] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the prosody transfer method as described above.
[0017] The prosodic transfer method, apparatus, electronic device, and storage medium provided by this invention effectively alleviate the problem of feature mixing by acquiring decoupled prosodic features based on the source prosodic speech and decoupled timbre features based on the target speaker's speech. This is because the decoupled prosodic features can more purely represent the prosody of the source speech and basically do not contain its timbre information; similarly, the decoupled timbre features can also more purely represent the timbre of the target speaker. In the synthesis process, combining a pure decoupled prosodic feature with a pure decoupled timbre feature can ensure that the final synthesized audio can highly maintain the purity of the target speaker's own timbre, thereby effectively improving the quality of cross-speaker prosodic transfer. Furthermore, by acquiring decoupled prosodic features, this invention can separate and retain more complete and refined prosodic information from complex speech signals. Since the interference of timbre components is avoided, the model can focus on modeling the prosody itself, thereby learning richer and more detailed prosodic representations. When generating the target speech vector sequence, this high-quality decoupled prosodic feature is used as an input condition, enabling the model to more accurately reproduce the prosodic style of the source speech. The final synthesized audio is closer to the source prosodic speech in terms of emotion, tone, and rhythm, enhancing expressiveness and making the listening experience more natural and vivid. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the prosody transfer scheme in traditional speech synthesis. Figure 2 This is a flowchart illustrating the prosody transfer method provided by the present invention; Figure 3This is a schematic diagram of the pluggable rhythm and timbre control module provided by the present invention; Figure 4 This is a schematic diagram of the process for constructing the pre-trained corpus provided by the present invention; Figure 5 This is a training schematic diagram of the speech vector encoder provided by the present invention; Figure 6 This is a schematic diagram of the input and output of the autoregressive generative model provided by the present invention; Figure 7 This is a schematic diagram of acoustic modeling and waveform synthesis provided by the present invention; Figure 8 This is a schematic diagram of the prosody transfer device provided by the present invention; Figure 9 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0021] Traditional speech synthesis systems primarily revolve around the task of "text-to-speech," with the core objective of accurately, clearly, and naturally reading the input text aloud. In early technical solutions, such as concatenation-based or rule-based TTS (Text-to-Speech) systems, the synthesized speech often sounded mechanical, rigid, and lacked emotion and expressiveness due to extremely limited control over speech prosody.
[0022] With the development of deep learning technology, end-to-end neural network models, such as Tacotron and FastSpeech, have gradually become mainstream, and speech synthesis technology has made significant progress in clarity, naturalness, and coherence. Building on this foundation, research is gradually expanding towards more expressive directions such as multi-speaker, multilingual, and multi-style approaches.
[0023] Among these, prosodic transfer has become a key research direction. This technology aims to transfer the speaking style (also known as intonation, rhythm, and emotional style) of a source speech to the target text or the speech of a target speaker, in order to generate synthesized speech with similar intonation, rhythm, and emotional style to the source speech. This task focuses not only on "what to say" and "who is saying," but also on "how to say it."
[0024] Specifically, based on the relationship between the source and target speakers, prosodic transfer tasks can be divided into two categories: single-speaker prosodic transfer and cross-speaker prosodic transfer. Single-speaker prosodic transfer refers to preserving the timbre characteristics of a speaker within the same speech synthesis system and transferring the prosodic style of a reference speech to new text, thereby achieving individualized tone adjustment and emotional expression. Cross-speaker prosodic transfer, on the other hand, refers to transferring prosodic features such as intonation, rhythm, and emotion from other speakers while maintaining the target speaker's timbre, enabling the target speaker to express various emotions and intonations not present in the training data, thus achieving prosodic style sharing between different speakers. Prosodic transfer technology has broad application prospects in fields such as virtual digital humans, audio content creation, and intelligent interaction.
[0025] Currently, most common speech prosody transfer schemes are small to medium scale. Due to limitations in computing resources and available data, researchers usually rely on manual prior design to explicitly decompose speech signals into prosodic features (such as fundamental frequency F0, duration, energy, etc.) and timbre features (such as spectral envelope, Mel-Cepstral coefficients, formants, etc.). Then, through rule-based, statistical, or neural network methods, the prosody in the source speech is transferred to the timbre of the target speaker. Figure 1 This is a flowchart illustrating the prosody transfer scheme in traditional speech synthesis, such as... Figure 1 As shown, this type of solution has a relatively simple overall structure, consisting of the following two core steps: (1) Separation and extraction of speech features. This stage mainly relies on domain knowledge of speech signal processing. For the source speech, the prosodic extraction module extracts prosodic feature vectors that can represent its prosodic style and changes. For example, the F0 trajectory is obtained through the fundamental frequency extraction algorithm to represent the intonation fluctuations and tone; the duration distribution of phoneme levels is analyzed through algorithms such as Dynamic Time Warping (DTW) to characterize speech rate and rhythm changes. For the target speech, the timbre extraction module extracts acoustic feature vectors related to timbre preservation. For example, the spectral envelope contour is extracted using Linear Predictive Coding (LPC), or Mel-Frequency Cepstral Coefficients (MFCC) are obtained based on the cepstral analysis method, thereby capturing the personalized voiceprint features of the target speaker, such as timbre and formant distribution.
[0026] (2) Feature mapping and speech synthesis. After obtaining the prosodic features of the source speaker and the timbre features of the target speaker, a mapping mechanism needs to be constructed to achieve the fusion and reconstruction of the two. The mapping process often adopts the following three strategies: ① Rule mapping, based on linguistic or phonetic prior rules, the prosodic features of the source speech are normalized and then directly replaced in the target speech structure. For example, the phoneme duration ratio of the source speech is mapped to the articulation unit of the target speaker; ② Statistical modeling methods, such as Gaussian Mixture Model (GMM) and Hidden Markov Model (HMM), model the correspondence between speech features by learning the joint probability distribution of source-target feature pairs; ③ Neural network methods, such as deep neural network (DNN) and long short-term memory (LSTM) deep learning models, based on a large amount of labeled data, use backpropagation to optimize network parameters and establish a nonlinear mapping relationship from source prosody to target timbre. After mapping, the fused features can be fed into a vocoder module (such as Griffin-Lim, WORLD, STRAIGHT, etc.) for waveform reconstruction, ultimately generating target synthesized speech that combines the timbre of the target speaker with the prosodic style of the source speaker.
[0027] Although small- to medium-sized prosodic transfer models have demonstrated some performance on multiple public datasets and are practical in simple scenarios, they still have the following shortcomings when faced with complex speech prosodic transfer scenarios: First, existing speech prosodic datasets are significantly insufficient in terms of scale, speaker diversity, and prosodic distribution. There are few publicly available high-quality datasets, and those that exist generally suffer from uneven distribution of prosodic styles and a lack of refined prosodic labels, making it difficult to support the modeling of complex prosodic phenomena and the demand for high-quality prosodic transfer.
[0028] Secondly, the ability to express rhythm is limited. Existing solutions struggle to accurately convey the intricate rhythmic details such as complex changes in speech rate, emotional shifts, and intonation fluctuations, resulting in synthesized speech that lacks expressiveness and sounds rather bland.
[0029] Secondly, timbre leakage is a significant issue during cross-speaker prosodic transfer. When performing cross-speaker prosodic transfer, the synthesized speech often contains the source speaker's timbre, resulting in an impure target speaker's timbre. This problem is particularly pronounced when multi-speaker corpora are scarce.
[0030] Finally, the granularity of prosody control is relatively coarse. Existing solutions rely heavily on acoustic features designed manually by researchers, and are limited by the model parameter size, resulting in insufficient modeling ability for speech. This makes it difficult to achieve fine-grained, high-fidelity control of prosody, which in turn affects the overall quality and expressiveness of synthesized speech.
[0031] To address the aforementioned shortcomings, this invention proposes a prosodic transfer method. By acquiring and utilizing decoupled prosodic and timbre features, it can accurately transfer the prosodic style of the source speech to the target speaker while maintaining the stability of the target speaker's timbre. This effectively solves the timbre leakage problem commonly found in cross-speaker prosodic transfer techniques. Furthermore, by introducing multimodal inputs (text, speech, prosodic, and timbre) to jointly guide the generation of speech vector sequences, the synthesized speech achieves high levels of accuracy in content, timbre similarity, and prosodic expressiveness, enhancing the naturalness and expressiveness of the synthesized speech. This provides a feasible technical path for achieving high-fidelity, highly controllable personalized speech synthesis.
[0032] Figure 2 This is a flowchart illustrating the prosody transfer method provided by the present invention, as shown below. Figure 2 As shown, the method includes: Step S10: Obtain decoupled prosodic features based on the source prosodic speech and decoupled timbre features based on the target speaker's speech, wherein the decoupled prosodic features characterize the prosody of the source prosodic speech and the decoupled timbre features characterize the timbre of the target speaker's speech.
[0033] It should be noted that the prosody transfer method provided in this embodiment of the invention can be applied to the field of artificial intelligence, particularly in scenarios such as speech synthesis, virtual digital humans, and audio content creation. The execution entity of this method can be a server, a terminal device (such as a personal computer or smartphone), a cloud computing platform, or any electronic device containing the functions of the aforementioned devices.
[0034] Specifically, in step S10, two types of input speech need to be processed first: source prosodic speech and target speaker speech. Here, source prosodic speech refers to a reference speech segment that provides the desired speaking style. The text content of this speech is not important; the key is that it contains the prosodic style that is to be transferred. Prosody here is a broad concept and can include, but is not limited to, a combination of suprasegmental features such as intonation (e.g., the flat intonation of a declarative sentence and the rising intonation of an interrogative sentence), rhythm, speech rate, stress, and emotional style (e.g., happiness, sadness, anger). For example, if a user wants to synthesize a speech with a "news broadcast" style, then the source prosodic speech could be a real news broadcast recording.
[0035] The target speaker's voice refers to a reference speech segment that provides the desired timbre. This voice segment is used to determine "who is speaking" in the final synthesized speech. Timbre here is a key acoustic characteristic that distinguishes different people's voices, primarily determined by the physical characteristics of the vocal organs, manifested in aspects such as spectral envelope and formant distribution; it can be understood as the speaker's personalized voiceprint characteristics. For example, if a user wants to synthesize the voice of "Zhang San," then the target speaker's voice could be one or more recordings of "Zhang San."
[0036] By processing the source prosodic speech and the target speaker's speech, we can obtain the decoupled prosodic features of the source prosodic speech and the decoupled timbre features of the target speaker's speech. Here, decoupled prosodic features refer to a digital representation extracted from the source prosodic speech that purely represents its prosodic information, and this representation has removed the timbre information of the source speaker as much as possible. The purpose of the decoupling operation is to prevent timbre leakage in subsequent synthesis, that is, to avoid the final synthesized speech containing both the timbre of the target speaker and the timbre of the source speaker.
[0037] Decoupled timbre features refer to a digital representation extracted from the speech of a target speaker that stably represents their timbre identity and is insensitive to prosodic changes. Similarly, the decoupling operation allows the timbre feature to more purely represent the speaker's identity, unaffected by the specific intonation or emotion of their speech.
[0038] There are several ways to obtain these two decoupled features. In one example, the source prosodic speech and the target speaker's speech can be input into a pre-trained feature decoupling model. This model is specially designed and trained to learn to separate the entangled prosodic and timbre information in the speech signal and output the corresponding feature vectors respectively. This model can be implemented based on a deep neural network architecture, for example, by using techniques such as adversarial training, contrastive learning, or information bottlenecks to make the prosodic and timbre orthogonal or uncorrelated in the feature space.
[0039] Step S20: Generate a target speech vector sequence based on the text features of the target text, the speech features of the target speaker, the decoupled prosodic features, and the decoupled timbre features.
[0040] Specifically, after obtaining the decoupled prosodic features that control prosodic style and the decoupled timbre features that control timbre, an intermediate representation, namely the target speech vector sequence, can be generated based on these features, as well as the text features of the input target text and the speech features of the target speaker.
[0041] Here, the target text refers to the specific text content that the synthesized speech is intended to be. For example, "The weather is really nice today." In order for the model to understand the semantics of the text, the target text needs to be converted into a machine-readable numerical representation. This is usually achieved through a text encoder, such as using a pre-trained language model, like BERT (Bidirectional Encoder Representation from Transformers), to extract the semantic embeddings of the text and obtain the text feature vector.
[0042] The speech features of the target speaker differ from the decoupled timbre features. They primarily serve as a global condition or identifier for the generative model, guiding it to generate speech that conforms to the overall acoustic characteristics of the target speaker. This feature can be extracted from the target speaker's speech using a separate speaker encoder or speech vector encoder. It can be a discrete speaker ID or a continuous embedding vector.
[0043] Specifically, by inputting the four features mentioned above (i.e., textual features of the target text, speech features of the target speaker, decoupled prosodic features, and decoupled timbre features) into a generative model, a target speech vector sequence can be generated. Here, the target speech vector sequence is a high-level, compressed intermediate representation of speech; it is not a directly audible audio signal or Mel spectrum, but rather a series of discrete tokens or continuous vectors. This sequence already contains the semantics of the text, the timbre of the target speaker, and the prosodic features of the source speech.
[0044] Preferably, the process of generating the target speech vector sequence can employ an autoregressive model, such as a generative pre-trained language model based on the Transformer architecture. This model fuses four input features and then predicts the speech vector for the next frame frame by frame or individually. Specifically, when predicting the speech vector for the t-th frame, the model uses all control information such as previous text, timbre, and prosody, as well as the already generated speech vectors from the previous t-1 frames, as conditions.
[0045] Step S30: Based on the target speech vector sequence, synthesize the target audio.
[0046] Specifically, this step converts the high-level intermediate representation generated in step S20 into the final playable audio waveform. This process typically consists of two stages: acoustic modeling and waveform synthesis. In the acoustic modeling stage, the target speech vector sequence can be input into an acoustic decoder, which is responsible for converting the high-level, discrete speech vector sequence into an acoustic representation that more closely resembles the physical world, such as the Mel spectrum. Here, the Mel spectrum is a two-dimensional image describing the energy of a sound signal as a function of frequency and time, and is a widely used intermediate representation in speech synthesis.
[0047] In the waveform synthesis stage, the Mel spectrum generated by the acoustic decoder is fed into a vocoder. A vocoder is a neural network model specifically designed to reconstruct the original audio waveform with high quality from acoustic representations (such as the Mel spectrum). For example, high-fidelity neural vocoders such as HiFi-GAN and BigVGAN can be used to accomplish this step, generating the final target audio.
[0048] The final output target audio is a continuous digital audio signal (such as a WAV format file). Its speech content is consistent with the target text, its timbre is consistent with the target speaker's speech, and its prosodic style is consistent with the source prosodic speech, thus realizing cross-speaker prosodic transfer.
[0049] The method provided in this invention effectively alleviates the problem of feature mixing by acquiring decoupled prosodic features based on the source prosodic speech and decoupled timbre features based on the target speaker's speech. This is because decoupled prosodic features can more purely represent the prosody of the source speech and basically do not contain its timbre information; similarly, decoupled timbre features can also more purely represent the timbre of the target speaker. In the synthesis process, combining a pure decoupled prosodic feature with a pure decoupled timbre feature can ensure that the final synthesized audio can highly maintain the purity of the target speaker's own timbre, thereby effectively improving the quality of cross-speaker prosodic transfer. Furthermore, by acquiring decoupled prosodic features, this invention can separate and retain more complete and refined prosodic information from complex speech signals. Since the interference of timbre components is avoided, the model can focus on modeling the prosody itself, thereby learning richer and more detailed prosodic representations. When generating the target speech vector sequence, this high-quality decoupled prosodic feature is used as an input condition, enabling the model to more accurately reproduce the prosodic style of the source speech. The final synthesized audio is closer to the source prosodic speech in terms of emotion, tone, and rhythm, enhancing expressiveness and making the listening experience more natural and vivid.
[0050] Based on any of the above embodiments, step S10 specifically includes: Step S11: Based on the speech vector encoder, the source prosodic speech is encoded into source speech features, and the source speech features are mapped to the prosodic feature space through the prosodic mapper to obtain prosodic mapping features. Step S12: Based on the speech vector encoder, the speech of the target speaker is encoded into target speech features, and the target speech features are mapped to the timbre feature space through the timbre mapper to obtain timbre mapping features; Step S13: Based on the prosody timbre decoupling module, the prosody mapping feature and the timbre mapping feature are decoupled respectively to obtain the decoupled prosody feature and the decoupled timbre feature.
[0051] It should be noted that, in order to achieve flexible adjustment of speech attributes and effectively support single-person prosodic transfer tasks and cross-person prosodic transfer tasks, this embodiment of the invention designs a pluggable prosodic and timbre control module. In single-person prosodic transfer scenarios, since the speaker's timbre remains consistent, this module can be selectively removed (or not removed) to improve the system's inference efficiency; while in cross-person prosodic transfer tasks, this module can be enabled to independently model and jointly control prosodic factors and timbre features, thereby improving the naturalness and consistency of the transferred speech.
[0052] Figure 3 This is a schematic diagram of the pluggable rhythm and timbre control module provided by the present invention, as shown below. Figure 3 As shown, in the prosodic branch, the source prosodic speech is first encoded into source speech features using a pre-trained speech vector encoder. This speech vector encoder is a powerful acoustic feature extraction network. It is typically pre-trained on massive, diverse speech data, enabling it to convert the original continuous speech waveform or spectral signal into a high-level, discrete vector representation. While this vector representation (i.e., the source speech features) has reduced dimensionality, it still contains a mixture of information such as prosody, timbre, speech content, and background noise. This encoder can be based on a convolutional neural network, Transformer, or a combination thereof, such as the encoder part in SoundStream, EnCodec, or VQ-VAE. Then, the source speech features are mapped to the prosodic feature space using a prosodic mapper, resulting in prosodic mapped features. The purpose of this step is to perform preliminary attribute separation on the mixed speech features, projecting them into a more targeted prosodic feature space.
[0053] Similarly, in the timbre branch, the target speaker's speech is first encoded into target speech features using a pre-trained speech vector encoder. This speech feature is also a high-level, discrete vector representation; although its dimensionality is reduced, it still contains a mixture of information such as prosody, timbre, speech content, and background noise. Then, a timbre mapper maps the target speech features to a timbre feature space, obtaining timbre-mapped features. The purpose of this step is to perform preliminary attribute separation on the mixed speech features, projecting them into a more targeted timbre feature space.
[0054] Understandably, a prosodic mapper is a network module specifically designed to process prosodic information. It receives source speech features extracted from source prosodic speech and transforms them into a prosodic feature space specifically designed to represent prosody. This process can be understood as refining the source speech features, amplifying the prosodic-related parts and suppressing parts irrelevant to timbre, content, etc., thus obtaining prosodic mapper features. Similarly, a timbre mapper is a network module specifically designed to process timbre information. It receives target speech features extracted from the target speaker's speech and transforms them into a timbre feature space designed to represent timbre, thus obtaining timbre mapper features. Structurally, prosodic and timbre mappers can be multilayer perceptrons, recurrent neural networks, or small Transformers, etc. Introducing these two mappers not only prepares the ground for subsequent fine decoupling but also enhances the system's modularity and scalability.
[0055] Finally, the prosodic and timbre decoupling module, which is pre-trained, decouples the prosodic mapping features and timbre mapping features respectively, thus obtaining the decoupled prosodic features and decoupled timbre features. Here, the prosodic and timbre decoupling module is the core of feature separation. It receives the prosodic and timbre mapping features obtained in the previous step, which have been preliminarily separated, and uses a specific training strategy (such as contrastive learning) to force the output feature vectors to be orthogonal or uncorrelated in terms of information. The output of this module is the decoupled prosodic features and decoupled timbre features required by this invention. The former retains pure prosodic information, and the latter retains pure timbre information, thus providing high-quality input for the leak-free prosodic transfer in step S20.
[0056] This invention provides a clear and modular feature decoupling architecture. Through a three-step process of general encoding, attribute mapping, and fine decoupling, it is possible to more systematically and robustly separate pure prosodic and timbre representations from complex speech signals. This structured design not only improves the decoupling effect but also makes the functions of each module clear, easy to optimize and train independently, thereby enhancing the controllability and stability of the entire prosodic transfer method.
[0057] Based on any of the above embodiments, in step S11, mapping the source speech features to the prosodic feature space using a prosodic mapper to obtain prosodic mapping features specifically includes: Step S111: Based on the text encoder, the source speech text is encoded to obtain the text features of the source speech text, and the source speech text corresponds to the source prosodic speech in content. Step S112: Based on the prosodic encoder, the source speech features and the text features of the source speech text are encoded to obtain the initial prosodic features; Step S113: Based on the prosodic mapper, the initial prosodic features are mapped to the prosodic feature space to obtain the prosodic mapping features.
[0058] It should be noted that in practice, prosodic expression is often closely related to semantic content (for example, the intonation of an interrogative sentence differs from that of a declarative sentence). Extracting prosodicity solely from acoustic signals may lead to prosodic drift due to a lack of semantic context, meaning the extracted prosodicity may not match the text content. To address this issue, embodiments of the present invention introduce textual information corresponding to the source prosodic speech.
[0059] Specifically, the source speech text is first encoded using a text encoder to obtain its text features. Here, source speech text refers to text that completely corresponds in content to the input source prosodic speech. For example, if the source prosodic speech is a recording of "The weather is really nice today," then the source speech text is the text "The weather is really nice today." A text encoder can then convert this text into a text feature vector containing rich semantic information.
[0060] Then, a prosodic encoder encodes the source speech features and the text features of the source speech text to obtain initial prosodic features. This prosodic encoder is a fusion module that effectively fuses the source speech features from the acoustic signal and the text features from the text content. Through this fusion, the model can learn the association between acoustic representations (such as pitch variations and duration distribution) and semantic information (such as words, syntactic structures, and punctuation). For example, the model can learn that a question mark at the end of a sentence usually corresponds to a rise in intonation. The fused output is a more accurate initial prosodic feature that incorporates semantic context. This prosodic encoder can be implemented using a cross-attention mechanism, where features from one modality serve as queries and features from the other modality serve as keys and values, thereby enabling information interaction and alignment.
[0061] Finally, the prosodic mapper maps the initial prosodic features to the prosodic feature space, yielding the prosodic mapped features. At this point, the input to the prosodic mapper is no longer simply the source speech features, but rather the initial prosodic features enhanced with textual information. Because the initial prosodic features have already incorporated semantic information, the prosodic mapper can map them to the prosodic feature space more accurately, resulting in higher-quality prosodic mapped features that better represent the true speaker's intent.
[0062] In this embodiment of the invention, the accuracy and robustness of prosodic feature extraction are enhanced by introducing textual information corresponding to the source prosodic speech. Combining acoustic and semantic information enables the model to understand the linguistic motivations behind the prosody, effectively avoiding prosodic drift caused by a lack of context. This makes the final extracted decoupled prosodic features not only acoustically similar to the source speech but also semantically more appropriate and reasonable, thus providing crucial support for generating more natural and expressive target audio.
[0063] Based on any of the above embodiments, the training steps of the speech vector encoder include: Step S100: The large-scale raw audio data is processed using an automated annotation method to obtain a pre-trained corpus that includes prosodic information and speaker information.
[0064] It should be noted that a powerful speech vector encoder with good generalization ability is the cornerstone of the entire technical solution of this invention. It is responsible for transforming the ever-changing raw speech signals into unified digital features that can be processed by downstream models. In order to train such an encoder, this embodiment of the invention proposes a scheme for pre-training based on large-scale raw audio data.
[0065] Specifically, before training, large-scale raw audio data can be processed using automated annotation methods to obtain a pre-training corpus that includes prosodic and speaker information. This corpus can not only serve as training data for the speech vector encoder, but also provide training data for subsequent training and fine-tuning of the prosodic timbre decoupling module, autoregressive generative model, etc.
[0066] Understandably, large-scale raw audio data refers to massive amounts of audio or video data collected from publicly available online sources (such as films, audiobooks, public speeches, online courses, podcasts, etc.) using automated tools like web scraping. This data covers multiple fields, multiple scenarios, and involves tens of thousands of speakers. This data is typically unprocessed and unstructured.
[0067] Since manual annotation of massive amounts of data is extremely costly and impractical, this invention employs an automated processing flow to annotate these raw audio files. This flow automatically identifies speech content, segments and aligns it, and extracts key prosodic information (such as pitch, duration, and energy) and speaker information (such as speaker ID). After automated annotation, a structured database, namely a pre-training corpus, is formed. Each data entry in the corpus contains an audio segment, its corresponding text, phoneme-level temporal alignment information, and extracted prosodic and speaker feature vectors.
[0068] Furthermore, step S100 specifically includes: Step S101: Perform speech recognition on the original audio data to generate audio-text data pairs; Step S102: Align the audio and text data pairs to obtain phoneme-level aligned data. Step S103: Extract prosodic features and speaker features from the phoneme-level aligned data to form the pre-trained corpus.
[0069] Figure 4 This is a schematic diagram of the pre-training corpus construction process provided by the present invention, as follows: Figure 4 As shown, the raw audio data can first undergo speech recognition to generate audio-text data pairs. The goal of this step is to match each valid audio segment with its corresponding text content. Specifically, the massive amount of raw audio data collected is first preprocessed, including but not limited to audio format conversion, resampling, noise reduction, silent segment removal, and deduplication. Then, it is determined whether the audio contains text. If original paired text exists (such as in videos with subtitles), it is used directly. If not, an Automatic Speech Recognition (ASR) engine is used for transcription. To improve the accuracy and reliability of the transcribed text, this embodiment of the invention preferably adopts a multi-engine cross-validation strategy, that is, the same audio segment is sent to multiple different ASR engines (such as Whisper, Paraformer, etc.) for recognition. Only when the recognition results of multiple engines are highly consistent is the audio and its transcribed text adopted, forming a reliable audio-text data pair.
[0070] Next, the text alignment module is used to align the audio and text data pairs, resulting in phoneme-level aligned data. The goal of this step is to establish a precise correspondence between each basic phoneme in the text and the audio signal on the time axis, which is crucial for subsequent extraction of prosodic features such as duration. The output of this step is phoneme-level aligned data, which records the specific timestamp of each phoneme in the audio.
[0071] After obtaining precise phoneme-level alignment, various fine-grained acoustic features can be easily extracted. Specifically, based on the phoneme-level aligned data, prosodic features such as duration, fundamental frequency F0, energy, and rhythm can be obtained through a prosodic feature extraction module. Similarly, speaker features such as speaker ID, spectral envelope, formants, and Mel-spectral coefficients can be obtained through a speaker feature extraction module. It should be understood that the structures of the prosodic feature extraction module and the speaker feature extraction module include, but are not limited to, multi-scale convolutions, attention-based LSTMs, and Transformer Encoders.
[0072] Finally, the original audio path, preprocessed audio data, recognized text, phoneme-level alignment information, extracted prosodic feature vectors, and speaker feature vectors are integrated to form a complete record, which is then stored in the database. Repeating this process constitutes the final pre-training corpus.
[0073] Step S200: Based on the pre-trained corpus, the speech vector encoder is trained to convert continuous speech signals into discrete speech features.
[0074] Specifically, once a large-scale pre-trained corpus is available, the speech vector encoder can be trained. Its core objective is to learn a compression mapping that encodes high-dimensional, continuous speech signals (such as waveforms) into low-dimensional, discrete speech features (also known as speech units or tokens). This discretized representation significantly reduces the complexity of subsequent modeling and makes it possible to leverage powerful models from the field of natural language processing.
[0075] Figure 5 This is a training diagram of the speech vector encoder provided by the present invention, as shown below. Figure 5 As shown, the preferred embodiment of this invention employs an end-to-end architecture of "Encoder + Quantizer + Decoder," such as SoundStream, EnCodec, or VQ-VAE. During training, the original speech (i.e., audio from the pre-training corpus) is input into the Encoder (typically a multi-layer network such as convolutional or Transformer networks) to generate high-order continuous speech sequence features. These features are then fed into a vector quantizer (such as Vector Quantizer or Residual Vector Quantizer), which projects the continuous features onto a fixed codebook to obtain its corresponding discrete vector representation. The Decoder then attempts to reconstruct the original audio signal based on this discrete vector. The entire model is optimized end-to-end by minimizing the difference between the reconstructed audio and the original audio (i.e., the reconstruction loss). By training on massive and diverse data, the Encoder learns to capture the most essential and representative acoustic structures of speech while ignoring irrelevant variables such as noise, thus generating discrete speech features with good generalization and representational capabilities.
[0076] Understandably, after the discretized vector representation is passed as input to the decoder, the decoder's output can be used to perform downstream tasks such as speech reconstruction, speech recognition, and timbre recognition. The specific structure of the decoder can be flexibly configured according to the task objectives. Common designs include neural network architectures symmetrical to the encoder, such as transposed convolutional networks or Transformer decoding modules. During the inference phase, it is only necessary to obtain the discretized vector representation of the original speech in the quantizer.
[0077] The method provided in this invention constructs a large-scale corpus through automated annotation, successfully addressing the problems of data scarcity and insufficient diversity encountered when training high-performance speech models. This method not only significantly reduces the cost of data preparation but also makes it possible to train a general and powerful speech vector encoder. This fully pre-trained encoder provides a high-quality feature extraction foundation for all subsequent downstream tasks such as prosody / timbre decoupling and speech generation.
[0078] Based on any of the above embodiments, the prosodic timbre decoupling module is trained using a contrastive learning strategy, and the training steps of the prosodic timbre decoupling module include: Construct positive and negative sample pairs for prosodic features and positive and negative sample pairs for timbre features; The prosodic feature positive and negative sample pairs and the timbre feature positive and negative sample pairs are input into the initial decoupling module to obtain the decoupling result output by the initial decoupling module. Based on the decoupling result, the prosodic contrast loss, timbre contrast loss and cross-contrast loss are calculated. The cross-contrast loss is used to suppress the correlation between the decoupled prosodic features and the decoupled timbre features in the decoupling result. Based on the prosodic contrast loss, the timbre contrast loss, and the cross-contrast loss, a combined loss is determined, and the parameters of the initial decoupling module are updated and iterated according to the combined loss to obtain the prosodic timbre decoupling module.
[0079] Specifically, to achieve effective separation of prosody and timbre features, this invention preferably employs a contrastive learning strategy to train the prosody-timbre decoupling module. First, it is necessary to construct positive and negative sample pairs for prosody features and timbre features; this is the foundation of contrastive learning. For prosody features, the model needs to be informed which speech segments have similar prosody (forming positive sample pairs) and which are dissimilar (forming negative sample pairs). Similarly, for timbre features, it is also necessary to define which speech segments have the same timbre (forming positive sample pairs) and which are different (forming negative sample pairs). The specific method for constructing the sample pairs will be described in detail in subsequent embodiments.
[0080] During training, samples can be selected from positive and negative pairs of prosodic features and positive and negative pairs of timbre features, and input into the initial decoupling module. This module outputs the decoupled prosodic features and decoupled timbre features extracted from the input samples. It should be understood that the initial decoupling module refers to the prosodic-timbre decoupling module that has not yet been trained or is in its initial state. It can be composed of components such as multi-layer convolutional networks, recurrent neural networks, or Transformers. The parameters (weights and biases) inside the initial decoupling module are randomized or initialized by a pre-trained model. At this point, this module does not yet have the ability to effectively separate prosodic and timbre.
[0081] Based on the decoupled features output by the module and pre-constructed sample pairs, three loss functions can be calculated: prosodic contrast loss, timbre contrast loss, and cross-contrast loss. The prosodic contrast loss function aims to optimize the prosodic feature space. For the decoupled prosodic features of an anchor sample, this loss will make its representation closer to the decoupled prosodic features of positive samples (samples with similar prosodicities), while making it more distant from the decoupled prosodic features of all negative samples (samples with different prosodicities).
[0082] For example, one sample can be taken from the pair of positive and negative prosodic features as an anchor sample. Then, this anchor sample, along with its corresponding positive and negative prosodic feature samples, are input into the module to obtain their respective decoupled prosodic features. Next, the distance (or similarity) between the anchor sample and the decoupled prosodic features of the positive sample is calculated, as is the distance between the anchor sample and the decoupled prosodic features of the negative sample. The prosodic contrast loss brings the decoupled prosodic features of the anchor sample and its positive sample closer together in the feature space, while keeping them further apart.
[0083] Similar to prosodic contrast loss, timbre contrast loss optimizes the timbre feature space by operating on pairs of positive and negative timbre features. It encourages timbre feature representations from the same speaker (positive sample pairs) to move closer together, while timbre feature representations from different speakers (negative sample pairs) move further apart. It should be understood that both prosodic contrast loss and timbre contrast loss can be implemented using typical contrastive learning loss functions such as InfoNCE loss or Triplet Loss.
[0084] Cross-contrast loss is crucial for achieving high-quality decoupling. Simply optimizing each feature space does not guarantee that prosodic features are free of timbre information, and vice versa. Cross-contrast loss aims to suppress the correlation between decoupled prosodic and timbre features. One approach is to treat the prosodic and timbre features of the same speech segment as a pair of negative samples. By minimizing the mutual information between them or maximizing their contrast loss, the model is forced to learn to encode irrelevant information in both feature spaces, thereby making the prosodic and timbre representations orthogonal in an information-theoretical sense, achieving deep decoupling.
[0085] By weighted summing the three loss functions mentioned above, we can obtain a final combined loss, which can be expressed as follows: in, Indicates loss of rhythmic contrast. Indicates loss of timbre contrast. Indicates cross-comparison loss. , , It is a hyperparameter used to balance the importance of different loss terms.
[0086] Subsequently, the gradient of the combination loss relative to the network parameters of the prosodic timbre decoupling module (and related prosodic and timbre mappers) is calculated using the backpropagation algorithm, and these parameters are updated using an optimizer (such as Adam, SGD, etc.) based on this gradient. This process is iterated through a large number of sample pairs until the model converges, thus obtaining the trained prosodic timbre decoupling module.
[0087] In this embodiment of the invention, by employing three carefully designed loss functions, the model not only learns discriminative prosodic and timbre features within their respective domains, but more importantly, by introducing cross-contrast loss, it effectively solves the problems of feature entanglement and information leakage. This training method provides a solid foundation for obtaining pure and orthogonal prosodic and timbre representations.
[0088] Based on any of the above embodiments, the construction of prosodic feature positive and negative sample pairs and timbre feature positive and negative sample pairs includes: Based on speech segments with the same prosodic style corresponding to different speakers, construct positive sample pairs of prosodic features; Based on speech segments with different prosodic styles corresponding to the same speaker, construct negative sample pairs of prosodic features; Based on speech segments with different text contents corresponding to the same speaker, construct positive sample pairs of timbre features; Based on the speech segments corresponding to different speakers, negative sample pairs of timbre features are constructed.
[0089] Specifically, in order to achieve effective decoupling of prosody and timbre, this embodiment of the invention constructs a large number of positive and negative prosodic feature and timbre feature sample pairs based on the pre-processed audio data in the pre-training corpus when training the prosody and timbre decoupling module.
[0090] For positive prosodic feature pairs, the construction principle is that the timbre differs but the prosody is the same. To obtain such data pairs, they can be constructed based on speech segments with the same prosodic style corresponding to different speakers. Specifically, based on the prosodic features and speaker features pre-extracted from the corpus, speech segments with different speaker IDs but similar prosodic features can be selected and paired into positive sample pairs. Alternatively, speech conversion technology can be used for artificial synthesis. For example, a speech segment can be selected as an anchor point, and then a trained speech conversion model can be used to convert its timbre to that of another speaker in the corpus, while maintaining its original prosody (intonation, rhythm, etc.). In this way, the original speech segment and the speech segment after speech conversion constitute an ideal pair of positive prosodic feature pairs. This method can effectively train the model to ignore timbre differences and focus on capturing consistent prosodic patterns across speakers.
[0091] For prosodic feature negative sample pairs, the construction principle is that they have the same timbre but different prosodic patterns. Specifically, they can be constructed based on speech segments of the same speaker with different prosodic styles. For example, two speech segments read by the same speaker from the corpus but with different emotions (such as happiness and sadness) or different word classes (such as declarative sentences and interrogative sentences) can constitute a pair of prosodic feature negative sample pairs.
[0092] For positive timbre feature pairs, the construction principle is that prosody can differ, but timbre must be the same. Specifically, this can be achieved by constructing speech segments from different text content corresponding to the same speaker. To further enhance the model's robustness, data augmentation can be performed. This involves applying slight perturbations to a segment of speech from the same speaker using technical means, such as adjusting speech rate, changing local stress rhythm, or applying minor pitch perturbations, generating a series of samples with slightly different prosody but unchanged timbre identity. Positive timbre feature pairs can be formed between the original samples and these augmented samples. This trains the model to be less sensitive to prosodic changes, extracting more stable timbre identity features.
[0093] The principle for constructing negative timbre feature pairs is based on the difference in timbre. This can be directly constructed from speech segments corresponding to different speakers. By arbitrarily selecting speech segments from two different speakers in the corpus, they can form a pair of negative timbre feature pairs.
[0094] Based on any of the above embodiments, step S20 specifically includes: Based on the text features of the target text, the speech features of the target speaker, the decoupled prosodic features, and the decoupled timbre features, a target speech vector sequence is generated through an autoregressive generation model. The autoregressive generation model is used to predict the speech vector of the next frame in an autoregressive manner. The autoregressive generation model is fine-tuned based on a labeled speech dataset.
[0095] Specifically, Figure 6 This is a schematic diagram of the input and output of the autoregressive generative model provided by the present invention, as shown below. Figure 6 As shown, by inputting multimodal inputs (including text features, speech features, decoupled prosodic features, decoupled timbre features, etc.) into the autoregressive generative model, the target speech vector sequence output by the model can be obtained. Here, the autoregressive generative model preferably adopts a large-scale generative model similar to the Transformer architecture in the field of natural language processing. The core feature of this type of model is autoregression, that is, when generating the next element in the sequence, it takes all previously generated elements and all control conditions as input. In the speech synthesis task, this means that when the model predicts the speech vector of frame t, it will rely on all control information and the speech vectors of the previous t-1 frames that have been generated. This mechanism enables the model to capture the complex long-range temporal dependencies in the speech signal and generate coherent and natural speech.
[0096] Understandably, the model's multimodal input includes text features of the target text, speech features of the target speaker's voice, decoupled prosodic features, and decoupled timbre features. Text features provide the semantic information of the text to be synthesized. Speech features, acting as a global identifier, guide the model in generating the voice of a specific speaker; this feature can be a discrete vector representation. Decoupled prosodic features serve as style control signals, guiding the model to generate speech with a specific speaking style. Decoupled timbre features act as finer-grained timbre control signals, ensuring the accurate reproduction of timbre details.
[0097] During inference (i.e., synthesis), the model fuses the four features mentioned above (e.g., through concatenation or cross-attention mechanisms), and then begins to autoregressively generate the target speech vector sequence frame by frame (or token by token) until a special marker representing the end of speech is generated. It should be understood that each vector in the target speech vector sequence is a speech vector carrying semantic, prosodic, timbre, and other information.
[0098] Furthermore, the autoregressive generative model can be fine-tuned by using manually labeled speech data to refine the pre-trained autoregressive model. Here, the manually labeled speech dataset is a relatively small but extremely high-quality dataset. Based on a pre-built pre-training corpus, data samples with clear speech, natural pronunciation, and accurate text alignment are selected as the labeled objects. Professional annotators then meticulously manually label these selected data samples, with label dimensions including, but not limited to, speech rate (fast, medium, slow), tone (statement, question, exclamation), and emotion (happy, sad, neutral). This dataset provides high-quality supervision signals for the model to learn precise style control capabilities.
[0099] An autoregressive model that has been pre-trained on massive amounts of unlabeled or pseudo-labeled data is further trained and fine-tuned on the aforementioned artificially labeled dataset. During the fine-tuning phase, the model's input consists of text features, speaker features, and features corresponding to the prosody / timbre labels in the labeled data; the model's learning objective is to predict the speech vector sequence corresponding to the real speech in the labeled data.
[0100] Based on any of the above embodiments, step S30 specifically includes: Step S31: Based on the acoustic decoder, the target speech vector sequence is converted into an intermediate acoustic representation, wherein the acoustic decoder converts the target speech vector sequence based on the decoupled prosodic features and the decoupled timbre features. Step S32: Based on the neural vocoder, the intermediate acoustic representation is synthesized into the target audio.
[0101] Specifically, Figure 7 This is a schematic diagram of acoustic modeling and waveform synthesis provided by the present invention, as shown below. Figure 7 As shown, the specific process of synthesizing the target audio includes the following two stages: First, using an acoustic decoder (i.e., Figure 7 The causal-acoustic model shown converts the target speech vector sequence into an intermediate acoustic representation of the speech (such as a Mel spectrum). Here, the function of the acoustic decoder is to convert the highly abstract target speech vector sequence generated by the autoregressive model in the previous step into an intermediate acoustic representation that is more easily interpreted in the physical world. In this embodiment of the invention, the Mel spectrum is preferably used as this intermediate representation. The acoustic decoder can employ a lightweight causal model architecture such as Transformer or Conformer to ensure efficient inference speed.
[0102] To further enhance the temporal consistency and the naturalness and stability of the prosodic style of the generated speech, this embodiment of the invention guides the decoding process of the acoustic decoder to be influenced by the decoupled prosodic features and decoupled timbre features. This means that at each step of converting the speech vector into a Mel spectrum, the model references these two global prosodic style and timbre control vectors. This conditional decoding method effectively prevents style drift or timbre degradation during the decoding process, ensuring that the final generated Mel spectrum strictly adheres to the preset prosodic and timbre instructions.
[0103] After obtaining a high-quality Mel spectrum, the final step is to convert it into a final, continuous audio waveform. Traditional methods (such as the Griffin-Lim algorithm) often produce poor sound quality with a noticeable electronic feel. This invention employs a high-fidelity neural vocoder to accomplish this task. For example, a generative adversarial network-based vocoder, such as HiFi-GAN or BigVGAN, can be used. These vocoders are specially trained to learn a complex nonlinear mapping from the Mel spectrum to a high-fidelity audio waveform. They are capable of generating detailed, natural, and smooth high-quality audio.
[0104] Through the above two-stage processing, the system finally outputs high-quality target audio that is consistent with the target text content, has the prosodic style of the source speech and the timbre of the target speaker.
[0105] Based on any of the above embodiments, and addressing the shortcomings of traditional prosodic transfer schemes in terms of model robustness, controllability, and transfer granularity—such as timbre leakage in cross-person prosodic transfer tasks, coarse-grained prosodic transfer, and inability to effectively distinguish emotions—and considering the scarcity of publicly available high-quality datasets, existing datasets generally suffer from uneven distribution of prosodic styles and a lack of fine-grained prosodic labels, this invention provides a prosodic transfer method based on a pre-trained speech model. This method aims to leverage the strong representation and modeling capabilities of large-scale speech models, combined with automated prosodic analysis tools, to construct large-scale pseudo-labeled speech data for pre-training, and then fine-tuning it using a small number of high-quality manually labeled samples, thereby achieving speaker-independent prosodic representation learning and controllable transfer. The method specifically includes the following steps: Step 1: Collection of speech prosody transfer data.
[0106] The key to speech prosody transfer lies in establishing a mapping relationship between the same text content under different prosodic styles. To achieve this goal, an ideal dataset should contain a large number of the same text content performed by different speakers in diverse prosodic styles, and have a wide range of prosodic variations, including different emotional states, speech rate differences, and tone categories.
[0107] However, existing data resources are insufficient to meet these requirements. On the one hand, natural language corpora lack multi-prosodic expression samples for uniform text, resulting in a scarcity of training data for prosodic modeling. On the other hand, current publicly available speech datasets are inadequate in terms of prosodic style diversity, annotation accuracy, and speech quality, making them difficult to directly use for training high-performance prosodic transfer models.
[0108] To address this, this invention proposes a two-stage data construction and model training method. First, a large-scale automatically prosodic-annotated speech dataset is constructed, and a basic prosodic perception model is trained based on this dataset to capture common prosodic variation features. Then, a high-quality dataset with manual annotation is introduced to fine-tune and optimize the model, improving its generation accuracy and style consistency in prosodic transfer tasks, thereby achieving an effective balance between data scale and annotation cost.
[0109] Step 11: Automated annotation of massive amounts of voice data.
[0110] In the first phase, data crawling tools are used to collect audio and video data from various fields online, including daily communication, speeches, education, healthcare, and film and television works, to obtain raw audio datasets. Then, as... Figure 4 As shown, standardized preprocessing is performed on the collected data, including noise reduction, deduplication, and audio processing. Differential processing is applied based on the data format: if original paired text exists, the text alignment module is used directly to align the text and audio; if no original paired text exists, a multi-ASR speech recognition engine is first used for recognition and cross-validation to obtain highly reliable audio-text data pairs before text and audio alignment. Subsequently, phoneme prediction, prosody, and speaker feature extraction are performed on the aligned text and audio, ultimately resulting in a pre-trained corpus for multi-person speech prosody transfer.
[0111] Step 12: Manual fine-tuning of the voice data.
[0112] In the second stage, to achieve the fine-grained control objective in the multi-speaker speech prosody transfer task, this embodiment of the invention further refines the prosodic features, explicitly dividing them into multiple sub-dimensions such as speech rate, tone, and rhythm. For each dimension, corresponding annotation definitions and execution standards are formulated to ensure operability and consistency in the data annotation process.
[0113] Based on the corpus built in the first stage, this stage prioritizes selecting data samples with clear speech, natural pronunciation, and accurate text alignment as the fine-calibration objects. In addition, the data distribution is reasonably configured according to the usage ratio of various scenarios in actual applications to improve the adaptability and generalization ability of model training.
[0114] During the annotation process, a dual-person independent annotation method was adopted, followed by a review by a third person for verification. This process not only effectively reduced subjective errors but also ensured the stability and consistency of the annotation results. Statistics show that the final high-quality annotated corpus exhibited over 90% consistency in the prosodic dimension, providing reliable supervised data support for the subsequent training of pluggable prosodic and timbre control modules.
[0115] Step 2: Based on a pre-trained large model speech prosody transfer architecture.
[0116] Step 21, Discretization and pre-training of speech vectors.
[0117] Since the original speech signal is continuous and high-dimensional, direct modeling is not only computationally expensive but also difficult to effectively extract reusable prosodic information. Therefore, this invention uses a speech encoder to represent and extract the speech signal, transforming it into finite discrete speech units. Generally, speech discretization techniques mainly include two implementation paths: one is an end-to-end "encoder + vector quantizer" architecture; the other is a "feature extraction + clustering discretization" method based on self-supervised learning. This invention preferably adopts the former. Figure 5 As shown, after the original speech signal enters the encoder, it obtains high-order continuous speech sequence features. These features are then fed into a vector quantizer, which projects the continuous vectors onto a fixed codebook table and obtains their corresponding discrete token sequences. During the training phase, the continuous features output by the encoder are discretized into a series of index values by the vector quantizer, with each index corresponding to a vector representation in the codebook. These vectors are then passed as input to the decoder to perform downstream tasks such as speech reconstruction, speech recognition, and timbre recognition. During the inference phase, it is only necessary to obtain the discrete tokens corresponding to the original speech in the quantizer, look up their corresponding vectors in the table, and input them into the decoder to reconstruct the speech signal.
[0118] To enable the quantized speech feature vectors to effectively encode implicit attribute information such as prosody and timbre in speech, and to possess good versatility and transferability, this invention uses a large-scale speech dataset with automated annotation from multiple domains and multiple speakers (i.e., the pre-training corpus obtained in step 1) during the training phase. This ensures that the model can learn diverse speech expressions during training, enabling it to learn representative prosody and timbre representations. Consequently, in subsequent speech synthesis and prosody transfer tasks, it can effectively restore and reconstruct the target style.
[0119] Step 22: The rhythm and timbre control module can be plugged in and removed.
[0120] To achieve flexible adjustment of speech attributes and effectively support prosodic transfer tasks for single and cross-speaker speakers, this invention provides a pluggable prosodic and timbre control module that is trained independently. Figure 3 As shown, the inputs to the prosody and timbre control module include source speech text, source prosodic speech, and target speaker speech. The source speech text and source prosodic speech correspond in content. The text encoder is responsible for extracting the embedded representation containing prosodic semantic information, providing necessary contextual information for subsequent prosodic modeling, thereby effectively preventing prosodic drift. Since the speech vector encoder trained in step 21 can extract speech information containing high-level features such as prosody, timbre, and speaking habits from speech, the speech vector encoder extracts features from the source prosodic speech and the target speaker speech respectively, and maps them to their respective feature spaces through the prosody mapper and timbre mapper, preparing for subsequent prosody and timbre separation. In addition, the design of the mapper also facilitates the introduction of more style variables, such as emotional expression and prosodic intensity.
[0121] During training, the prosody-timbre decoupling module employs a contrastive learning strategy, aiming to output decoupled prosodic features and decoupled timbre features separately. The decoupled prosodic features are highly similar to the features output by the prosodic mapper, and the decoupled timbre features are highly similar to the features output by the timbre mapper. However, the decoupled prosodic features do not contain any timbre-related information, and the decoupled timbre features do not contain any prosodic-related information, thus achieving effective decoupling of prosody and timbre.
[0122] To achieve effective decoupling of prosody and timbre, this invention constructs a large number of positive and negative sample pairs for prosody features and timbre features during the training of the decoupling module. The specific construction of these sample pairs can be referred to in the above embodiments and will not be repeated here. For the loss function, this invention introduces contrastive learning loss functions for both the prosody and timbre branches to enhance the discriminative power of their respective feature spaces. Typical contrastive losses such as InfoNCE or Triplet Loss can be used for the prosody and timbre branches to narrow the distance between positive sample pairs and increase the discriminative power between negative sample pairs. Furthermore, to further avoid information overlap between the two feature spaces, a cross-contrast constraint (i.e., cross-contrast loss) is specifically introduced. By suppressing the correlation between prosody features and timbre features, this promotes orthogonality in the representation of the two types of features, thereby improving the decoupling effect and ensuring that each feature possesses good independence and expressive purity in downstream tasks.
[0123] Step 23, fine-tuning of the large speech model.
[0124] Similar to text modeling approaches in natural language processing, this invention employs a large-scale generative model with an autoregressive structure to model and predict the vectorized representation of speech. For example... Figure 6 As shown, the model's input includes the text features of the target text, the discrete identity code of the target speaker (i.e., speech features), and the decoupled prosodic and timbre features. During training, the model predicts the speech vector features of the next frame in an autoregressive manner, with the training input being the precisely labeled data from step 1. This aims to model the dynamic relationship between the current frame and the context frames, and to fuse a joint representation of text semantics, speaker features, prosodic style, and timbre. During inference, the system can flexibly incorporate different prosodic or timbre modules to achieve fine-grained control over the style dimension of the generated speech.
[0125] Unlike traditional end-to-end acoustic modeling methods, this structure explicitly separates semantic, stylistic, and individual feature factors in the speech generation process, making speech modeling more controllable and pluggable.
[0126] Step 24, acoustic modeling and waveform synthesis.
[0127] After completing high-level speech vector modeling, it is necessary to map it into real, playable audio information. In this stage, embodiments of the present invention employ a two-stage acoustic modeling strategy: first, an acoustic decoder converts the target speech vector sequence into an intermediate representation of speech (such as a Mel spectrum), and then a high-fidelity neural vocoder reconstructs it into waveform audio. For example... Figure 7 As shown, in the acoustic modeling stage, a causal acoustic model can be used. This model receives the speech vector generated by the autoregressive model in the previous step, combines it with its corresponding decoupled prosodic and timbre features to further enhance its temporal consistency and prosodic naturalness, and outputs the intermediate result, the Mel spectrum. Subsequently, the generated Mel spectrum is input into a neural vocoder. The vocoder, acting as a spectrum-to-waveform mapper, is responsible for high-fidelity reconstruction of the spectrum, generating a detailed, natural, and fluent speech waveform.
[0128] Through the above process, not only is the naturalness of the speech output and the consistency of prosodic expression ensured, but the generalization ability and controllability of the model in multi-speaker and multi-style speech generation tasks are also improved.
[0129] The prosody transfer device provided by the present invention is described below. The prosody transfer device described below can be referred to in correspondence with the prosody transfer method described above.
[0130] Based on any of the above embodiments Figure 8 This is a schematic diagram of the prosody transfer device provided by the present invention, as shown below. Figure 8 As shown, the device includes: The acquisition unit 810 is used to acquire decoupled prosodic features based on source prosodic speech and decoupled timbre features based on target speaker speech, wherein the decoupled prosodic features characterize the prosody of the source prosodic speech and the decoupled timbre features characterize the timbre of the target speaker speech. The generation unit 820 is used to generate a target speech vector sequence based on the text features of the target text, the speech features of the target speaker's speech, the decoupled prosodic features, and the decoupled timbre features; Synthesis unit 830 is used to synthesize target audio based on the target speech vector sequence.
[0131] The apparatus provided in this invention effectively alleviates the problem of feature mixing by acquiring decoupled prosodic features based on the source prosodic speech and decoupled timbre features based on the target speaker's speech. This is because the decoupled prosodic features can more purely represent the prosody of the source speech and basically do not contain its timbre information; similarly, the decoupled timbre features can also more purely represent the timbre of the target speaker. In the synthesis process, combining a pure decoupled prosodic feature with a pure decoupled timbre feature can ensure that the final synthesized audio can highly maintain the purity of the target speaker's own timbre, thereby effectively improving the quality of cross-speaker prosodic transfer. Furthermore, by acquiring decoupled prosodic features, this invention can separate and retain more complete and refined prosodic information from complex speech signals. Since the interference of timbre components is avoided, the model can focus on modeling the prosody itself, thereby learning richer and more detailed prosodic representations. When generating the target speech vector sequence, this high-quality decoupled prosodic feature is used as an input condition, enabling the model to more accurately reproduce the prosodic style of the source speech. The final synthesized audio is closer to the source prosodic speech in terms of emotion, tone, and rhythm, enhancing expressiveness and making the listening experience more natural and vivid.
[0132] Based on any of the above embodiments, the acquisition unit 810 includes: The encoding subunit is used to encode the source prosodic speech and the target speaker speech into source speech features and target speech features, respectively, based on a speech vector encoder; The prosody mapping subunit is used to map the source speech features to the prosody feature space based on the prosody mapper to obtain the prosody mapping features; The timbre mapping subunit is used to map the target speech features to the timbre feature space based on the timbre mapper to obtain timbre mapping features; The decoupling subunit is used to decouple the prosodic mapping feature and the timbre mapping feature based on the prosodic timbre decoupling module to obtain the decoupled prosodic feature and the decoupled timbre feature.
[0133] Based on any of the above embodiments, the prosody mapping subunit is specifically used for: Based on a text encoder, the source speech text is encoded to obtain the text features of the source speech text, and the source speech text corresponds to the source prosodic speech in terms of content. Based on the prosodic encoder, the source speech features and the text features of the source speech text are encoded to obtain the initial prosodic features; Based on the prosodic mapper, the initial prosodic features are mapped to the prosodic feature space to obtain the prosodic mapping features.
[0134] Based on any of the above embodiments, the prosodic timbre decoupling module is trained using a contrastive learning strategy. Accordingly, the device further includes a decoupling module training unit, which specifically includes: The sample construction subunit is used to construct positive and negative sample pairs of prosodic features and positive and negative sample pairs of timbre features; The loss calculation subunit is used to calculate prosodic contrast loss, timbre contrast loss and cross-contrast loss based on the prosodic feature positive and negative sample pairs and the timbre feature positive and negative sample pairs. The cross-contrast loss is used to suppress the correlation between the decoupled prosodic features and the decoupled timbre features. The parameter update subunit is used to determine the combination loss based on the prosodic contrast loss, the timbre contrast loss, and the cross-contrast loss, and to update the parameters of the prosodic timbre decoupling module according to the combination loss.
[0135] Based on any of the above embodiments, the sample construction subunit is specifically used for: Based on speech segments with the same prosodic style corresponding to different speakers, construct positive sample pairs of prosodic features; Based on speech segments with different prosodic styles corresponding to the same speaker, construct negative sample pairs of prosodic features; Based on speech segments with different text contents corresponding to the same speaker, construct positive sample pairs of timbre features; Based on the speech segments corresponding to different speakers, negative sample pairs of timbre features are constructed.
[0136] Based on any of the above embodiments, the device further includes a speech encoder training unit, wherein the speech encoder training unit specifically includes: The corpus construction subunit is used to process large-scale raw audio data through automated annotation methods to obtain a pre-trained corpus that includes prosodic information and speaker information. The training subunit is used to train the speech vector encoder based on the pre-trained corpus to convert continuous speech signals into discrete speech features.
[0137] Based on any of the above embodiments, the corpus construction subunit is specifically used for: The original audio data is subjected to speech recognition to generate audio-text data pairs; Align the audio and text data pairs to obtain phoneme-level aligned data; Prosodic features and speaker features are extracted from the phoneme-level aligned data to form the pre-trained corpus.
[0138] Based on any of the above embodiments, the generation unit 820 is specifically used for: Based on the text features of the target text, the speech features of the target speaker, the decoupled prosodic features, and the decoupled timbre features, a target speech vector sequence is generated through an autoregressive generation model. The autoregressive generation model is used to predict the speech vector of the next frame in an autoregressive manner. The autoregressive generation model is fine-tuned based on a labeled speech dataset.
[0139] Based on any of the above embodiments, the synthesis unit 830 is specifically used for: Based on an acoustic decoder, the target speech vector sequence is converted into an intermediate acoustic representation, wherein the acoustic decoder converts the target speech vector sequence based on the decoupled prosodic features and the decoupled timbre features; The intermediate acoustic representation is synthesized into the target audio based on a neural vocoder.
[0140] Figure 9 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 9 As shown, the electronic device may include: a processor 910, a communication interface 920, a memory 930, and a communication bus 940, wherein the processor 910, the communication interface 920, and the memory 930 communicate with each other through the communication bus 940. The processor 910 can call logical instructions in the memory 930 to execute a prosodic transfer method, which includes: acquiring decoupled prosodic features based on source prosodic speech and decoupled timbre features based on target speaker speech, wherein the decoupled prosodic features characterize the prosody of the source prosodic speech and the decoupled timbre features characterize the timbre of the target speaker speech; generating a target speech vector sequence based on text features of the target text, speech features of the target speaker speech, the decoupled prosodic features, and the decoupled timbre features; and synthesizing target audio based on the target speech vector sequence.
[0141] Furthermore, the logical instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to related technologies, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0142] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the prosodic transfer method provided by the above methods. The method includes: acquiring decoupled prosodic features based on source prosodic speech and decoupled timbre features based on target speaker speech, wherein the decoupled prosodic features characterize the prosody of the source prosodic speech and the decoupled timbre features characterize the timbre of the target speaker speech; generating a target speech vector sequence based on text features of target text, speech features of the target speaker speech, the decoupled prosodic features, and the decoupled timbre features; and synthesizing target audio based on the target speech vector sequence.
[0143] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the prosodic transfer method provided by the above methods. The method includes: acquiring decoupled prosodic features based on source prosodic speech and decoupled timbre features based on target speaker speech, wherein the decoupled prosodic features characterize the prosody of the source prosodic speech and the decoupled timbre features characterize the timbre of the target speaker speech; generating a target speech vector sequence based on text features of target text, speech features of the target speaker speech, the decoupled prosodic features, and the decoupled timbre features; and synthesizing target audio based on the target speech vector sequence.
[0144] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0145] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A prosodic transfer method, characterized in that, include: Obtain decoupled prosodic features based on source prosodic speech and decoupled timbre features based on target speaker speech, wherein the decoupled prosodic features characterize the prosody of the source prosodic speech and the decoupled timbre features characterize the timbre of the target speaker speech; Based on the text features of the target text, the speech features of the target speaker, the decoupled prosodic features, and the decoupled timbre features, a target speech vector sequence is generated; Based on the target speech vector sequence, the target audio is synthesized.
2. The prosodic transfer method according to claim 1, characterized in that, The acquisition of decoupled prosodic features based on the source prosodic speech and decoupled timbre features based on the target speaker's speech includes: Based on the speech vector encoder, the source prosodic speech is encoded into source speech features, and the source speech features are mapped to the prosodic feature space through the prosodic mapper to obtain prosodic mapping features. Based on the speech vector encoder, the speech of the target speaker is encoded into target speech features, and the target speech features are mapped to the timbre feature space through the timbre mapper to obtain timbre mapping features; Based on the prosodic timbre decoupling module, the prosodic mapping feature and the timbre mapping feature are decoupled respectively to obtain the decoupled prosodic feature and the decoupled timbre feature.
3. The prosodic transfer method according to claim 2, characterized in that, The step of mapping the source speech features to the prosodic feature space using a prosodic mapper to obtain prosodic mapped features includes: Based on a text encoder, the source speech text is encoded to obtain the text features of the source speech text, and the source speech text corresponds to the source prosodic speech in terms of content. Based on the prosodic encoder, the source speech features and the text features of the source speech text are encoded to obtain the initial prosodic features; Based on the prosodic mapper, the initial prosodic features are mapped to the prosodic feature space to obtain the prosodic mapping features.
4. The prosodic transfer method according to claim 2, characterized in that, The prosodic timbre decoupling module is trained using a contrastive learning strategy. The training steps of the prosodic timbre decoupling module include: Construct positive and negative sample pairs for prosodic features and positive and negative sample pairs for timbre features; The prosodic feature positive and negative sample pairs and the timbre feature positive and negative sample pairs are input into the initial decoupling module to obtain the decoupling result output by the initial decoupling module. Based on the decoupling result, the prosodic contrast loss, timbre contrast loss and cross-contrast loss are calculated. The cross-contrast loss is used to suppress the correlation between the decoupled prosodic features and the decoupled timbre features in the decoupling result. Based on the prosodic contrast loss, the timbre contrast loss, and the cross-contrast loss, a combined loss is determined, and the parameters of the initial decoupling module are updated and iterated according to the combined loss to obtain the prosodic timbre decoupling module.
5. The prosodic transfer method according to claim 2, characterized in that, The training steps of the speech vector encoder include: By using automated annotation methods, large-scale raw audio data is processed to obtain a pre-trained corpus that includes prosodic information and speaker information; Based on the pre-trained corpus, the speech vector encoder is trained to convert continuous speech signals into discrete speech features.
6. The prosodic transfer method according to claim 5, characterized in that, The process involves using automated annotation methods to process large-scale raw audio data, resulting in a pre-trained corpus that includes prosodic information and speaker information, comprising: The original audio data is subjected to speech recognition to generate audio-text data pairs; Align the audio and text data pairs to obtain phoneme-level aligned data; Prosodic features and speaker features are extracted from the phoneme-level aligned data to form the pre-trained corpus.
7. The prosodic transfer method according to any one of claims 1 to 6, characterized in that, The generation of a target speech vector sequence based on the text features of the target text, the speech features of the target speaker, the decoupled prosodic features, and the decoupled timbre features includes: Based on the text features of the target text, the speech features of the target speaker, the decoupled prosodic features, and the decoupled timbre features, a target speech vector sequence is generated through an autoregressive generation model. The autoregressive generation model is used to predict the speech vector of the next frame in an autoregressive manner. The autoregressive generation model is fine-tuned based on a labeled speech dataset.
8. The prosodic transfer method according to any one of claims 1 to 6, characterized in that, The process of synthesizing target audio based on the target speech vector sequence includes: Based on an acoustic decoder, the target speech vector sequence is converted into an intermediate acoustic representation, wherein the acoustic decoder converts the target speech vector sequence based on the decoupled prosodic features and the decoupled timbre features; The intermediate acoustic representation is synthesized into the target audio based on a neural vocoder.
9. A rhythm transfer device, characterized in that, include: The acquisition unit is used to acquire decoupled prosodic features based on source prosodic speech and decoupled timbre features based on target speaker speech, wherein the decoupled prosodic features characterize the prosody of the source prosodic speech and the decoupled timbre features characterize the timbre of the target speaker speech. The generation unit is used to generate a target speech vector sequence based on the text features of the target text, the speech features of the target speaker's speech, the decoupled prosodic features, and the decoupled timbre features; A synthesis unit is used to synthesize target audio based on the target speech vector sequence.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the prosody transfer method as described in any one of claims 1 to 8.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the prosody transfer method as described in any one of claims 1 to 8.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the prosody transfer method as described in any one of claims 1 to 8.
Citation Information
Cited By
Multi-round interaction emotion speech synthesis method and system based on context self-adaption
CN122024704A
A speech synthesis and smoothing method based on multi-lingual migration
CN122454954A
A speech synthesis and smoothing method based on multi-lingual migration
CN122454954B