Audio synthesis method, device and equipment and readable storage medium
By extracting and dynamically combining emotional and timbre features, the problem of inconsistent emotional perception in cross-linguistic speech synthesis is solved, achieving natural emotional expression and tone changes, which is suitable for multilingual emotion synthesis applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-03
AI Technical Summary
Existing speech synthesis algorithms struggle to maintain consistency in emotional expression and intonation rhythm across language scenarios, failing to effectively depict the natural emotional fluctuations that occur in real human speech over time.
By acquiring emotion reference audio and timbre reference audio, extracting features using emotion encoder and timbre encoder, and combining attention weighting and cross-linguistic emotion mapping function, dynamic combination and fusion of emotion and timbre are achieved to generate audio with natural emotional expression.
It achieves improved cross-language emotional consistency and speech naturalness, maintaining similar emotional intensity and expression across different languages, and is suitable for scenarios such as virtual humans, digital anchors, voice assistants, and emotional companion robots.
Smart Images

Figure CN121789631A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio processing technology, and in particular to an audio synthesis method, apparatus, device, and readable storage medium. Background Technology
[0002] Traditional speech synthesis algorithms primarily focus on intelligibility and naturalness, aiming for clear and fluent pronunciation. However, with the rise of interactive voice, virtual humans, intelligent customer service, emotional companion robots, and cross-language speech synthesis applications, user expectations have evolved from simply being able to speak to understanding what is being spoken. In multilingual scenarios, speech must not only accurately convey content but also maintain consistent emotional tone and rhythm across different languages. Existing methods using emotion categories and fixed intensities can generate emotionally charged speech. However, these methods struggle to capture the natural emotional fluctuations that occur in human speech over time.
[0003] In conclusion, how to effectively solve problems such as making the generated audio have more natural emotional expression is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0004] The purpose of this application is to provide an audio synthesis method, apparatus, device, and readable storage medium for generating audio with natural emotional expression.
[0005] To solve the above-mentioned technical problems, this application provides the following technical solution:
[0006] A cross-language audio synthesis method, comprising:
[0007] Obtain emotional reference audio, timbre reference audio, and target text; wherein the emotional reference audio and the timbre reference audio correspond to different language types.
[0008] Emotional features are extracted from the emotional reference audio using an emotion encoder, timbre features are extracted from the timbre reference audio using a timbre encoder, and semantic features are extracted from the target text.
[0009] The emotional features and the semantic features are fused to obtain the fused features;
[0010] The fusion feature and the timbre feature are synthesized to obtain a synthesized audio; the synthesized audio has an emotion corresponding to the emotion reference audio, a timbre corresponding to the timbre reference audio, and semantics corresponding to the target text.
[0011] Preferably, the fusion feature and the timbre feature are synthesized to obtain a synthesized audio, including:
[0012] By using attention weighting, the emotion vector in the fusion feature and the timbre vector in the timbre feature are dynamically combined along the time dimension so that the pitch and intonation of the synthesized audio change dynamically with the emotion.
[0013] Preferably, attention-weighted combination is used to dynamically combine the emotion vector in the fusion feature and the timbre vector in the timbre feature along the time dimension, including:
[0014] An emotion recognizer is used to identify the emotion in the synthesized audio to obtain the emotional state.
[0015] If the emotional state is insufficient in terms of emotional expression relative to the emotional reference audio, then during dynamic combination, the fundamental frequency and energy adjustment are increased;
[0016] If the emotional state is too strong relative to the emotional reference audio, then the intonation change will be reduced during dynamic combination.
[0017] Preferably, attention-weighted combination is used to dynamically combine the emotion vector in the fusion feature and the timbre vector in the timbre feature along the time dimension, including:
[0018] Attention weights are determined using a cross-lingual sentiment mapping function; the sentiment mapping function is established based on a cross-lingual sentiment discriminator and is used to pull the sentiment embedding distribution of different languages;
[0019] Based on the attention weights, the emotion vector in the fusion features and the timbre vector in the timbre features are dynamically combined along the time dimension.
[0020] Preferably, the emotion encoder is used to extract emotion features from the emotion reference audio, including:
[0021] The energy envelope, fundamental frequency curve, and prosodic rhythm features in the emotion reference audio are extracted using the convolutional and attention layers in the emotion encoder.
[0022] The energy envelope, the fundamental frequency curve, and the rhythmic features are converted into an emotion vector.
[0023] The emotion vector is determined as the emotion feature.
[0024] Preferably, during training, a language recognizer is used as an adversarial discriminator so that the emotion encoder removes language features by minimizing language recognition loss.
[0025] Preferably, extracting timbre features from the timbre reference audio using a timbre encoder includes:
[0026] In the process of extracting the timbre features from the timbre reference audio using the timbre encoder, the emotion interference suppression layer in the timbre encoder is used to prevent the timbre features from being interfered with by emotion changes by introducing emotion invariance constraints.
[0027] An audio synthesis device, comprising:
[0028] The input acquisition module is used to acquire emotion reference audio, timbre reference audio, and target text; wherein the emotion reference audio and the timbre reference audio correspond to different language types.
[0029] The feature extraction module is used to extract emotional features from the emotional reference audio using an emotion encoder, extract timbre features from the timbre reference audio using a timbre encoder, and extract semantic features from the target text.
[0030] The feature fusion module is used to fuse the emotion features and the semantic features to obtain fused features;
[0031] An audio generation module is used to synthesize the fusion feature and the timbre feature to obtain a synthesized audio; the synthesized audio has an emotion corresponding to the emotion reference audio, a timbre corresponding to the timbre reference audio, and semantics corresponding to the target text.
[0032] An electronic device, comprising:
[0033] Memory, used to store computer programs;
[0034] A processor is used to implement the steps of the above-described audio synthesis method when executing the computer program.
[0035] A readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described audio synthesis method.
[0036] The method provided in this application provides for obtaining an emotion reference audio, a timbre reference audio, and a target text; wherein the emotion reference audio and the timbre reference audio correspond to different language types; an emotion encoder is used to extract emotion features from the emotion reference audio, a timbre encoder is used to extract timbre features from the timbre reference audio, and semantic features are extracted from the target text; the emotion features and semantic features are fused to obtain fused features; the fused features and timbre features are synthesized to obtain synthesized audio; the synthesized audio has the emotion corresponding to the emotion reference audio, the timbre corresponding to the timbre reference audio, and the semantics corresponding to the target text.
[0037] In this application, during audio synthesis, the emotion reference audio and timbre reference audio can each correspond to different language types, and the semantic content is derived separately from the target text. Specifically, an emotion encoder can be used to extract emotion features from the emotion reference audio, a timbre encoder can be used to extract timbre features from the timbre reference audio, and semantic features can be extracted from the target text. Then, firstly, the emotion features and semantic features are fused to obtain fused features. Finally, the fused features are synthesized with the timbre features, thus obtaining synthesized audio with the emotion corresponding to the emotion reference audio, the timbre corresponding to the timbre reference audio, and the semantics corresponding to the target text.
[0038] Compared to synthesizing audio directly based on clear labels, the emotions in this application are derived from emotion reference audio, providing a more natural emotional guidance for the synthesized audio. Furthermore, the emotion reference audio and timbre reference audio correspond to different languages, enabling cross-language emotion and speech synthesis. This results in more natural synthesized audio when facing cross-language audio synthesis.
[0039] Specifically, in scenarios involving virtual humans and digital anchors: it enables control over tone changes, making performances more closely resemble the emotional rhythm of real people; in scenarios involving voice assistants and emotional companion robots, it can automatically adjust tone based on the content of the conversation, such as offering comfort, encouragement, and empathy; in scenarios involving audiobooks and film dubbing, it can control the emotional fluctuations of characters, enhancing immersion; and in scenarios involving cross-language emotional transfer, it can preserve similar emotional intensity and expression methods across different languages.
[0040] Accordingly, embodiments of this application also provide an audio synthesis apparatus, device, and readable storage medium corresponding to the above-described audio synthesis method, which have the aforementioned technical effects, and will not be repeated here. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a flowchart illustrating an audio synthesis method as described in this application.
[0043] Figure 2 This is a schematic diagram of audio generation input and output in an embodiment of this application;
[0044] Figure 3 This is a schematic diagram of feature fusion in an embodiment of this application;
[0045] Figure 4 This is a schematic diagram of an emotion feedback closed-loop mechanism in an embodiment of this application;
[0046] Figure 5 This is a schematic diagram of the structure of an audio synthesis device according to an embodiment of this application;
[0047] Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application;
[0048] Figure 7 This is a schematic diagram of the specific structure of an electronic device in an embodiment of this application. Detailed Implementation
[0049] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0050] Please refer to Figure 1 , Figure 1 This is a flowchart of an audio synthesis method according to an embodiment of this application, which can be applied to an audio synthesis device or model. The method includes the following steps:
[0051] S101. Obtain the emotion reference audio, the timbre reference audio, and the target text.
[0052] Among them, the emotion reference audio and the timbre reference audio correspond to different language types.
[0053] The target text can be obtained by semantic extraction from emotion reference audio or timbre reference audio. Alternatively, for easier subsequent synthesis, the extracted semantics can be directly saved as semantic features for later fusion. In other words, in practical applications, fusion can be performed using only two audio samples.
[0054] Emotional reference audio refers to speech with emotional representation, while timbre reference audio refers to speech with a human voice. Emotional reference audio and timbre reference audio can correspond to different language types; for example, emotional reference audio could be spoken in Chinese, while timbre reference audio could be spoken in English. The content does not need to be identical.
[0055] The emotion reference audio is the one from which the emotion of the subsequently generated synthesized audio is expected to correspond to the emotion in the emotion reference audio. The timbre reference audio is the one from which the timbre of the subsequently generated synthesized audio is expected to correspond to the timbre reference audio. The target text is the semantic content that is expected to be expressed in the subsequently generated synthesized audio.
[0056] In other words, the emotion reference audio, the timbre reference audio, and the target text can respectively guide the final synthesized audio from different dimensions of emotion, timbre, and semantics.
[0057] S102. Extract emotional features from emotional reference audio using an emotion encoder, extract timbre features from timbre reference audio using a timbre encoder, and extract semantic features from the target text.
[0058] To facilitate the fusion of multi-dimensional features, in this embodiment, features can be extracted from the emotion reference audio, the timbre reference audio, and the target text separately to obtain the corresponding multi-dimensional features.
[0059] Specifically, an emotion encoder can be used to extract emotional features from an emotion reference audio, a timbre encoder can be used to extract timbre features from a timbre reference audio, and semantic features can be extracted from the target text.
[0060] In this way, one can reference emotions solely from emotion reference audio, timbre from timbre reference audio, and semantics solely from the target text.
[0061] In one specific embodiment of this application, an emotion encoder is used to extract emotion features from an emotion reference audio, including:
[0062] By utilizing the convolutional and attention layers in the emotion encoder, we extract the energy envelope, fundamental frequency curve, and prosodic features from the emotion reference audio.
[0063] Convert energy envelope, fundamental frequency curve, and rhythmic features into emotion vectors;
[0064] The emotion vector is defined as the emotion feature.
[0065] During training, a language recognizer is used as an adversarial discriminator so that the emotion encoder can remove language features by minimizing the language recognition loss.
[0066] In other words, the emotion encoder is responsible for extracting language-independent emotional features from the emotion reference audio. It can employ a structure combining multi-layer convolution and attention mechanisms to extract the energy envelope, fundamental frequency curve, and prosodic rhythm features from speech, and encode them into a low-dimensional emotion vector representation.
[0067] To avoid the semantic or phonetic patterns of a specific language being included in the emotion vector, a language self-adversarial alignment mechanism is introduced: during the training process of the model, a language recognizer is introduced as an adversarial discriminator, enabling the emotion encoder to erase language features by minimizing the language recognition loss, thereby ensuring that the extracted emotion representation truly has cross-language universality. This process enables the Chinese word "高兴" and the English word "happy" to map to a similar region in the emotion latent space, serving as the basis for interchangeable emotion expressions.
[0068] In a specific implementation manner of the present application, a timbre encoder is used to extract timbre features from a timbre reference audio, including:
[0069] During the process of using the timbre encoder to extract timbre features from the timbre reference audio, an emotion interference suppression layer in the timbre encoder is used to prevent the timbre features from being interfered by emotion changes by introducing an emotion invariance constraint.
[0070] That is, the encoder is responsible for extracting stable speaker features, namely timbre, from the timbre reference audio. Different from the emotion encoder, the goal of the timbre encoder is to capture the inherent vocal features of the speaker. The emotion interference suppression layer can be used to prevent the timbre features from being interfered by emotion changes by introducing an emotion invariance constraint. During training, the system uses an adversarial loss function to ensure that the timbre embedding remains consistent under different emotion conditions, thereby achieving a complete decoupling of emotion and timbre. This mechanism ensures that during the subsequent cross-language emotion transfer process, even if the emotion signal changes, the timbre features can remain stable, and it always sounds like the same person.
[0071] Extract semantic features from the target text, that is, convert the target text into a semantic vector.
[0072] S103. Fuse the emotion features and the semantic features to obtain fused features.
[0073] After the feature extraction is completed, in order to make the synthesized audio more natural and fluent, in this embodiment, the emotion features and the semantic features are first fused to obtain fused features.
[0074] That is, map the emotion embedding of the source language (emotion reference audio) to the vocal feature space of the target language (the language system of the target text) to ensure the consistency of the emotion expression in the final output speech.
[0075] S104. Synthesize the fused features and the timbre features to obtain a synthesized audio.
[0076] Among them, the synthesized audio has the emotion corresponding to the emotion reference audio, the timbre corresponding to the timbre reference audio, and the semantics corresponding to the target text.
[0077] After obtaining the fusion features, the audio still lacks a distinctive timbre. Synthetic audio can be produced by combining the fusion features with the timbre features.
[0078] In this way, the synthesized audio acquires emotion, timbre, and semantics.
[0079] In one specific embodiment of this application, the fusion features and timbre features are synthesized to obtain synthesized audio, including:
[0080] By using attention weighting, the emotion vector in the fusion feature and the timbre vector in the timbre feature are dynamically combined along the time dimension so that the pitch and intonation of the synthesized audio change dynamically with the emotion.
[0081] In other words, at the fusion layer, an attention-weighted mechanism can be used to dynamically combine emotion vectors and timbre vectors along the time dimension, allowing the model to adjust synthesis parameters at each time step based on the current emotion intensity and timbre stability. Specifically, when the emotion curve rises, the system enhances the fluctuation range of energy and pitch; when the emotion is calm or declining, it reduces the amplitude of intonation changes to maintain overall coherence. In this way, frame-level emotion controllability is achieved during audio generation, thereby simulating the natural phenomenon of emotions gradually changing with semantic progression in human speech.
[0082] In one specific embodiment of this application, attention weighting is used to dynamically combine the emotion vector in the fusion features and the timbre vector in the timbre features along the time dimension, including:
[0083] An emotion recognizer is used to identify emotions in synthesized audio to obtain emotional states.
[0084] If the emotional state is insufficient relative to the emotional reference audio, then the fundamental frequency and energy adjustment are increased during dynamic combination.
[0085] If the emotional state is too strong relative to the emotional reference audio, then the intonation changes will be reduced during dynamic combination.
[0086] To further enhance stability and controllability, an emotion feedback closed-loop mechanism is added during the decoding stage. A lightweight emotion recognizer is simultaneously activated during synthesis to detect the emotional state of the currently output synthesized audio in real time and compare it with the target emotion (i.e., the emotion in the emotion reference audio). For example, if insufficient emotional expression is detected during synthesis, the model will appropriately increase the fundamental frequency and energy modulation; if excessive emotion is detected, the tone changes will automatically converge. This closed-loop control mechanism enables the system to possess self-awareness, allowing it not only to generate emotionally charged speech but also to continuously optimize the accuracy of emotional expression during the generation process.
[0087] In a specific implementation manner of the present application, by using attention weighting, the emotion vector in the fusion feature and the timbre vector in the timbre feature are dynamically combined in the time dimension, including:
[0088] Determine the attention weight by using a cross - language emotion mapping function; the emotion mapping function is established based on a cross - language emotion discriminator, and the emotion mapping function is used to拉近 the emotion embedding distributions of different languages;
[0089] Based on the attention weight, the emotion vector in the fusion feature and the timbre vector in the timbre feature are dynamically combined in the time dimension.
[0090] In this embodiment, multi - speaker emotion synthesis under cross - language transfer is also supported. The system constructs a unified emotion latent space using multilingual and multi - speaker corpora during the training phase, enabling samples in different languages to share the same emotion dimensions. For example, the system can use the "happy" voice sample of Chinese speaker A as a reference and transfer the same "happy" emotion to the timbre of English speaker B to generate the synthetic speech "Hello", thus achieving the separation and recombination of the three elements of "emotion from A, timbre from B, and language determined by the text". This mechanism provides a general technical framework for applications such as cross - language virtual humans, bilingual anchors, and multilingual emotion robots.
[0091] Specifically, a lightweight Transformer structure can be used as the backbone network to balance efficiency and expressiveness. The model can dynamically weight the importance of emotion features during the attention calculation process, thereby achieving emotion - driven context modeling. This design significantly improves the intonation naturalness and rhythm of the synthetic speech, making the generated speech not only accurate in content but also maintain consistent emotion perception across different languages. For example, when the input is "你好呀" and the emotion reference is "开心", whether the output is in Chinese or English, the system can maintain a brisk and bright tone rhythm.
[0092] Apply the method provided by the embodiment of the present application to obtain an emotion reference audio, a timbre reference audio, and a target text; wherein, the emotion reference audio and the timbre reference audio correspond to different language types respectively; use an emotion encoder to extract emotion features from the emotion reference audio, use a timbre encoder to extract timbre features from the timbre reference audio, and extract semantic features from the target text; fuse the emotion features and the semantic features to obtain a fusion feature; synthesize the fusion feature and the timbre feature to obtain a synthetic audio; the synthetic audio has the emotion corresponding to the emotion reference audio, the timbre corresponding to the timbre reference audio, and the semantics corresponding to the target text.
[0093] In this application, during audio synthesis, the emotion reference audio and timbre reference audio can each correspond to different language types, and the semantic content is derived separately from the target text. Specifically, an emotion encoder can be used to extract emotion features from the emotion reference audio, a timbre encoder can be used to extract timbre features from the timbre reference audio, and semantic features can be extracted from the target text. Then, firstly, the emotion features and semantic features are fused to obtain fused features. Finally, the fused features are synthesized with the timbre features, thus obtaining synthesized audio with the emotion corresponding to the emotion reference audio, the timbre corresponding to the timbre reference audio, and the semantics corresponding to the target text.
[0094] Compared to synthesizing audio directly based on clear labels, the emotions in this application are derived from emotion reference audio, providing a more natural emotional guidance for the synthesized audio. Furthermore, the emotion reference audio and timbre reference audio correspond to different languages, enabling cross-language emotion and speech synthesis. This results in more natural synthesized audio when facing cross-language audio synthesis.
[0095] Specifically, in scenarios involving virtual humans and digital anchors: it enables control over tone changes, making performances more closely resemble the emotional rhythm of real people; in scenarios involving voice assistants and emotional companion robots, it can automatically adjust tone based on the content of the conversation, such as offering comfort, encouragement, and empathy; in scenarios involving audiobooks and film dubbing, it can control the emotional fluctuations of characters, enhancing immersion; and in scenarios involving cross-language emotional transfer, it can preserve similar emotional intensity and expression methods across different languages.
[0096] To facilitate those skilled in the art to better understand and implement the audio synthesis method provided in the embodiments of this application, the audio synthesis method will be described in detail below with specific scenario examples.
[0097] As can be seen from the above, the audio synthesis method provided in this application embodiment allows the emotions in the synthesized audio to change naturally with the semantic rhythm, so that the emotions can be adjusted in real time and gradually evolved during the speech generation process, and maintain the same auditory experience and speaker consistency in different languages, so as to truly achieve a speech generation experience that is "emotional, warm, and consistent across languages".
[0098] Specifically, this application elevates "emotional intensity" from a static label to a function, introducing prior knowledge of prosody, semantics, and syntax at the text level, and steady-state constraints on speaker timbre at the acoustic level. By introducing language self-adversarial alignment and cross-lingual emotion mapping modules, it achieves unified emotional features across different languages, forming an editable, trackable, and feedback-enabled end-to-end control loop. Furthermore, it maintains stable naturalness, intelligibility, and speaker similarity under cross-linguistic and multi-speaker conditions.
[0099] The following example illustrates the application of this audio synthesis method in a video synthesis model.
[0100] Please refer to Figure 2 The video synthesis model has two inputs and one output. The inputs are the emotion reference audio from Speaker A and the timbre reference audio from Speaker B. After passing through the core EGE-TTS cross-language emotion transfer model, the target audio (i.e., the synthesized audio) is obtained at the output.
[0101] The input end receives reference speech signals from different speakers. Speaker A (e.g., a Chinese speaker) provides speech samples with specific emotions as emotional reference audio, along with a target text (e.g., "Hello"). Speaker B (e.g., an English speaker) provides speech samples including timbre features as timbre reference audio. After entering the EGE TTS model, the two signals are processed by the emotion encoder and timbre encoder respectively to extract multi-dimensional latent features. The model then fuses the two signals in a shared latent space, and the emotion features are transferred and remapped through a cross-language alignment module. Finally, the decoder generates output speech that simultaneously possesses "Speaker A's emotional expression" and "Speaker B's timbre features," i.e., "emotion = A, timbre = B," resulting in cross-language emotion-consistent speech.
[0102] By introducing a multi-level emotion control mechanism and a cross-lingual feature alignment module into the speech synthesis model, high-fidelity and high-consistency emotion transfer is achieved across different languages. The core idea of this invention is to extend "emotion intensity" from a static label to a time-series function, and to achieve dynamic and continuous emotion regulation throughout the speech generation process. Simultaneously, by decoupling the emotion representation space from the timbre feature space, the model can achieve "emotion and timbre separation" under cross-lingual conditions, thereby transferring emotional expression from one language to another while maintaining the speaker's timbre stability, achieving truly cross-lingual emotion-consistent synthesis.
[0103] Most related emotional TTS systems generate speech in the way of "emotional category + fixed intensity", and their emotional changes only stay at the level of macroscopic category differentiation (such as "happy", "sad", "angry"), unable to reflect the delicate fluctuations of emotions over time. In this invention, by modeling emotional signals as time functions, the system can control the emotional intensity curve with frame-level precision, achieving a continuous gradual change from "slightly happy" to "extremely happy", thus being more in line with the emotional rhythm when humans speak. At the same time, the model analyzes the prosody, semantics and syntactic structure of the input text, and combines the intensity dynamics in the emotional reference signal. For example, when the input text contains emotional features such as questions, exclamations or transitions, the system can automatically adjust the rising speed and peak moment of the emotional curve, making the output speech show a natural sense of emotional flow.
[0104] Please refer to Figure 3 , in terms of the internal structure of the model, the EGE TTS framework includes an emotion encoding module, a timbre encoding module, a cross-language emotion alignment module, an emotion-timbre fusion module and a decoder.
[0105] The emotion encoding module is responsible for extracting language-independent emotional features from the emotional reference audio of Speaker A. This module adopts a structure combining multi-layer convolution and attention mechanism to extract the energy envelope, fundamental frequency curve and prosody rhythm features in the emotional reference audio, and encodes them into a low-dimensional emotional vector representation. In order to avoid including specific language semantic or speech patterns in this emotional vector, a language self-adversarial alignment mechanism can be introduced: that is, a language recognizer is introduced as an adversarial discriminator during the training process of the model, so that the emotion encoder "erases" the language features by minimizing the language recognition loss, thus ensuring that the extracted emotional representation truly has cross-language universality. This process enables "happy" in Chinese and "happy" in English to map to similar regions in the emotional latent space, becoming the basis for interchangeable emotional expressions.
[0106] The timbre encoding module is responsible for extracting stable speaker features, that is, timbre features, from the reference audio of Speaker B. Different from the emotion encoder, the goal of the timbre encoder is to capture the inherent vocalization features of the speaker. An emotion interference suppression layer can be designed, and by introducing an emotion invariance constraint, the timbre features are prevented from being interfered by emotional changes. During training, the system uses an adversarial loss function to make the timbre embedding remain consistent under different emotional conditions, thus achieving a complete decoupling of emotion and timbre. This mechanism ensures that during the subsequent cross-language emotion transfer process, even if the emotional signal changes, the timbre features can remain stable and the listener always feels it is the same person.
[0107] After completing emotion and timbre encoding, the cross-lingual emotion alignment module begins its work. This module maps the emotion embeddings of the source language into the vocal feature space of the target language to ensure consistency in emotional expression in the output speech. To achieve this, an adversarial learning-based emotion alignment method is used: by introducing a cross-lingual emotion discriminator, the distributions of emotion embeddings from different languages are brought closer together, allowing the model to learn language-independent emotion projections in the latent space. Specifically, the system is trained using a large amount of cross-lingual data to minimize the distance between samples of the same emotion category in the high-dimensional space and maximize the distance between different emotion categories, thereby establishing a cross-lingual emotion mapping function. This process effectively solves the problem of inconsistent emotional expression between different languages due to differences in prosodic systems.
[0108] In the stage of fusing emotion and timbre features, a multi-layer adaptive fusion network can be introduced. The fusion layer uses an attention-weighted mechanism to dynamically combine the emotion vector and timbre vector along the time dimension, allowing the model to adjust the synthesis parameters at each time step based on the current emotion intensity and timbre stability. Specifically, when the emotion curve rises, the system enhances the fluctuation range of energy and pitch; when the emotion is flat or declining, it reduces the amplitude of intonation changes to maintain overall coherence. In this way, the system achieves frame-level emotion controllability during the generation process, thereby simulating the natural phenomenon of emotion gradually changing with semantic progression in human speech.
[0109] Please refer to Figure 4 To further enhance the model's stability and controllability, an emotion feedback closed-loop mechanism can be incorporated into the decoding stage. During synthesis, a lightweight emotion recognizer is simultaneously activated to detect the emotional state of the current output speech in real time and compare it with the target emotion. For example, if insufficient emotional expression is detected during synthesis, the model will appropriately increase the fundamental frequency and energy modulation; if excessive emotion is detected, the model will automatically converge the intonation changes. This closed-loop control mechanism enables the system to possess "self-awareness," allowing it not only to generate emotionally charged speech but also to continuously optimize the accuracy of emotional expression during the generation process.
[0110] Furthermore, this EGE TTS model supports multi-speaker emotion synthesis under cross-language transfer. During the training phase, a unified emotion latent space is constructed using multilingual, multi-speaker corpora, allowing samples from different languages to share the same emotional dimension. For example, the system can use a "happy" voice sample from Chinese speaker A as a reference, transferring the same "happy" emotion to the timbre of English speaker B to generate the synthesized speech "Hello," thus achieving the separation and recombination of the three elements: "emotion from A, timbre from B, and language determined by the text." This mechanism provides a general technical framework for applications such as cross-lingual virtual humans, bilingual anchors, and multilingual emotion robots.
[0111] In specific implementations, a lightweight Transformer structure can be adopted as the backbone network to balance efficiency and expressiveness. The model can dynamically weight the importance of emotional features during the attention calculation process, thereby achieving emotion-driven context modeling. This design significantly improves the intonation naturalness and rhythm of the synthesized speech, making the generated speech not only accurate in content but also maintain consistent emotional perception across different languages. For example, when the input is "你好呀" and the emotional reference is "开心", whether the output is in Chinese or English, the system can maintain a brisk and bright tone rhythm.
[0112] That is to say, the audio synthesis method provided by this application has the following technical effects when actually implemented.
[0113] Cross-language emotion decoupling and dynamic migration mechanism: By introducing a language self-adversarial alignment and cross-language emotion mapping module, the emotional features across different languages are unified, making "高兴 in Chinese" and "happy in English" equivalent in the latent space, thus achieving true cross-language emotion consistent expression.
[0114] Frame-level emotional intensity function modeling: Expand the emotional intensity from static labels to functions, enabling the system to dynamically adjust the emotional intensity during speech generation, significantly improving the naturalness and emotional fluency of the speech.
[0115] Adaptive fusion and emotion closed-loop feedback: Achieve fine-grained dynamic combination of emotion and timbre through a multi-layer attention fusion mechanism, and introduce an emotion recognition feedback loop, enabling the system to have self-awareness and self-correction capabilities, ensuring accurate and stable emotion expression of the generated speech.
[0116] Corresponding to the above method embodiment, the embodiment of this application also provides an audio synthesis device, and the audio synthesis device described below can be mutually referred to the audio synthesis method described above.
[0117] See Figure 5 As shown, the device includes the following modules:
[0118] An input acquisition module 101, configured to acquire an emotion reference audio, a timbre reference audio, and a target text; wherein, the emotion reference audio and the timbre reference audio respectively correspond to different language types;
[0119] A feature extraction module 102, configured to extract emotion features from the emotion reference audio by using an emotion encoder, extract timbre features from the timbre reference audio by using a timbre encoder, and extract semantic features from the target text;
[0120] A feature fusion module 103, configured to fuse the emotion features and the semantic features to obtain fused features;
[0121] The audio generation module 104 is used to synthesize the fusion features and timbre features to obtain the synthesized audio; the synthesized audio has the emotion corresponding to the emotion reference audio, the timbre corresponding to the timbre reference audio, and the semantics corresponding to the target text.
[0122] Using the apparatus provided in the embodiments of this application, an emotion reference audio, a timbre reference audio, and a target text are obtained; wherein the emotion reference audio and the timbre reference audio correspond to different language types; an emotion encoder is used to extract emotion features from the emotion reference audio, a timbre encoder is used to extract timbre features from the timbre reference audio, and semantic features are extracted from the target text; the emotion features and semantic features are fused to obtain fused features; the fused features and timbre features are synthesized to obtain synthesized audio; the synthesized audio has the emotion corresponding to the emotion reference audio, the timbre corresponding to the timbre reference audio, and the semantics corresponding to the target text.
[0123] In this application, during audio synthesis, the emotion reference audio and timbre reference audio can each correspond to different language types, and the semantic content is derived separately from the target text. Specifically, an emotion encoder can be used to extract emotion features from the emotion reference audio, a timbre encoder can be used to extract timbre features from the timbre reference audio, and semantic features can be extracted from the target text. Then, firstly, the emotion features and semantic features are fused to obtain fused features. Finally, the fused features are synthesized with the timbre features, thus obtaining synthesized audio with the emotion corresponding to the emotion reference audio, the timbre corresponding to the timbre reference audio, and the semantics corresponding to the target text.
[0124] Compared to synthesizing audio directly based on clear labels, the emotions in this application are derived from emotion reference audio, providing a more natural emotional guidance for the synthesized audio. Furthermore, the emotion reference audio and timbre reference audio correspond to different languages, enabling cross-language emotion and speech synthesis. This results in more natural synthesized audio when facing cross-language audio synthesis.
[0125] Specifically, in scenarios involving virtual humans and digital anchors: it enables control over tone changes, making performances more closely resemble the emotional rhythm of real people; in scenarios involving voice assistants and emotional companion robots, it can automatically adjust tone based on the content of the conversation, such as offering comfort, encouragement, and empathy; in scenarios involving audiobooks and film dubbing, it can control the emotional fluctuations of characters, enhancing immersion; and in scenarios involving cross-language emotional transfer, it can preserve similar emotional intensity and expression methods across different languages.
[0126] In one specific embodiment of this application, the audio generation module is specifically used to dynamically combine the emotion vector in the fusion feature and the timbre vector in the timbre feature according to the time dimension using attention weighting, so that the pitch and intonation of the synthesized audio change dynamically with the emotion.
[0127] In one specific embodiment of this application, the audio generation module is specifically used to perform emotion recognition on the synthesized audio using an emotion recognizer to obtain the emotion state;
[0128] If the emotional state is insufficient relative to the emotional reference audio, then the fundamental frequency and energy adjustment are increased during dynamic combination.
[0129] If the emotional state is too strong relative to the emotional reference audio, then the intonation changes will be reduced during dynamic combination.
[0130] In one specific embodiment of this application, the audio generation module is specifically used to determine attention weights using a cross-lingual emotion mapping function; the emotion mapping function is established based on a cross-lingual emotion discriminator and is used to pull together the emotion embedding distribution of different languages;
[0131] Based on attention weights, the emotion vector in the fusion features and the timbre vector in the timbre features are dynamically combined according to the time dimension.
[0132] In one specific embodiment of this application, the feature extraction module is specifically used to extract the energy envelope, fundamental frequency curve, and prosodic rhythm features from the emotional reference audio by utilizing the convolution and attention layers in the emotion encoder.
[0133] Convert energy envelope, fundamental frequency curve, and rhythmic features into emotion vectors;
[0134] The emotion vector is defined as the emotion feature.
[0135] In one specific embodiment of this application, during the training process, a language recognizer is used as an adversarial discriminator so that the emotion encoder removes language features by minimizing language recognition loss.
[0136] In one specific embodiment of this application, the feature extraction module is specifically used to prevent the timbre features from being interfered with by emotional changes by using the emotion interference suppression layer in the timbre encoder during the process of extracting timbre features from the timbre reference audio using the timbre encoder, by introducing emotion invariance constraints.
[0137] Corresponding to the above method embodiments, this application also provides an electronic device. The electronic device described below and the audio synthesis method described above can be referred to in correspondence.
[0138] See Figure 6 As shown, the electronic device includes:
[0139] Memory 332 is used to store computer programs;
[0140] The processor 322 is used to implement the steps of the audio synthesis method of the above method embodiment when executing a computer program.
[0141] For details, please refer to Figure 7 , Figure 7 This is a schematic diagram of the specific structure of an electronic device provided in this embodiment. The electronic device can vary significantly due to differences in configuration or performance. It may include one or more central processing units (CPUs) (e.g., one or more processors) and a memory 332. The memory 332 stores one or more computer programs 342 or data 344. The memory 332 can be temporary or permanent storage. The program stored in the memory 332 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the data processing device. Furthermore, the processor 322 may be configured to communicate with the memory 332 and execute the series of instruction operations stored in the memory 332 on the electronic device 301.
[0142] Electronic device 301 may also include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341.
[0143] The steps in the audio synthesis method described above can be implemented by the structure of an electronic device.
[0144] Corresponding to the above method embodiments, this application also provides a readable storage medium. The readable storage medium described below can be referred to in conjunction with the audio synthesis method described above.
[0145] A readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the audio synthesis method described in the above method embodiments.
[0146] The readable storage medium can specifically be a USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, or any other readable storage medium capable of storing program code.
[0147] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0148] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0149] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0150] Finally, it should be noted that in this document, relationships such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "include," "contain," or any other variations are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.
[0151] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An audio synthesis method, characterized in that, include: Obtain emotional reference audio, timbre reference audio, and target text; wherein the emotional reference audio and the timbre reference audio correspond to different language types. Emotional features are extracted from the emotional reference audio using an emotion encoder, timbre features are extracted from the timbre reference audio using a timbre encoder, and semantic features are extracted from the target text. The emotional features and the semantic features are fused to obtain the fused features; The fusion feature and the timbre feature are synthesized to obtain a synthesized audio; the synthesized audio has an emotion corresponding to the emotion reference audio, a timbre corresponding to the timbre reference audio, and semantics corresponding to the target text.
2. The method according to claim 1, characterized in that, The fusion feature and the timbre feature are combined to obtain a synthesized audio, including: By using attention weighting, the emotion vector in the fusion feature and the timbre vector in the timbre feature are dynamically combined along the time dimension so that the pitch and intonation of the synthesized audio change dynamically with the emotion.
3. The method according to claim 2, characterized in that, Using attention weighting, the emotion vector in the fusion feature and the timbre vector in the timbre feature are dynamically combined along the time dimension, including: An emotion recognizer is used to identify the emotion in the synthesized audio to obtain the emotional state. If the emotional state is insufficient in terms of emotional expression relative to the emotional reference audio, then during dynamic combination, the fundamental frequency and energy adjustment are increased; If the emotional state is too strong relative to the emotional reference audio, then the intonation change will be reduced during dynamic combination.
4. The method according to claim 2, characterized in that, Using attention weighting, the emotion vector in the fused features and the timbre vector in the timbre features are dynamically combined along the time dimension, including: Attention weights are determined using a cross-lingual sentiment mapping function; the sentiment mapping function is established based on a cross-lingual sentiment discriminator and is used to pull the sentiment embedding distribution of different languages; Based on the attention weights, the emotion vector in the fusion features and the timbre vector in the timbre features are dynamically combined along the time dimension.
5. The method according to claim 1, characterized in that, Extracting emotional features from the emotional reference audio using an emotion encoder includes: The energy envelope, fundamental frequency curve, and prosodic rhythm features in the emotion reference audio are extracted using the convolutional and attention layers in the emotion encoder. The energy envelope, the fundamental frequency curve, and the rhythmic features are converted into an emotion vector. The emotion vector is determined as the emotion feature.
6. The method according to claim 5, characterized in that, During training, a language recognizer is used as an adversarial discriminator so that the emotion encoder can remove language features by minimizing the language recognition loss.
7. The method according to claim 1, characterized in that, Extracting timbre features from the timbre reference audio using a timbre encoder includes: In the process of extracting the timbre features from the timbre reference audio using the timbre encoder, the emotion interference suppression layer in the timbre encoder is used to prevent the timbre features from being interfered with by emotion changes by introducing emotion invariance constraints.
8. An audio synthesis device, characterized in that, include: The input acquisition module is used to acquire emotion reference audio, timbre reference audio, and target text; wherein the emotion reference audio and the timbre reference audio correspond to different language types. The feature extraction module is used to extract emotional features from the emotional reference audio using an emotion encoder, extract timbre features from the timbre reference audio using a timbre encoder, and extract semantic features from the target text. The feature fusion module is used to fuse the emotion features and the semantic features to obtain fused features; An audio generation module is used to synthesize the fusion feature and the timbre feature to obtain a synthesized audio; the synthesized audio has an emotion corresponding to the emotion reference audio, a timbre corresponding to the timbre reference audio, and semantics corresponding to the target text.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the audio synthesis method as described in any one of claims 1 to 7 when executing the computer program.
10. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the audio synthesis method as described in any one of claims 1 to 7.