Emotional speech synthesis method, system, device and medium
By transferring the timbre of the target object and controlling its voice attributes, high-quality, multi-emotional long audio is generated, solving the problem of insufficient emotional expression in existing technologies. This improves the flexibility and subtlety of emotional speech synthesis, enhancing the realism and appeal of dubbed works.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN JIANSHI TECH CO LTD
- Filing Date
- 2026-03-25
- Publication Date
- 2026-05-15
AI Technical Summary
Existing dubbing technologies based on zero-sample voice cloning and TTS lack flexibility, subtlety, and adaptability in emotional expression, especially in cross-language dubbing, resulting in insufficient emotional expressiveness of synthesized speech and affecting the realism and appeal of dubbed works.
By acquiring audio samples of speech with various emotions and the target voice of the target object, timbre transfer processing is performed to generate long audio recordings of multiple emotions. Based on the voice attribute control parameters input by the user, speech synthesis is carried out to achieve high-quality, multi-type emotional speech synthesis.
It improves the flexibility, subtlety, and adaptability of synthesized speech in emotional expression, enhances emotional expressiveness, and improves the realism and appeal of dubbed works.
Smart Images

Figure CN122050355A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer-aided voice-over technology, and in particular to a method, system, device and medium for synthesizing emotional speech. Background Technology
[0002] Currently, AI-based cross-language dubbing for films, television dramas, and short dramas primarily employs a combination of zero-sample voice cloning and text-to-speech (TTS) techniques. The typical process involves extracting a certain length of speaker audio samples from the film or drama to extract and replicate the speaker's vocal characteristics; then, using the extracted vocal timbre and emotion, it is combined with the translated text to synthesize the target language speech, which is then synchronized to the video feed, thus completing the dubbing process.
[0003] Within this technical framework, the acquisition and reproduction of vocal features for characters in video content such as movies, TV dramas, and short dramas are relatively mature, achieving a high degree of vocal imitation similarity. However, the system still has significant shortcomings in emotional expression, mainly due to the uncertain sample duration of different emotions contained in the original audio, resulting in limited feature information and thus affecting the emotional expressiveness of the synthesized speech.
[0004] Specifically, although the system can reproduce the character's voice relatively well, when there are few audio samples that need to express emotions, the synthesized speech often sounds stiff, unnatural, or inconsistent with the character's actual psychological state. This problem is particularly prominent in cross-language dubbing, severely limiting the overall realism and appeal of the dubbed work.
[0005] Therefore, while existing dubbing technologies based on zero-sample voice cloning and TTS have achieved certain results in timbre imitation, they are significantly lacking in the flexibility, subtlety, and adaptability of emotional expression, and urgently need further improvement and enhancement. Summary of the Invention
[0006] To overcome the significant shortcomings of existing speech synthesis technologies in terms of the flexibility, subtlety, and adaptability of emotional expression, this application provides an emotional speech synthesis method, system, device, and medium.
[0007] Firstly, in order to solve the aforementioned technical problems, this application provides a method for synthesizing emotional speech, including: Acquire audio samples of speech with various emotions and the target voice of the target object; Based on the target vocal timbre, multiple speaking sample audios are processed by timbre transfer to obtain a long audio file with multiple target emotions. Obtain the voice attribute control parameters input by the user; Speech synthesis is performed based on speech attribute control parameters and target multi-emotion long audio to obtain target emotion speech.
[0008] Secondly, this application also provides an emotion-based speech synthesis system, comprising: The sample acquisition module is used to acquire audio samples of speech with various emotions and the target voice of the target object; The timbre transfer module is used to perform timbre transfer processing on multiple speaking sample audios based on the target voice to obtain a long audio file with multiple target emotions. The control parameter acquisition module is used to acquire the voice attribute control parameters input by the user. The speech synthesis module is used to synthesize speech based on speech attribute control parameters and target multi-emotion long audio to obtain the target emotional speech.
[0009] Thirdly, this application also provides a computing device, including a memory, a processor, and a program stored in the memory and running on the processor, wherein the processor executes the program to implement the steps of the emotional speech synthesis method described above.
[0010] Fourthly, this application also provides a computer-readable storage medium storing instructions that, when executed on a terminal device, cause the terminal device to perform the steps of an emotion-based speech synthesis method.
[0011] The beneficial effects of this application are: based on the target voice of the target object, the audio samples of speech with multiple emotions are processed by timbre transfer to obtain target multi-emotion long audio, and based on the voice attribute control parameters input by the user and the target multi-emotion long audio, speech is synthesized to obtain target emotion speech. In this way, rich emotional information can be intuitively understood through the target emotion speech, thereby improving the flexibility, subtlety and adaptability of synthesized speech in emotional expression. Attached Figure Description
[0012] Figure 1 This is a flowchart illustrating an exemplary embodiment of the present application of an emotional speech synthesis method; Figure 2 This is a schematic diagram of the timbre transfer process for a target multi-emotion long audio file in an exemplary embodiment of this application; Figure 3 This is a schematic diagram illustrating the structure of an emotion-based speech synthesis system, as shown in an exemplary embodiment of this application. Figure 4 This is a schematic diagram of the structure of a computer system of a computing device, illustrating an exemplary embodiment of this application. Detailed Implementation
[0013] The following embodiments are further explanations and supplements to this application and do not constitute any limitation on this application.
[0014] Currently, cross-language dubbing technology based on voice cloning still has significant shortcomings in achieving high-quality emotional expression, mainly in the following four aspects: Firstly, emotional feature extraction is limited by the length of the original audio: because the original audio of movies, TV dramas, and short dramas only contains emotional states corresponding to specific plot points, the diversity of the emotional feature sample library is insufficient. When it is necessary to generate emotional expressions with limited sample audio data from within the drama, the system struggles to output subtle and natural emotional changes, affecting the realism of the dubbing.
[0015] Secondly, the emotional expression lacks subtlety and dynamic variation: the emotional output of existing technologies is often fixed in a limited number of modes, making it difficult to achieve gradual changes in the intensity of emotions and dynamic flow within sentences. This results in the templated emotional expression of synthesized speech, lacking the subtle layers of real speech.
[0016] Thirdly, speech quality deteriorates under high-intensity emotions: When synthesizing high-intensity emotions such as intense anger and extreme excitement, speech distortion and reduced clarity are likely to occur, which restricts the application effect of the system in dramatic conflict scenarios.
[0017] Fourthly, there is a reliance on high-quality emotional data and a tendency towards homogenization in expression: existing methods rely on a large amount of accurately labeled emotional speech data, but in practical applications, the cost of acquiring such data is high, leading to a tendency for the system output to be "mediocre," lacking individual expressiveness and failing to reach the artistic level of professional dubbing.
[0018] To address the aforementioned issues, embodiments of this application provide a method, system, device, and medium for synthesizing emotional speech, which will be described in detail below.
[0019] The emotion-based speech synthesis method provided in this application can be specifically executed by a server. It should be noted that the server can be a standalone server, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. No limitation is imposed here.
[0020] This application aims to generate synthetic speech with rich emotional expression using a limited number of original speech samples of the target subject. The core idea is to transfer the target subject's timbre to high-quality emotional speech audio recorded by a professional voice actor, thus achieving high-quality, multi-type emotional speech synthesis without requiring a large number of emotional samples from the target subject.
[0021] Please see Figure 1 , Figure 1An exemplary embodiment of this application illustrates a method for synthesizing emotional speech, such as... Figure 1 As shown, this application provides a method for synthesizing emotional speech, including: S11, acquire audio samples of speech with various emotions and the target voice of the target object; S12, based on the target vocal timbre, perform timbre transfer processing on multiple speaking sample audios to obtain a target multi-emotion long audio; S13, Obtain the voice attribute control parameters input by the user; S14: Based on the speech attribute control parameters and the target multi-emotion long audio, speech synthesis is performed to obtain the target emotion speech.
[0022] The emotional speech synthesis method provided in this application performs timbre transfer processing on audio samples of various emotions based on the target voice of the target object to obtain a target multi-emotion long audio. Then, based on the voice attribute control parameters input by the user and the target multi-emotion long audio, speech synthesis is performed to obtain the target emotional speech. In this way, rich emotional information can be intuitively understood through the target emotional speech, thereby improving the flexibility, subtlety and adaptability of synthesized speech in emotional expression.
[0023] Optionally, obtain audio samples of speech with various emotions and the target voice of the target object, including: Acquire audio samples of speech with various emotions, and obtain corresponding film and television works for the target audience; Extracting the original spoken audio of the target subject from film and television works; The timbre of the original speech audio is extracted to obtain the target voice of the target object.
[0024] In the embodiment provided in this application, the original spoken audio of the target object is extracted from the film and television works corresponding to the target object, and the timbre of the original spoken audio is extracted to obtain the target voice of the target object. This achieves targeted acquisition of the voice, ensures the accuracy of the target voice, and thus improves the matching degree between the final synthesized target emotional speech and the target object. The film and television works include TV dramas (TV series, movies) and short dramas. If there are inconsistencies in the voices of the target object extracted from different film and television works, the voice that appears most frequently corresponding to the target object is determined as the target voice.
[0025] In an exemplary embodiment provided in this application, the specific steps for obtaining audio samples of speech with various emotions and the target voice of the target object are as follows: Target object timbre extraction: Collect a relatively long (longer than the set duration, which can be any value from 1 second to 1 minute) original audio samples of the target object in movies, TV series and short dramas, without distinguishing emotions, to extract the vocal lines containing the timbre features of the target object; Create a multi-emotional voice set for the target audience: Obtain all the corresponding emotional speech samples needed for the target audience in the text-to-speech process from a high-quality voice sample library pre-recorded by professional voice actors and covering a variety of emotional types.
[0026] Optionally, based on the target vocal timbre, timbre transfer processing is performed on multiple speech sample audios to obtain a long audio file with multiple target emotions, including: Multiple speaking sample audios are spliced and fused together to obtain a long sample audio with multiple emotions; Using a pre-trained speech recognition model, timbre transfer processing is performed on long audio samples with multiple emotions based on the target vocal tone to obtain target long audio samples with multiple emotions.
[0027] In the embodiment provided in this application, multiple speaking sample audios are spliced and fused to merge the sample audios recorded by professional voice actors with various corresponding emotional expressions into a single long audio with multiple emotional expressions, thus obtaining a multi-emotion long sample audio. Then, using a pre-trained speech recognition model, the multi-emotion long sample audio is processed by timbre transfer based on the target voice to obtain the target multi-emotion long audio. This not only meets the voice requirements of users when performing speech synthesis, but also enables high-quality, multi-type emotional speech synthesis without the need for a large number of target person multi-emotion samples.
[0028] In an exemplary embodiment provided in this application, the multi-emotion long sample audio is pre-recorded with segmentation points based on the duration and merging order of each speech sample audio. After the migration is completed, the migrated multi-emotion long audio is divided into several "target timbre-multi-emotion" segmented audios with different emotions according to the segmentation points and segmentation rules. This generates a new set of multi-emotion speech samples with target timbre but carrying professional emotional expression. That is, it is used to extract the multi-emotion voice set of the target object that has both target timbre features and all the corresponding emotions required in the text-to-speech process, so as to facilitate the direct acquisition for speech use and / or selective retrieval for speech synthesis, thereby improving the efficiency of speech use and reducing the difficulty of speech synthesis.
[0029] Please see Figure 2 , Figure 2 This is a schematic diagram of the timbre transfer process for a target multi-emotion long audio file in an exemplary embodiment of this application, as shown below. Figure 2 As shown, the timbre transfer process for target multi-emotion long audio is as follows: Extract sample audio (original speaking audio) of a target object of a certain duration from film and television dramas and short dramas, perform timbre extraction, and obtain the target object's voice; Retrieve professionally recorded sample audio (speaking sample audio) of all the emotion types required for text-to-speech conversion of the target object from within the system. Merge all the professionally recorded sample audios corresponding to the required emotion types and record the segmentation points to obtain long sample audios with multiple emotions; The merged audio is subjected to vocal transfer to obtain a merged sample audio that is consistent with the timbre of the target object, while completely preserving the emotional prosodic features of the merged sample audio before transfer, thus obtaining the target multi-emotion long audio. Based on the segmentation points, the merged sample audio after voice transfer is segmented into a multi-emotion sample audio set that is consistent with the timbre of the target object; A multi-emotion vocal line set of the target object is created using a segmented audio set of multi-emotion samples that matches the timbre of the target object.
[0030] Optionally, using a pre-trained speech recognition model, timbre transfer processing is performed on long audio samples with multiple emotions based on the target vocal tone to obtain target long audio samples with multiple emotions, including: Input the long audio samples with multiple emotions into the pre-trained speech recognition model, extract features from the long audio samples with multiple emotions, and obtain the semantic features of each word in the long audio samples with multiple emotions. The semantic features are adjusted according to the target vocal features to obtain the target character's pronunciation. Multiple target word speech samples are aligned based on multi-emotion long audio samples to obtain target multi-emotion long audio.
[0031] In the embodiment provided in this application, firstly, multi-emotion long-sample audio is input into a pre-trained speech recognition model. Feature extraction is performed on the multi-emotion long-sample audio to obtain the semantic features of each character in the multi-emotion long-sample audio. Then, the semantic features are adjusted according to the target vocal timbre to obtain the target character's speech. Next, multiple target character speech samples are aligned based on the multi-emotion long-sample audio to obtain the target multi-emotion long audio, thereby achieving speech synthesis that meets the user's emotional needs. The semantic features include emotional features and character vocal timbre features. Emotional features include emotional prosody, and character vocal timbre features include specific pronunciation and duration.
[0032] In an exemplary embodiment provided in this application, the vocal timbre of a merged multi-emotion long sample audio is transferred to the vocal timbre of a target object. The specific steps are as follows: A speech recognition model pre-trained on a large-scale dataset is used as a feature extractor to extract semantic features from the multi-emotion long sample audio. These features are similar to phonemes, including the specific pronunciation, duration, and emotional prosody of each word in the corresponding text, while simultaneously separating the original speaker information. Then, using the target object's vocal audio as a condition and the semantic features of the original audio as input, a diffusion model is used to generate the timbre features of the target object. Finally, the timbre transfer is completed, and the pronunciation, duration, emotion, and prosody features are aligned with the original audio, resulting in the target multi-emotion long audio. Because a single long audio contains a variety of different emotional expressions, the loss is smaller compared to transferring each emotion sample audio separately. Therefore, while ensuring the consistency of timbre after multi-emotion sample transfer, the emotional prosody features of the merged sample audio before transfer are also completely preserved.
[0033] Optionally, the speech attribute control parameters include the input text, emotion attributes, and emotion level; speech synthesis is performed based on the speech attribute control parameters and the target multi-emotion long audio to obtain the target emotion speech, including: Feature extraction is performed on long audio clips with multiple emotions to obtain emotional features that match both emotional attributes and emotional levels. Feature extraction is performed on long audio files with multiple emotions to obtain the vocal features of each character in the input text. The input text is synthesized based on emotional features and the vocal features of multiple words to obtain the target emotional speech.
[0034] In the embodiment provided in this application, feature extraction is performed on the target multi-emotion long audio to obtain emotional features that match both emotional attributes and emotional levels, as well as the vocal features of each character in the input text. The input text is then synthesized according to the emotional features and multiple vocal features to obtain the target emotional speech. In this way, rich emotional information can be intuitively understood through the target emotional speech, thereby improving the flexibility, subtlety and adaptability of synthesized speech in emotional expression.
[0035] In this embodiment, the emotional attributes include 26 emotional types: neutral, happy, sad, angry, disgusted, fearful, surprised, expectant, guilty, shy, humorous, serious, firm, pleading, resentful, tense, excited, calm, tired, gentle, proud, helpless, gritted teeth, smug, doubtful, and chattering. Except for neutral emotions, each emotion has an emotional level of low, medium, and high.
[0036] Optionally, the speech attribute control parameters include the input text, emotion attributes, and emotion level; speech synthesis is performed based on the speech attribute control parameters and the target multi-emotion long audio to obtain the target emotion speech, including: Identify candidate emotional voices that match both emotional attributes and emotional levels from a range of long audio recordings with multiple emotions. Extract the initial word speech that matches each word in the input text from the candidate emotional speech; Multiple initial word speech samples are concatenated and fused according to the order of the input text to obtain the initial emotional speech sample; The initial emotional speech is processed to achieve word coherence, thus obtaining the target emotional speech.
[0037] In the embodiment provided in this application, firstly, candidate emotional speech that matches both the emotional attribute and the emotional level is found from the target multi-emotion long audio, and initial word speech that matches each word in the input text is extracted from the candidate emotional speech. Secondly, the multiple initial word speech are concatenated and fused according to the order of the input text, and then word-to-word coherence processing is performed to obtain the target emotional speech. In this way, rich emotional information can be intuitively understood through the target emotional speech, thereby improving the flexibility, subtlety, and adaptability of synthesized speech in emotional expression.
[0038] In an exemplary embodiment provided in this application, a practical application example of emotion-based speech merging is as follows: In practical applications, the user first provides input text and selects an emotional attribute from 26 given emotion types. Each emotion, except for neutral emotions, has three levels: low, medium, and high, resulting in 76 possible combinations. This determines the emotional attribute and level, forming the speech attribute control parameters. The system automatically matches the target voice and the target multi-emotion long audio file corresponding to the emotional combination to synthesize speech, outputting a synthesized result that meets the requirements and has a high degree of timbre-emotion matching, thus obtaining the target emotional speech.
[0039] Optionally, an emotion-based speech synthesis method further includes: Get frequently used modal words input by the user; By adapting and integrating commonly used interjections with the target emotional speech, a new target emotional speech is obtained.
[0040] In the embodiment provided in this application, commonly used interjections input by the user are adapted and fused with the target emotional speech to obtain a new target emotional speech, which can improve the degree of customization of speech synthesis and thus improve the user's satisfaction with speech synthesis.
[0041] In summary, the key improvements of the emotional speech synthesis method in this application include timbre transfer based on multiple emotional samples, decoupling and recombination of timbre and emotion, construction of multiple emotional speech samples, and fine-grained emotion modulation.
[0042] Specifically, the timbre transfer method based on multiple emotional samples involves transferring the timbre features of a target character from films, television dramas, or short dramas to multiple emotional sample audio recordings made by professional voice actors. This generates a variety of high-quality emotional voices that the target character originally lacked, effectively solving the problem of limited emotional expression caused by the single emotional sample in the original drama. This innovative process of transferring the target character's timbre to professional multi-emotional samples effectively utilizes high-quality emotional sample resources, overcoming the limitations imposed by the limited emotional types in the original drama's voice on the synthesis effect, and significantly improving the realism and richness of the emotional expression in the dubbing.
[0043] The decoupling and recombination mechanism of timbre and emotion adopts a "merge-transfer, then segment and use" approach. First, multiple emotional sample audios are merged and subjected to unified timbre transfer. Then, they are segmented and used as needed. This process preserves the original emotional prosodic features while achieving timbre replacement, thus supporting the expression of multiple natural emotions under the same timbre. In this way, through the merging-transfer-segmentation mechanism, effective separation and recombination of timbre and emotion are achieved while maintaining timbre consistency. This allows the same timbre to flexibly carry multiple natural emotions, improving the vividness and expressiveness of synthesized speech.
[0044] A highly efficient multi-emotion speech sample construction process: Relying only on a small amount of original speech from the target subject, combined with a reusable professional emotion sample library, a multi-emotion speech dataset with rich emotional expressiveness can be quickly constructed through a merging, migration, and segmentation process. This significantly reduces reliance on original multi-emotion data. Thus, a multi-emotion speech synthesis system can be rapidly built using only a small amount of original speech from the target subject, combined with a reusable professional emotion sample library, greatly reducing dependence on multi-emotion data from specific speakers and improving the applicability and efficiency of technology deployment.
[0045] Fine-grained emotion control capability: Supports detailed control of multiple basic emotion types and their varying intensity levels, distinguishing up to 26 emotion types: neutral, happy, sad, angry, disgusted, fearful, surprised, expectant, guilty, shy, humorous, serious, determined, pleading, resentful, tense, excited, calm, tired, gentle, arrogant, helpless, clenching teeth, smug, doubtful, and talkative. Except for neutral emotions, each emotion has low, medium, and high levels, totaling 76 combinations, enabling refined and customized output of synthesized speech emotions. This fine-grained control over multiple emotion types and their intensities provides more accurate and expressive speech synthesis support for dubbing content in films, television dramas, and short dramas, expanding the application scenarios of the technology.
[0046] Please see Figure 3 , Figure 3 An exemplary embodiment of this application illustrates an emotion-based speech synthesis system, such as... Figure 3 As shown, this application provides an emotion-based speech synthesis system 300, comprising: The sample acquisition module 301 is used to acquire audio samples of speech with various emotions and the target voice of the target object; The timbre transfer module 302 is used to perform timbre transfer processing on multiple speaking sample audios based on the target voice to obtain a target multi-emotion long audio; The control parameter acquisition module 303 is used to acquire the voice attribute control parameters input by the user; The speech synthesis module 304 is used to synthesize speech based on speech attribute control parameters and target multi-emotion long audio to obtain target emotional speech.
[0047] The emotional speech synthesis system 300 provided in this application utilizes a timbre transfer module 302 to perform timbre transfer processing on the speech sample audio of various emotions obtained by the sample acquisition module 301 based on the target voice of the target object obtained by the sample acquisition module 301, thereby obtaining a target multi-emotion long audio. Then, the speech synthesis module 304 uses the speech attribute control parameters input by the user obtained by the control parameter acquisition module 303 and the target multi-emotion long audio to perform speech synthesis, thereby obtaining the target emotional speech. In this way, rich emotional information can be intuitively understood through the target emotional speech, thereby improving the flexibility, subtlety and adaptability of synthesized speech in emotional expression.
[0048] Optionally, the sample acquisition module is specifically used for: Acquire audio samples of speech with various emotions, and obtain corresponding film and television works for the target audience; Extracting the original spoken audio of the target subject from film and television works; The timbre of the original speech audio is extracted to obtain the target voice of the target object.
[0049] Optionally, the tone transfer module 302 is specifically used for: Multiple speaking sample audios are spliced and fused together to obtain a long sample audio with multiple emotions; Using a pre-trained speech recognition model, timbre transfer processing is performed on long audio samples with multiple emotions based on the target vocal tone to obtain target long audio samples with multiple emotions.
[0050] Optionally, the tone transfer module 302 is specifically used for: Input the long audio samples with multiple emotions into the pre-trained speech recognition model, extract features from the long audio samples with multiple emotions, and obtain the semantic features of each word in the long audio samples with multiple emotions. The semantic features are adjusted according to the target vocal features to obtain the target character's pronunciation. Multiple target word speech samples are aligned based on multi-emotion long audio samples to obtain target multi-emotion long audio.
[0051] Optionally, the voice attribute control parameters include input text, emotion attributes, and emotion level; the speech synthesis module 304 is specifically used for: Feature extraction is performed on long audio clips with multiple emotions to obtain emotional features that match both emotional attributes and emotional levels. Feature extraction is performed on long audio files with multiple emotions to obtain the vocal features of each character in the input text. The input text is synthesized based on emotional features and the vocal features of multiple words to obtain the target emotional speech.
[0052] Optionally, the voice attribute control parameters include input text, emotion attributes, and emotion level; the speech synthesis module 304 is specifically used for: Identify candidate emotional voices that match both emotional attributes and emotional levels from a range of long audio recordings with multiple emotions. Extract the initial word speech that matches each word in the input text from the candidate emotional speech; Multiple initial word speech samples are concatenated and fused according to the order of the input text to obtain the initial emotional speech sample; The initial emotional speech is processed to achieve word coherence, thus obtaining the target emotional speech.
[0053] It should be noted that the emotion-based speech synthesis system and the emotion-based speech synthesis method provided in the above embodiments belong to the same concept. The specific methods by which each module and unit performs its operations have been described in detail in the method embodiments and will not be repeated here. In practical applications, the emotion-based speech synthesis system provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above. This is not a limitation here.
[0054] A computing device according to an embodiment of this application includes a memory, a processor, and a program stored in the memory and running on the processor. When the processor executes the program, it implements some or all of the steps of the above-described emotional speech synthesis method.
[0055] The computing device can be a computer, and the corresponding program is computer software. The parameters and steps in the computing device described above can be referred to the parameters and steps in the embodiment of the emotion speech synthesis method above, and will not be repeated here.
[0056] Figure 4 A schematic diagram of a computer system suitable for implementing the embodiments of this application is shown. It should be noted that... Figure 4 The computer system 400 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0057] like Figure 4 As shown, the computer system 400 includes a Central Processing Unit (CPU) 401, which can perform various appropriate actions and processes, such as executing the methods described in the above embodiments, based on programs stored in Read-Only Memory (ROM) 402 or programs loaded from Storage Unit 408 into Random Access Memory (RAM) 403. The RAM 403 also stores various programs and data required for system operation. The CPU 401, ROM 402, and RAM 403 are interconnected via a bus 404. An Input / Output (I / O) interface 405 is also connected to the bus 404.
[0058] The following components are connected to I / O interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to I / O interface 405 as needed. A removable medium 411, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 410 as needed so that computer programs read from it can be installed into storage section 408 as needed.
[0059] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 409, and / or installed from removable medium 411. When the computer program is executed by central processing unit (CPU) 401, it performs various functions defined in the system of this application.
[0060] This application embodiment provides a computer-readable storage medium storing instructions that, when executed, perform the steps of the aforementioned emotion-based speech synthesis method. The computer-readable storage medium can be either transient or non-transient.
[0061] The technical solutions of this disclosure can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of this disclosure. The aforementioned computer-readable storage medium can be a non-transitory computer-readable storage medium, including: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, and other media capable of storing program code; it can also be a transient computer-readable storage medium.
[0062] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0063] Those skilled in the art will recognize that this application can be implemented as a system, method, or computer program product. Therefore, this disclosure can be implemented in the following forms: it can be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "module" or "system." Furthermore, in some embodiments, this application can also be implemented as a computer program product contained in one or more computer-readable media, which contains computer-readable program code. Computer-readable storage media can be, for example, but not limited to—electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof.
[0064] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0065] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A method for synthesizing emotional speech, characterized in that, include: Acquire audio samples of speech with various emotions and the target voice of the target object; Based on the target vocal timbre, multiple speaking sample audios are subjected to timbre transfer processing to obtain a target multi-emotion long audio; Obtain the voice attribute control parameters input by the user; Speech synthesis is performed based on the speech attribute control parameters and the target multi-emotion long audio to obtain the target emotion speech.
2. The method according to claim 1, characterized in that, The acquisition of audio samples of speech with multiple emotions and the target voice of the target object includes: Acquire audio samples of speech with various emotions, and obtain corresponding film and television works for the target audience; Extract the original speaking audio of the target object from the aforementioned film and television works; The timbre of the original spoken audio is extracted to obtain the target voice of the target object.
3. The method according to claim 1, characterized in that, The step of performing timbre transfer processing on multiple spoken audio samples based on the target vocal timbre to obtain a target multi-emotion long audio file includes: Multiple spoken audio samples are spliced and fused together to obtain a long audio sample with multiple emotions. Using a pre-trained speech recognition model, the target voice is used to perform timbre transfer processing on the long audio sample of multiple emotions to obtain the target long audio of multiple emotions.
4. The method according to claim 3, characterized in that, The process of using a pre-trained speech recognition model to perform timbre transfer processing on the multi-emotion long audio sample based on the target vocal timbre to obtain the target multi-emotion long audio includes: The long sample audio of multiple emotions is input into the pre-trained speech recognition model, and the features of the long sample audio of multiple emotions are extracted to obtain the semantic features of each word in the long sample audio of multiple emotions. The semantic features are adjusted according to the target vocal timbre to obtain the target character pronunciation of the corresponding character. Based on the multi-emotion long sample audio, multiple target word speech are aligned to obtain the target multi-emotion long audio.
5. The method according to any one of claims 1 to 4, characterized in that, The voice attribute control parameters include input text, emotion attributes, and emotion level; the process of voice synthesis based on the voice attribute control parameters and the target multi-emotion long audio to obtain target emotion voice includes: Feature extraction is performed on the target multi-emotion long audio to obtain emotional features that match the emotional attributes and the emotional level. Feature extraction is performed on the target multi-emotion long audio to obtain the vocal features of each character that matches the input text; The input text is synthesized using the emotional features and multiple vocal features of the characters to obtain the target emotional speech.
6. The method according to any one of claims 1 to 4, characterized in that, The voice attribute control parameters include input text, emotion attributes, and emotion level; the process of voice synthesis based on the voice attribute control parameters and the target multi-emotion long audio to obtain target emotion voice includes: From the target multi-emotion long audio, find alternative emotional voices that match both the emotional attribute and the emotional level; Extract the initial word speech that matches each word in the input text from the candidate emotional speech; Multiple initial word speech samples are concatenated and fused according to the order of the input text to obtain the initial emotional speech. The initial emotional speech is processed to achieve word coherence, thereby obtaining the target emotional speech.
7. An emotion-based speech synthesis system, characterized in that, include: The sample acquisition module is used to acquire audio samples of speech with various emotions and the target voice of the target object; The timbre transfer module is used to perform timbre transfer processing on multiple speaking sample audios based on the target voice to obtain a target multi-emotion long audio; The control parameter acquisition module is used to acquire the voice attribute control parameters input by the user. The speech synthesis module is used to synthesize speech based on the speech attribute control parameters and the target multi-emotion long audio to obtain the target emotional speech.
8. The system according to claim 7, characterized in that, The sample acquisition module is specifically used for: Acquire audio samples of speech with various emotions, and obtain corresponding film and television works for the target audience; Extract the original speaking audio of the target object from the aforementioned film and television works; The timbre of the original spoken audio is extracted to obtain the target voice of the target object.
9. A computing device, comprising a memory, a processor, and a program stored in the memory and running on the processor, characterized in that, When the processor executes the program, it implements the steps of an emotional speech synthesis method as described in any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a terminal device, cause the terminal device to perform the steps of an emotion-based speech synthesis method as described in any one of claims 1 to 6.