Speech synthesis method and device, electronic equipment and storage medium

By fusing the features of the acquired emotional intensity command and the pre-trained model, and selecting the matching reference audio to generate the target speech spectrum, the problem of not being able to control the emotional intensity of speech in the existing technology is solved, and the precise control of emotional intensity and the realism of speech synthesis are achieved.

CN121565138APending Publication Date: 2026-02-24GUANGZHOU HUYA TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511684363.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing speech synthesis technology cannot effectively control the intensity of emotional expression in speech, resulting in unrealistic emotional expressions in the generated speech.

Method used

By acquiring the text to be synthesized and the emotion intensity instruction, a pre-trained emotion text processing model is used for feature extraction and fusion. A matching emotion reference audio is selected to generate the target speech spectrum, thereby achieving fine control over the emotion intensity.

Benefits of technology

It achieves precise control over the emotional intensity of speech, and the generated speech can realistically express different emotional intensities, thus improving the realism of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121565138A_ABST
    Figure CN121565138A_ABST
Patent Text Reader

Abstract

The invention relates to the field of voice processing, in particular to a voice synthesis method and device, electronic equipment and a storage medium, and the method comprises the steps: respectively extracting to-be-synthesized text features and emotion instruction features of a to-be-synthesized text and an emotion intensity instruction; processing the to-be-synthesized text features and the emotion instruction features through an emotion text processing model to obtain emotion text features; selecting a target emotional audio matched with the emotional label of the emotional intensity instruction; generating a target speech spectrum according to the target emotional audio and the emotional text features; and performing acoustic reconstruction on the target speech spectrum to obtain a target emotion speech. Compared with the prior art, the emotion intensity instruction added with the emotion intensity level is obtained, the to-be-synthesized text feature and the emotion intensity instruction are fused by utilizing the pre-trained emotion text processing model, so that the emotion label and the emotion intensity level are fused into the emotion text feature, and the emotion intensity is effectively controlled; and target emotional voices capable of truly expressing different emotions and fine intensities are generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech processing, and more specifically, to a speech synthesis method, apparatus, electronic device, and storage medium. Background Technology

[0002] Speech synthesis technology is a technique that converts text into speech. During the synthesis process, a target timbre and emotional style can be specified by inputting reference audio, thereby making the generated speech more closely match the desired expressive effect. Speech synthesis technology can efficiently generate speech, reducing the cost of manual dubbing. Therefore, it is widely used in various fields, such as live video streaming and map navigation.

[0003] However, current speech synthesis technologies typically only allow specifying the target emotion, without controlling the intensity of that emotion. For example, one might want to control the generation of angry speech, but cannot achieve fine-grained control over the degree of anger. This results in unrealistic emotional expression in the speech generated by existing speech synthesis technologies.

[0004] Therefore, there is an urgent need for a method that can effectively control the emotional intensity of speech synthesis. Summary of the Invention

[0005] This invention provides a speech synthesis method, apparatus, electronic device, and storage medium for effectively controlling the emotional intensity of speech synthesis, making the generated speech more realistic.

[0006] According to a first aspect of this application, a speech synthesis method is provided, the method comprising: Obtain the text to be synthesized and the sentiment intensity instruction; the sentiment intensity instruction includes sentiment tags and the sentiment intensity level corresponding to the sentiment tags; Feature extraction is performed on the text to be synthesized and the emotion intensity command respectively to obtain the features of the text to be synthesized and the features of the emotion command; The emotional text features are obtained by fusing the features of the text to be synthesized and the emotional instruction features using a pre-trained emotional text processing model. Select a target emotional audio that matches the emotional tag of the emotional intensity instruction from the preset emotional reference audio; A target speech spectrum is generated based on the target emotional audio and the emotional text features.

[0007] Optionally, the training of the sentiment text processing model includes: Acquire training audio, and obtain the emotion tags and training text corresponding to the training audio; Construct the emotional intensity command for the trained speech audio; Extract the reference sentiment text features from the training speech audio; Feature extraction is performed on the training speech text and emotional intensity command of the training speech audio respectively to obtain training speech text features and training emotional command features; The training speech text features and training emotional command features corresponding to the training speech audio are fused by the emotional text processing model to be trained to obtain the training emotional text features of the training speech audio. The feature loss between the reference sentiment text features and the training sentiment text features is calculated according to a preset loss function, and the sentiment text processing model is updated according to the feature loss.

[0008] Optionally, the instruction for constructing the emotional intensity of the training speech audio includes: Extract the effective speech audio containing human voices from the training speech audio; Extract the fundamental frequency of each audio frame in the effective speech audio, and obtain the average fundamental frequency of the effective speech audio based on the fundamental frequency of the audio frames; Extract the energy of each audio frame in the effective speech audio, and obtain the average energy of the effective speech audio based on the energy of the audio frames; Based on the training speech text, obtain the audio rate of the effective speech audio; The emotional intensity level of the effective speech audio is obtained based on the average fundamental frequency, average energy, and audio speed of the effective speech audio. Based on the emotional intensity level and emotional label corresponding to the training speech audio, construct the emotional intensity instruction of the training speech audio.

[0009] Optionally, extracting the effective speech audio containing human voices from the training speech audio includes: Perform sound detection on the training audio to obtain the start and end positions of the human voice portion in the training audio; Based on the start position and the end position, extract valid speech audio containing human voice from the training speech audio.

[0010] Optionally, obtaining the audio rate of the effective speech audio based on the training speech text includes: Remove punctuation from the training speech text to obtain the preprocessed speech text; Count the number of characters in the preprocessed speech text; Based on the valid audio recordings, obtain the valid audio duration; The audio speed of the training speech text is calculated based on the effective speech duration and the number of characters.

[0011] Optionally, obtaining the emotional intensity level of the effective speech audio based on its average fundamental frequency, average energy, and speech rate includes: The base frequency level of the effective speech audio is obtained according to the preset base frequency level division strategy and the average base frequency; The energy level of the effective speech audio is obtained based on the preset energy level classification strategy and the average energy. Based on the preset speech rate level classification strategy and the speech rate of the audio, the speech rate level of the effective speech audio is obtained; One or more of the fundamental frequency level, the energy level, and the speech rate level can be used as the emotional intensity level corresponding to the effective speech audio.

[0012] Optionally, generating the target speech spectrum based on the target emotional audio and the emotional text features includes: Extract the initial speech spectrum of the target emotional audio; Based on the initial speech spectrum of the target emotional audio, extract the speech semantic features and target timbre features of the target emotional audio; The pre-trained emotional speech generation model performs feature fusion on the speech semantic features, the target timbre features, and the emotional text features to obtain fused features, and generates the target speech spectrum based on the fused features.

[0013] According to a second aspect of this application, a speech synthesis apparatus is provided, the apparatus comprising: The input acquisition module is used to acquire the input text to be synthesized and the emotion intensity instruction; the emotion intensity instruction includes an emotion tag and the emotion intensity level corresponding to the emotion tag; The feature extraction module is used to extract features from the text to be synthesized and the emotion intensity command respectively, so as to obtain the features of the text to be synthesized and the emotion command features. The feature fusion module is used to fuse the features of the text to be synthesized and the features of the emotional instruction through a pre-trained emotional text processing model to obtain emotional text features. The reference audio extraction module is used to select target emotional audio that matches the emotional tag of the emotional intensity instruction from the preset emotional reference audio. The spectrogram generation module is used to generate a target speech spectrum based on the target emotional audio and the emotional text features.

[0014] According to a third aspect of this application, an electronic device is provided, comprising: Memory, used to store one or more computer programs; A processor, when the one or more computer programs are executed by the processor, implements the speech synthesis method described in the first aspect above.

[0015] According to a fourth aspect of this application, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the speech synthesis method described in the first aspect above.

[0016] Based on any of the above aspects, the speech synthesis method, apparatus, electronic device, and computer storage medium provided in this application embodiment acquires an emotional intensity instruction with added emotional intensity level, and uses a pre-trained emotional text processing model to fuse the features of the text to be synthesized and the emotional instruction features, so that the emotional label and emotional intensity level can be fused with the feature items of the text to be synthesized, effectively realizing the control of emotional intensity, and thus obtaining emotional text features that can express both emotions and subtle emotional intensities, so that the target emotional speech obtained based on the emotional text features can more realistically express different emotions. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is an illustrative application scenario diagram of the speech synthesis method provided in this embodiment.

[0019] Figure 2 This is a flowchart illustrating the steps of the speech synthesis method provided in this embodiment.

[0020] Figure 3 This is a schematic diagram illustrating the steps involved in training the emotional text processing model provided in this embodiment.

[0021] Figure 4 This is a flowchart illustrating the steps involved in constructing the emotion intensity command provided in this embodiment.

[0022] Figure 5 This is a flowchart illustrating the steps for obtaining audio speech rate in this embodiment.

[0023] Figure 6 This is a schematic diagram of the functional modules of the speech synthesis device provided in this embodiment.

[0024] Figure 7 This is a schematic diagram of the structure of the electronic device provided in this embodiment. Detailed Implementation

[0025] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this application. To better illustrate the following embodiments, some components in the drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions of the product; it is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0026] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0028] Speech synthesis technology is a technique that converts text into speech. During the synthesis process, a target timbre and emotional style can be specified by inputting reference audio, thereby making the generated speech more closely match the desired expressive effect. Speech synthesis technology can efficiently generate speech, reducing the cost of manual dubbing. Therefore, it is widely used in various fields, such as live video streaming and map navigation.

[0029] However, current speech synthesis technologies typically only allow specifying the target emotion, without controlling the intensity of that emotion. For example, while aiming to generate angry speech, it's impossible to achieve fine-grained control over the degree of anger. This is because existing models cannot generate feature representations of different emotional intensities. The model doesn't understand the difference between "somewhat angry" and "quite angry," so the emotional features generated by the model can only contain the characteristics of the emotion itself. This results in speech expressions based on the same emotion having the same emotional expression, leading to unrealistic emotional portrayals.

[0030] This embodiment provides a technical solution that can solve the above problems. The specific implementation of this application will be described in detail below with reference to the accompanying drawings.

[0031] An exemplary diagram illustrating an application scenario of a speech synthesis method provided in this application embodiment is shown below. Figure 1 As shown, the application scenario includes at least a server 100 and a terminal 200 that can communicate with the server 100.

[0032] Understandably, the server 100 can be an independent electronic device or a cluster of multiple electronic devices; the terminal 200 can be a smartphone terminal, personal computer, tablet computer, vehicle terminal, etc., but is not limited to these.

[0033] In one possible implementation, server 100 and terminal 200 may respectively execute the speech synthesis method provided in the embodiments of this application, or, optionally, the speech synthesis method provided in the embodiments of this application may be partially executed in server 100 and partially executed in terminal 200.

[0034] like Figure 2 As shown, this embodiment provides a speech synthesis method, which may include the following steps: S1: Obtain the text to be synthesized and the emotion intensity instruction; In this embodiment, the text to be synthesized is a text reference used for speech synthesis, and the emotion intensity instruction includes an emotion tag and the emotion intensity level corresponding to the emotion tag; In one embodiment, the emotional intensity level includes one or more of a fundamental frequency level, an energy level, and a speech rate level, which can be set as the emotional intensity level according to actual needs. Preferably, the emotional intensity level includes a fundamental frequency level, an energy level, and a speech rate level.

[0035] Here, fundamental frequency can be understood as the pitch of speech, and the fundamental frequency level can be divided according to a preset fundamental frequency level division strategy to obtain several fundamental frequency levels; energy can be understood as the amplitude of sound in speech, which can be obtained by calculating the average of the square of the amplitude of speech per unit time, and the energy level can be divided according to a preset energy level division strategy to obtain several energy levels; speech rate is the rate at which human voices speak, and the speech rate level can be divided according to a preset speech rate level division strategy to obtain several speech rate levels. By preset corresponding level division strategies, subtle facial expressions can be uniformly expressed.

[0036] The emotional intensity instruction in step S1 can be generated based on a preset instruction format, using emotional tags, and one or more of fundamental frequency level, energy level, and speech rate level. For example, the emotional intensity instruction can be expressed as: happy, fundamental frequency level (pitch): level four (higher), energy level: level five (highest), speech rate level: level three (medium).

[0037] By acquiring the emotional intensity instruction that includes the emotional intensity level, it can serve as a basis for subsequent adjustments and control of emotional intensity.

[0038] S2: Extract features from the text to be synthesized and the emotion intensity instruction respectively to obtain the features of the text to be synthesized and the features of the emotion instruction; In this embodiment, a pre-trained feature extraction model can be used to extract features from the text to be synthesized and the sentiment intensity instruction.

[0039] S3: The features of the text to be synthesized and the features of the emotional instruction are fused using a pre-trained emotional text processing model to obtain the emotional text features; In this embodiment, the emotional text processing model can effectively fuse the emotional instruction features with the text features to be synthesized through training, and integrate the emotional tags and emotional intensity in the emotional instruction features into the text features to be synthesized, so that the resulting emotional text features can not only express the emotional type of the target, but also achieve control over subtle emotional intensity.

[0040] S4: Select a target emotional audio that matches the emotional tag of the emotional intensity instruction from the preset emotional reference audio; In this embodiment, the emotional reference audio for various emotions can be obtained through the following steps: By using a pre-trained emotion embedding model, the preset timbre reference audio and the emotion reference audio corresponding to the emotion are synthesized to obtain the first candidate emotion reference audio. By adjusting the pre-trained emotional instruction adjustment model and the emotional intensity instruction corresponding to the emotion, the preset timbre reference audio is adjusted to obtain the second candidate emotional reference audio. Analyze the first candidate emotional reference audio and the second candidate emotional reference audio, and select the third candidate emotional reference audio with a higher degree of matching with the emotion from the first candidate emotional reference audio and the second candidate emotional reference audio; The volume of the third candidate emotional reference audio is normalized, and the audio segment containing human voice in the normalized third candidate emotional reference audio is extracted as the emotional reference audio of the emotion.

[0041] Understandably, the timbre reference audio can be understood as a voice recording of a speaker with a specific timbre speaking, which contains the timbre information of the speaker with that specific timbre; the emotion reference audio can be understood as a voice recording of any speaker speaking with a certain emotion, which contains the emotion information of that emotion.

[0042] By generating multiple candidate emotional reference audios using different generation methods, and then selecting the emotional reference audio that best matches the emotion, the matching accuracy between the emotional reference audio and the emotion can be effectively improved, thereby providing a better basis for selecting the target emotional audio in step S4.

[0043] S5: Generate the target speech spectrum based on the target emotional audio and the emotional text features; Preferably, the target speech spectrum can be the Mel spectrum of the target. Compared with other conventional linear spectra, the Mel spectrum is more consistent with the perceptual characteristics of the human auditory system. Therefore, by generating the Mel spectrum of the target as the target speech spectrum, the human voice can be better reconstructed and restored.

[0044] In this embodiment, step S5 may include the following steps: Extract the initial speech spectrum of the target emotional audio; based on the initial speech spectrum of the target emotional audio, extract the speech semantic features and target timbre features of the target emotional audio; perform feature fusion on the speech semantic features, the target timbre features and the emotional text features through the pre-trained emotional speech generation model to obtain fused features, and generate the target speech spectrum based on the fused features.

[0045] In one implementation, the extraction of the initial speech spectrum may include resampling the target emotional audio to a 16kHz, 16-bit, single-channel format, and extracting the initial speech spectrum based on the resampled target emotional audio. Resampling the target emotional audio ensures audio format consistency while balancing storage overhead and avoiding multi-channel interference.

[0046] In one implementation, the extraction of speech semantic features may include processing the initial speech spectrum using a pre-trained speech semantic extraction model to extract the speech semantic features; and obtaining the fused features by extracting and fusing the speech semantic features, so that the fused features can include fine-grained prior information on emotion and prosodic style in the target emotional audio, guiding the generated speech to align with the reference audio in terms of prosody, rhythm, and emotional expression, thereby enhancing the naturalness and expressiveness of the emotional synthesis.

[0047] In one embodiment, the extraction of the target timbre features may include processing the initial speech spectrum using a pre-trained timbre feature extraction model to extract the target timbre features; and obtaining the fused features by extracting the target timbre features and fusing them, so that the fused features can guide the generated speech to match the timbre of the target.

[0048] In one embodiment of this invention, after obtaining the target speech spectrum, acoustic reconstruction can be performed on the target speech spectrum to obtain the target emotional speech. Preferably, the target speech spectrum can be acoustically reconstructed using a device or module with acoustic reconstruction function, such as a vocoder, to obtain the target emotional speech.

[0049] In this embodiment, as Figure 3 As shown, training the emotional text processing model may include the following steps: A1: Obtain the training audio, and obtain the emotion tag and training audio text corresponding to the training audio; Understandably, the training audio can be a recording of a speaker with a specific timbre speaking with a specific emotion, which contains the corresponding emotional information as well as the speaker's timbre information. The training audio can be used to extract both timbre information and emotional information simultaneously.

[0050] In one embodiment, the training speech audio can be obtained as described above, by having a speaker with a specific timbre speak with a specific emotion and recording the speech to obtain the training speech audio; in some embodiments, the training speech audio can be obtained by referring to the steps for obtaining the emotion reference audio described above, which will not be elaborated further here.

[0051] In this embodiment, the emotion tag identifies the emotional information contained in the training audio, such as happiness, anger, sadness, etc.; the training audio text is the text content corresponding to the speech in the training audio.

[0052] A2: Construct the emotional intensity instructions for the trained speech audio; In this embodiment, the training phase aims to train the emotional text processing model's ability to process emotional intensity commands, enabling it to recognize the emotional intensity levels within the commands and generate features suitable for speech synthesis. Therefore, unlike the execution and inference phase of the emotional text processing model, the training phase requires constructing training emotional intensity commands based on the training audio recordings to ensure the model can accurately identify the emotional intensity levels.

[0053] In this embodiment, as Figure 4 As shown, the construction of the emotional intensity instruction for the training speech audio may include the following sub-steps: A21: Extract the effective speech audio containing human voices from the training speech audio; In one implementation, the extraction of the valid speech audio may include: Sound detection is performed on the training audio to obtain the start and end positions of the human voice portion in the training audio; based on the start and end positions, valid audio containing human voice is extracted from the training audio.

[0054] Specifically, Voice Activity Detection (VAD) can be used to detect sound in the training audio and extract the human voice portion. By processing the extracted valid audio, noise interference can be effectively reduced, enabling the emotional text processing model to extract information from the training audio more accurately.

[0055] A22: Extract the fundamental frequency of each audio frame in the effective speech audio, and obtain the average fundamental frequency of the effective speech audio based on the fundamental frequency of the audio frame; In this embodiment, the base frequencies of all the audio frames can be added together to obtain the total base frequency of the effective speech audio, and the total base frequency can be divided by the number of audio frames contained in the effective speech audio to obtain the average base frequency.

[0056] A23: Extract the energy of each audio frame in the effective speech audio, and obtain the average energy of the effective speech audio based on the energy of the audio frames; In this embodiment, the total energy of the effective speech audio can be obtained by adding the energy of all the audio frames, and the total energy can be divided by the number of audio frames contained in the effective speech audio to obtain the average energy.

[0057] A24: Based on the training speech text, obtain the audio speed of the effective speech audio; In this embodiment, as Figure 5 As shown, obtaining the audio rate of the effective speech audio may include the following sub-steps: A241: Remove punctuation from the training speech text to obtain preprocessed speech text; A242: Count the number of characters in the preprocessed speech text; A243: Obtain the effective voice duration based on the effective voice audio; A244: The audio speed of the training speech text is calculated based on the effective speech duration and the number of characters.

[0058] Understandably, by removing punctuation from the training speech text, it is possible to better count the number of characters actually pronounced in the training speech text, and thus effectively calculate the audio speed of the training speech text based on the effective speech duration of the effective speech audio.

[0059] The pre-processed speech text can be processed using a pre-trained character statistics model or a character statistics algorithm to more accurately obtain the number of characters in the pre-processed speech text.

[0060] A25: Obtain the emotional intensity level of the effective speech audio based on the average fundamental frequency, average energy, and audio speed of the effective speech audio; As described above, the emotional intensity level may include one or more of the fundamental frequency level, energy level, and speech rate level. Therefore, in this embodiment, step A25 may include the following steps: The base frequency level of the effective speech audio is obtained according to the preset base frequency level division strategy and the average base frequency; The energy level of the effective speech audio is obtained based on the preset energy level classification strategy and the average energy. Based on the preset speech rate level classification strategy and the speech rate of the audio, the speech rate level of the effective speech audio is obtained; One or more of the fundamental frequency level, the energy level, and the speech rate level can be used as the emotional intensity level corresponding to the effective speech audio.

[0061] A26: Construct the emotional intensity instruction for the training speech audio based on the emotional intensity level and emotional label corresponding to the training speech audio.

[0062] In this embodiment, it can be constructed based on a preset instruction format and according to the emotion tag corresponding to the training speech audio, as well as one or more of the fundamental frequency level, energy level and speech rate level.

[0063] A3: Extract the reference emotional text features from the training speech audio; Understandably, the reference emotional text features directly extracted from the training audio are used as a reference for the emotional text processing model. A4: Extract features from the training speech text and emotional intensity command of the training speech audio respectively to obtain training speech text features and training emotional command features; A5: The training speech text features and training emotional command features corresponding to the training speech audio are fused by the emotional text processing model to be trained to obtain the training emotional text features of the training speech audio. A6: Calculate the feature loss between the reference sentiment text features and the training sentiment text features according to the preset loss function, and update the sentiment text processing model according to the feature loss.

[0064] Understandably, the training objective of the emotional text processing model is to generate training emotional text features that express the corresponding emotions and emotional intensities by using the features of the input training speech text and the features of the emotional intensity command. Then, the generated training emotional text features can be used to reconstruct the training speech audio that expresses the corresponding emotions and emotional intensities. The reference emotional text features are directly extracted from the training speech audio and contain features of actual emotions and emotional intensities. Therefore, by calculating the feature loss between the reference emotional text features and the training emotional text features, and updating the emotional text processing model through the feature loss, the emotional text processing model can be trained in a direction that generates reference emotional text features that are closer to reality.

[0065] Understandably, in this embodiment, if the feature loss does not meet the preset conditions, such as the loss function not converging, then the steps A1-A6 described above need to be repeated until the feature loss meets the preset conditions to obtain the trained sentiment text processing model.

[0066] This process involves pre-selecting several training audio files based on different emotions and emotional intensities, and obtaining corresponding training audio text for each training audio file. The timbre of the training audio files can be the same or different. A training set is constructed based on these training audio files and their corresponding training audio texts, enabling the emotional text processing model to be trained based on different emotions and emotional intensities, thereby improving the model's ability to generate different emotions and emotional intensities. Therefore, in step A1, training audio files can be obtained from the training set, along with the corresponding emotion tags and training audio text.

[0067] like Figure 6 As shown in the illustration, this application also provides a speech synthesis device. Optionally, the speech synthesis device may include: The input acquisition module 11 is used to acquire the input text to be synthesized and the emotion intensity instruction; the emotion intensity instruction includes an emotion tag and the emotion intensity level corresponding to the emotion tag; In this embodiment, the input acquisition module 11 can be used to perform... Figure 2 For a detailed description of the input acquisition module 11 shown in step S1, please refer to the description of step S1.

[0068] Feature extraction module 12 is used to extract features from the text to be synthesized and the emotion intensity instruction respectively, so as to obtain the features of the text to be synthesized and the features of the emotion instruction; In this embodiment, the feature extraction module 12 can be used to perform... Figure 2 For a detailed description of the feature extraction module 12 shown in step S2, please refer to the description of step S2.

[0069] Feature fusion module 13 is used to fuse the features of the text to be synthesized and the features of the emotional instruction through a pre-trained emotional text processing model to obtain emotional text features; In this embodiment, the feature fusion module 13 can be used to perform... Figure 2 For a detailed description of the feature fusion module 13 shown in step S3, please refer to the description of step S3.

[0070] The reference audio extraction module 14 is used to select a target emotional audio that matches the emotional tag of the emotional intensity instruction from a preset emotional reference audio. In this embodiment, the reference audio extraction module 14 can be used to perform... Figure 2 For a detailed description of the reference audio extraction module 14 shown in step S4, please refer to the description of step S4.

[0071] The spectrogram generation module 15 is used to generate a target speech spectrum based on the target emotional audio and the emotional text features; In this embodiment, the sound spectrum generation module 15 can be used to perform... Figure 2 For a detailed description of the acoustic spectrum generation module 15, please refer to the description of step S5 shown in step S5.

[0072] In one embodiment, the apparatus may further include a model training module 16, which is used to train the emotional text processing model. In this embodiment, the model training module 16 can be used to perform... Figure 3 For a detailed description of the model training module 16, see steps A1-A6 shown below.

[0073] It is understood that the above-described device embodiments and method embodiments can correspond to each other, and similar descriptions of the device embodiments can be referred to the method embodiments. To avoid repetition, further details are omitted here. The speech synthesis device provided in this application can execute a speech synthesis method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of executing the method. The functional modules of the speech synthesis device can be implemented in hardware, in software instructions, or in a combination of hardware and software modules.

[0074] Specifically, the steps of the method embodiments of this application can be implemented by integrated logic circuits in the processor hardware and / or instructions in software form. The steps of the speech synthesis method in conjunction with the embodiments of this application can be directly implemented by a hardware encoding processor, or by a combination of hardware and software modules in the encoding processor. Optionally, the software module can be located in random access memory, and storage media such as read-only memory, programmable read-only memory, flash memory, electrically erasable programmable memory, and registers are all acceptable. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps in the above method embodiments.

[0075] This application provides an electronic device with the following structure: Figure 7 As shown. The electronic device can be as described in this embodiment. Figure 1 The server 100 or terminal 200 shown.

[0076] The electronic device includes a memory 21, a processor 22, a communication module 23, and an input / output interface 24, etc. Optionally, the memory 21, the processor 22, the communication module 23, and the input / output interface 24 can be connected and communicate with each other through a bus 25.

[0077] The memory 21 is used to store one or more computer programs and to transfer the code of the computer programs to the processor 22; when the one or more computer programs are executed by the processor 22, the speech synthesis method in the embodiments of this application is implemented.

[0078] Optionally, the electronic device can be connected to a network via communication module 23 to communicate with other devices, such as terminals or servers, to achieve data interaction. The electronic device can be various forms of digital computers, exemplarily such as desktop computers, servers, workbenches, mainframes, or other types of computers. The electronic device can also be various forms of mobile terminals, exemplarily such as smartphones, tablets, wearable devices (such as helmets, glasses, watches, etc.), and other similar mobile terminals.

[0079] Optionally, the electronic device can connect to required input / output devices, such as a keyboard or display device, via the input / output interface 24. The electronic device itself may have a display device, and other display devices can also be connected externally via the input / output interface 24. Optionally, a storage device, such as a hard disk, can also be connected via the input / output interface 24 to store data from the electronic device, read data from the storage device, or store data from the storage device in the memory 21. It is understood that the input / output interface 24 can be a wired interface or a wireless interface. Depending on the actual application scenario, the device connected to the input / output interface 24 can be a component of the electronic device or an external device connected to the electronic device when needed.

[0080] Optionally, the memory 21 may be a volatile memory and / or a non-volatile memory. The volatile memory may be a random access memory, etc., and the non-volatile memory may be a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, or a flash memory, etc.

[0081] Optionally, the computer program stored in the processor 22 can be divided into one or more modules, which are stored in the memory 21 and executed by the processor 22 to perform the method provided in this embodiment. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the electronic device.

[0082] Optionally, the processor 22 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 22 include, but are not limited to, a central processing unit, a graphics processing unit, a digital signal processor, various special-purpose artificial intelligence computing chips, various processors running machine learning model algorithms, and can also be any suitable controller, microcontroller, processor, etc. The processor 22 executes the various methods and processes of this embodiment, exemplarily, such as a speech synthesis method according to an embodiment of this application.

[0083] Optionally, the bus 25 may include a path for transmitting information. Depending on its function, the bus 25 may be divided into an address bus, a data bus, a control bus, etc.

[0084] In an optional implementation, this application embodiment also provides a computer storage medium storing a computer program thereon, which, when executed by a computer, enables the computer to perform the methods described in the above-described method embodiments. Part or all of the computer program can be loaded and / or installed on the memory 21 of an electronic device. When the computer program is executed by the processor 22, one or more steps of a speech synthesis method according to an embodiment of this application can be performed.

[0085] Optionally, the computer-readable storage medium may be a random access memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, etc.

[0086] Obviously, the above embodiments of this application are merely examples for clearly illustrating the technical solution of this application, and are not intended to limit the specific implementation of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the claims of this application should be included within the protection scope of the claims of this application.

Claims

1. A speech synthesis method, characterized in that, The method includes: Obtain the text to be synthesized and the sentiment intensity instruction; the sentiment intensity instruction includes sentiment tags and the sentiment intensity level corresponding to the sentiment tags; Feature extraction is performed on the text to be synthesized and the emotion intensity command respectively to obtain the features of the text to be synthesized and the features of the emotion command; The emotional text features are obtained by fusing the features of the text to be synthesized and the emotional instruction features using a pre-trained emotional text processing model. Select a target emotional audio that matches the emotional tag of the emotional intensity instruction from the preset emotional reference audio; A target speech spectrum is generated based on the target emotional audio and the emotional text features.

2. The speech synthesis method according to claim 1, characterized in that, The training of the sentiment text processing model includes: Acquire training audio, and obtain the emotion tags and training text corresponding to the training audio; Construct the emotional intensity command for the trained speech audio; Extract the reference sentiment text features from the training speech audio; Feature extraction is performed on the training speech text and emotional intensity command of the training speech audio respectively to obtain training speech text features and training emotional command features; The training speech text features and training emotional command features corresponding to the training speech audio are fused by the emotional text processing model to be trained to obtain the training emotional text features of the training speech audio. The feature loss between the reference sentiment text features and the training sentiment text features is calculated according to a preset loss function, and the sentiment text processing model is updated according to the feature loss.

3. The speech synthesis method according to claim 2, characterized in that, The instructions for constructing the emotional intensity of the training speech audio include: Extract the effective speech audio containing human voices from the training speech audio; Extract the fundamental frequency of each audio frame in the effective speech audio, and obtain the average fundamental frequency of the effective speech audio based on the fundamental frequency of the audio frames; Extract the energy of each audio frame in the effective speech audio, and obtain the average energy of the effective speech audio based on the energy of the audio frames; Based on the training speech text, obtain the audio rate of the effective speech audio; The emotional intensity level of the effective speech audio is obtained based on the average fundamental frequency, average energy, and audio speed of the effective speech audio. Based on the emotional intensity level and emotional label corresponding to the training speech audio, construct the emotional intensity instruction of the training speech audio.

4. The speech synthesis method according to claim 3, characterized in that, The extraction of effective speech audio containing human voices from the training speech audio includes: Perform sound detection on the training audio to obtain the start and end positions of the human voice portion in the training audio; Based on the start position and the end position, extract valid speech audio containing human voice from the training speech audio.

5. The speech synthesis method according to claim 3, characterized in that, The step of obtaining the audio rate of the effective speech audio based on the training speech text includes: Remove punctuation from the training speech text to obtain the preprocessed speech text; Count the number of characters in the preprocessed speech text; Based on the valid audio recordings, obtain the valid audio duration; The audio speed of the training speech text is calculated based on the effective speech duration and the number of characters.

6. The speech synthesis method according to claim 3, characterized in that, The step of obtaining the emotional intensity level of the effective speech audio based on the average fundamental frequency, average energy, and audio speech rate of the effective speech audio includes: The base frequency level of the effective speech audio is obtained according to the preset base frequency level division strategy and the average base frequency; The energy level of the effective speech audio is obtained based on the preset energy level classification strategy and the average energy. Based on the preset speech rate level classification strategy and the speech rate of the audio, the speech rate level of the effective speech audio is obtained; One or more of the fundamental frequency level, the energy level, and the speech rate level can be used as the emotional intensity level corresponding to the effective speech audio.

7. The speech synthesis method according to any one of claims 1-5, characterized in that, The step of generating the target speech spectrum based on the target emotional audio and the emotional text features includes: Extract the initial speech spectrum of the target emotional audio; Based on the initial speech spectrum of the target emotional audio, extract the speech semantic features and target timbre features of the target emotional audio; The pre-trained emotional speech generation model performs feature fusion on the speech semantic features, the target timbre features, and the emotional text features to obtain fused features, and generates the target speech spectrum based on the fused features.

8. A speech synthesis device, characterized in that, The device includes: The input acquisition module is used to acquire the input text to be synthesized and the emotion intensity instruction; the emotion intensity instruction includes an emotion tag and the emotion intensity level corresponding to the emotion tag; The feature extraction module is used to extract features from the text to be synthesized and the emotion intensity command respectively, so as to obtain the features of the text to be synthesized and the emotion command features. The feature fusion module is used to fuse the features of the text to be synthesized and the features of the emotional instruction through a pre-trained emotional text processing model to obtain emotional text features. The reference audio extraction module is used to select target emotional audio that matches the emotional tag of the emotional intensity instruction from the preset emotional reference audio. The spectrogram generation module is used to generate a target speech spectrum based on the target emotional audio and the emotional text features.

9. An electronic device, characterized in that, include: Memory, used to store one or more computer programs; A processor, when the one or more computer programs are executed by the processor, implements the speech synthesis method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the speech synthesis method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Speech synthesis method and equipment thereof

    CN111048062A

  • Speech synthesis method and device, electronic equipment and storage medium

    CN113096640A

  • Emotional speech synthesis method and synthesis device

    CN114927122A

  • Model training method and device, equipment, storage medium and program product

    CN117216532A

  • Speech synthesis method and device, electronic equipment, storage medium and program product

    CN120199228A