Speech synthesis method, device, electronic device and storage medium

By training the emotion encoding model and utilizing the minimum mutual information loss and contrastive learning loss, the problem of insufficient emotion control in the existing speech synthesis model is solved, the control of fine-grained emotions is achieved, and the flexibility and controllability of speech synthesis are improved.

CN119763545BActive Publication Date: 2025-09-23IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411906533.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-09-23
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Existing speech synthesis models lack fine-grained control over emotions, resulting in insufficient control over speech synthesis.

Method used

By training the emotion encoding model and utilizing the minimum mutual information loss and contrastive learning loss, irrelevant information in speech and text content is extracted and removed to achieve fine-grained emotion control.

Benefits of technology

It improves the control of speech synthesis models over emotions, can output fine-grained emotional features, and enhances the flexibility and control of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119763545B_ABST
    Figure CN119763545B_ABST
Patent Text Reader

Abstract

The present invention provides a speech synthesis method, device, electronic device, and storage medium, relating to the field of speech technology, wherein the method comprises: inputting the acquired text to be synthesized and emotional attributes into a speech synthesis model to obtain a target speech output by the speech synthesis model; wherein the speech synthesis model is trained based on a first sample text corresponding to a first sample speech and a first sample emotional feature corresponding to the first sample speech, the first sample emotional feature is obtained by inputting the first sample speech into an emotional coding model, and the emotional coding model is trained based on the minimum mutual information loss of the target code and the second sample emotional feature. The present invention can obtain an emotional coding model based on the minimum mutual information loss training, so that the emotional features output by the emotional coding model do not include irrelevant information such as timbre and text content, so that the speech synthesis model can achieve fine-grained emotional control, thereby improving the controllability of speech synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech technology, and in particular to a speech synthesis method, device, electronic device and storage medium. Background Art

[0002] In recent years, Text To Speech (TTS) has been widely used in scenarios such as smart assistants, speakers, in-car applications, novel reading, and (short) video dubbing. These open-domain speech synthesis needs have put forward new requirements for the controllability and diversity of speech synthesis.

[0003] In related technologies, high-quality sound library recording data is usually used to train speech synthesis models, but the trained speech synthesis models lack fine-grained control over emotions, thereby reducing the controllability of speech synthesis. Summary of the Invention

[0004] The present invention provides a speech synthesis method, device, electronic device and storage medium, which are used to solve the defect of reducing the controllability of speech synthesis in the prior art.

[0005] The present invention provides a speech synthesis method, comprising the following steps.

[0006] Obtain the text to be synthesized and its emotional attributes;

[0007] Inputting the text to be synthesized and the emotional attribute into a speech synthesis model to obtain a target speech output by the speech synthesis model;

[0008] In which, the speech synthesis model is trained based on the first sample text corresponding to the first sample speech and the first sample emotional feature corresponding to the first sample speech, the first sample emotional feature is obtained after the first sample speech is input into the emotional coding model, the emotional coding model is trained based on the minimum mutual information loss of the target coding and the second sample emotional feature, the target coding includes the timbre coding of the second sample speech and / or the text content coding of the second sample text of the second sample speech, and the second sample emotional feature is obtained after the second sample speech is input into the initial emotional coding model.

[0009] According to a speech synthesis method provided by the present invention, obtaining emotional attributes includes:

[0010] When receiving the attribute description text input by the user, inputting the attribute description text into the text macro model to obtain the emotional attribute output by the text macro model;

[0011] When the emotion template voice input by the user is received, the emotion template voice is input into the emotion coding model to obtain the emotion feature output by the emotion coding model, and the emotion feature is determined as the emotion attribute.

[0012] According to a speech synthesis method provided by the present invention, the emotion coding model is trained based on the following method:

[0013] Inputting the second sample speech into the speech coding network of the initial emotion coding model to obtain semantic features output by the speech coding network;

[0014] Inputting the preset learnable emotion space and the semantic features into the speech coding network of the initial emotion coding model to obtain the second sample emotion feature output by the emotion coding network;

[0015] Determining a first minimum mutual information loss between the second sample emotion feature and the timbre encoding, and determining a second minimum mutual information loss between the second sample emotion feature and the text content encoding;

[0016] Based on the first minimum mutual information loss and the second minimum mutual information loss, network parameters of the emotion coding network are updated to obtain the emotion coding model.

[0017] According to a speech synthesis method provided by the present invention, updating the network parameters of the emotion coding network based on the first minimum mutual information loss and the second minimum mutual information loss to obtain the emotion coding model includes:

[0018] Inputting the emotion description text corresponding to the second sample speech into the description text encoding network of the initial emotion encoding model to obtain the emotion description code output by the description text encoding network;

[0019] determining a contrastive learning loss based on the second sample emotion feature and the emotion description code;

[0020] Based on the first minimum mutual information loss, the second minimum mutual information loss, and the contrastive learning loss, network parameters of the emotion encoding network and network parameters of the description text encoding network are updated to obtain the emotion encoding model.

[0021] According to a speech synthesis method provided by the present invention, the speech synthesis model is trained based on the following method:

[0022] Inputting the first sample text and the first sample emotion feature into a speech feature extraction network of an initial speech synthesis model to obtain sample speech features output by the speech feature extraction network;

[0023] Inputting the sample speech features into the speech decoding network of the initial speech synthesis model to obtain a first predicted speech output by the speech decoding network;

[0024] Based on the first predicted speech and the first sample speech, the network parameters of the speech feature extraction network and the network parameters of the speech decoding network are updated to obtain the speech synthesis model.

[0025] According to a speech synthesis method provided by the present invention, inputting the text to be synthesized and the emotional attribute into a speech synthesis model to obtain a target speech output by the speech synthesis model includes:

[0026] Inputting the text to be synthesized, the emotional attribute, and other attributes into the speech synthesis model to obtain the target speech output by the speech synthesis model; the other attributes include at least one of the following: target speech environment, target sound quality level, target language, and target speech style; the sample speech feature is obtained by inputting the first sample text, the first sample emotional feature, and other sample attributes corresponding to the first sample speech into the speech feature extraction network;

[0027] Among them, the other sample attributes include at least one of the following: a sample voice environment corresponding to the first sample voice, a sample sound quality level corresponding to the first sample voice, a sample language corresponding to the first sample voice, and a sample voice style corresponding to the first sample voice.

[0028] According to a speech synthesis method provided by the present invention, the network parameters of the speech feature extraction network and the network parameters of the speech decoding network are updated based on the first predicted speech and the first sample speech to obtain the speech synthesis model, including:

[0029] Based on the first predicted speech and the first sample speech, updating network parameters of the speech feature extraction network and network parameters of the speech decoding network to obtain a reference speech synthesis model;

[0030] Obtaining a third sample text corresponding to a third sample voice and a third sample emotion feature corresponding to the third sample voice, wherein the third sample voice is a voice recorded by a sample subject at a sample location;

[0031] Inputting the third sample text, the third sample emotional feature, and the identifier of the sample object into the reference speech synthesis model to obtain a second predicted speech output by the reference speech synthesis model;

[0032] Based on the second predicted speech and the third sample speech, the network parameters of the speech feature extraction network and the network parameters of the speech decoding network in the reference speech synthesis model are updated to obtain the speech synthesis model.

[0033] The present invention also provides a speech synthesis device, comprising:

[0034] An acquisition unit, used to acquire the text to be synthesized and the emotional attributes;

[0035] A synthesis unit, configured to input the text to be synthesized and the emotional attribute into a speech synthesis model to obtain a target speech output by the speech synthesis model;

[0036] In which, the speech synthesis model is trained based on the first sample text corresponding to the first sample speech and the first sample emotional feature corresponding to the first sample speech, the first sample emotional feature is obtained after the first sample speech is input into the emotional coding model, the emotional coding model is trained based on the minimum mutual information loss of the target coding and the second sample emotional feature, the target coding includes the timbre coding of the second sample speech and / or the text content coding of the second sample text of the second sample speech, and the second sample emotional feature is obtained after the second sample speech is input into the initial emotional coding model.

[0037] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-described speech synthesis methods when executing the computer program.

[0038] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements any of the above-mentioned speech synthesis methods when executed by a processor.

[0039] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned speech synthesis methods.

[0040] The speech synthesis method, device, electronic device, and storage medium provided by the present invention input the acquired text to be synthesized and the emotional attributes into a trained speech synthesis model to obtain the target speech output by the speech synthesis model. The speech synthesis model is trained based on a first sample text corresponding to a first sample speech and a first sample emotional feature corresponding to the first sample speech. The first sample emotional feature is obtained by inputting the first sample speech into an emotional coding model. The emotional coding model is trained based on the minimum mutual information loss between a target code and a second sample emotional feature. The target code includes a timbre code of the second sample speech and / or a text content code of the second sample text of the second sample speech. The second sample emotional feature is obtained by inputting the second sample speech into an initial emotional coding model. It can be seen that the present invention can train an emotional coding model based on the minimum mutual information loss between the target code and the second sample emotional feature, so that the emotional features output by the emotional coding model do not include irrelevant information such as timbre and text content. This enables the speech synthesis model trained based on the emotional features output by the emotional coding model and the sample text to achieve fine-grained emotional control, thereby improving the controllability of speech synthesis. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0042] Figure 1 It is a flowchart of the speech synthesis method provided by an embodiment of the present invention.

[0043] Figure 2 This is an example diagram of speech synthesis provided by an embodiment of the present invention.

[0044] Figure 3 This is one of the schematic diagrams of the training process of the emotion coding model provided by an embodiment of the present invention.

[0045] Figure 4 Schematic diagram of a public speech training set provided by an embodiment of the present invention.

[0046] Figure 5 This is the second schematic diagram of the training process of the emotion coding model provided by an embodiment of the present invention.

[0047] Figure 6 3 is a schematic diagram of the framework of the initial emotion coding model provided by an embodiment of the present invention.

[0048] Figure 7This is one of the schematic diagrams of the training process of the speech synthesis model provided by an embodiment of the present invention.

[0049] Figure 8 3 is a schematic diagram of the framework of the initial speech synthesis model provided by an embodiment of the present invention.

[0050] Figure 9 This is the second schematic diagram of the training process of the speech synthesis model provided by an embodiment of the present invention.

[0051] Figure 10 It is a structural diagram of a speech synthesis device provided by an embodiment of the present invention.

[0052] Figure 11 It is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0053] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0054] The following combination Figures 1-9 The speech synthesis method of the present invention is described. The execution subject of the speech synthesis method can be an electronic device such as a terminal, a tablet computer, or a computer, or a speech synthesis device provided in the electronic device. The speech synthesis device can be implemented by software, hardware, or a combination of both.

[0055] Figure 1 FIG. 1 is a flow chart of a speech synthesis method according to an embodiment of the present invention. Figure 1 As shown, the speech synthesis method includes the following steps:

[0056] Step 101: Obtain the text to be synthesized and the emotional attributes.

[0057] For example, when a user has a need for speech synthesis, the text to be synthesized and the emotional attributes that need to be synthesized can be input into the electronic device, so that the electronic device obtains the text to be synthesized and the emotional attributes, or the electronic device obtains the previously stored text to be synthesized and the emotional attributes from the memory, for example, the emotional attributes include sad emotions or happy emotions, etc.

[0058] Step 102: input the text to be synthesized and the emotional attribute into a speech synthesis model to obtain a target speech output by the speech synthesis model.

[0059] In which, the speech synthesis model is trained based on the first sample text corresponding to the first sample speech and the first sample emotional feature corresponding to the first sample speech, the first sample emotional feature is obtained after the first sample speech is input into the emotional coding model, the emotional coding model is trained based on the minimum mutual information loss of the target coding and the second sample emotional feature, the target coding includes the timbre coding of the second sample speech and / or the text content coding of the second sample text of the second sample speech, and the second sample emotional feature is obtained after the second sample speech is input into the initial emotional coding model.

[0060] For example, when the text to be synthesized and the emotional attributes are obtained, the text to be synthesized and the emotional attributes are input into a trained speech synthesis model, and the target speech is generated and output based on the text to be synthesized and the emotional attributes by the speech synthesis model, so that the output target speech includes not only the semantics of the text to be synthesized, but also the emotions corresponding to the emotional attributes.

[0061] It should be noted that the identifier of the target speaker may also be input into the speech synthesis model so that the speech synthesis model ultimately outputs the target speech obtained using the timbre of the target speaker, and the present invention does not limit this.

[0062] The speech synthesis method provided by the present invention inputs the acquired text to be synthesized and the emotional attributes into a trained speech synthesis model to obtain the target speech output by the speech synthesis model. The speech synthesis model is trained based on the first sample text corresponding to the first sample speech and the first sample emotional feature corresponding to the first sample speech. The first sample emotional feature is obtained by inputting the first sample speech into an emotional coding model. The emotional coding model is trained based on the minimum mutual information loss between the target coding and the second sample emotional feature. The target coding includes the timbre coding of the second sample speech and / or the text content coding of the second sample text of the second sample speech. The second sample emotional feature is obtained by inputting the second sample speech into the initial emotional coding model. It can be seen that the present invention can train the emotional coding model based on the minimum mutual information loss between the target coding and the second sample emotional feature, so that the emotional features output by the emotional coding model do not include irrelevant information such as timbre and text content. Furthermore, the speech synthesis model trained based on the emotional features output by the emotional coding model and the sample text can achieve fine-grained emotional control, thereby improving the controllability of speech synthesis.

[0063] In one embodiment, obtaining the emotional attributes in step 101 may be achieved in the following manner:

[0064] The first way is to input the attribute description text input by the user into the text big model to obtain the emotional attribute output by the text big model.

[0065] For example, the user inputs the emotional attributes of speech synthesis in the form of attribute description text. For the attribute description text, the attribute description text can be input into the text big model, and the attribute description text is parsed by the text big model. Through prompt design, the text big model outputs the emotional attributes and other attributes involved in the attribute description text. For example, the attribute description text input by the user is "Generate high-quality speech, spoken in English, news style, happy", and the text big model parses it out as "sound quality level: high, language: English, speech style: news, emotion: happy". For attributes not involved in the attribute description text, the system default attributes can be used, such as "speech environment: recording studio". The parsed emotional attributes, other attributes and default attributes are then input into the speech synthesis model together to obtain the target speech output by the speech synthesis model.

[0066] The second way is to input the emotion template voice into the emotion coding model when receiving the emotion template voice input by the user, obtain the emotion feature output by the emotion coding model, and determine the emotion feature as the emotion attribute.

[0067] For example, when receiving user input of an emotion template speech, indicating that the user wishes to control the emotion of the synthesized target speech using the emotion template speech, the speech environment, sound quality level, language, and voice style adopt the user-specified specific or default values ​​of the speech attributes. For the emotion attribute, the emotion template speech is input into a trained emotion encoding model (also known as a fine-grained emotion encoding model). The emotion encoding model extracts the emotional features from the emotion template speech and determines these emotional features as emotion attributes. The parsed emotion attributes and the user-specified specific or default values ​​of the speech attributes are then input into the speech synthesis model to produce the target speech output by the speech synthesis model.

[0068] In the third method, the user directly enters the emotional attribute and other attributes. The electronic device directly inputs the specific values ​​of the emotional attribute and other attributes entered by the user into the speech synthesis model to obtain the target speech output by the speech synthesis model. For attributes not specified by the user, the default value is used.

[0069] Figure 2 This is an example diagram of speech synthesis provided by an embodiment of the present invention. Figure 2 As shown, any of the three methods mentioned above can be used to obtain different speech attributes, and the speech attributes, the text to be synthesized and the identifier of the target speaker are input into the speech synthesis model, so that the speech synthesis model finally outputs the target speech obtained using the timbre of the target speaker.

[0070] In this embodiment, the emotional attributes that the user wants to control in speech synthesis can be obtained based on different forms of user input, thereby improving the flexibility of speech synthesis.

[0071] In one embodiment, Figure 3 This is one of the training flow diagrams of the emotion coding model provided by the embodiment of the present invention. Figure 3 As shown, the emotion coding model is trained based on the following method:

[0072] Step 301: Input the second sample speech into the speech coding network of the initial emotion coding model to obtain the semantic features output by the speech coding network.

[0073] For example, each sample speech in the public speech training set and the high-quality speech training set can be used as the second sample speech, and the second sample speech can be input into the speech coding network of the initial emotion coding model, and the semantic features in the second sample speech can be extracted through the speech coding network.

[0074] It should be noted that the public speech training set pretrain-data can be obtained in the following ways: the public speech training set does not have high requirements on speech quality, but it needs to consider the coverage of various speech attributes. The public speech can be recorded or crawled from the Internet. The public speech needs to cover a variety of speech environments, speech styles, speech sources, etc. The present invention can classify speech attributes into five categories: speech environment, sound quality level, language, speech style, and fine-grained emotion, and obtain speech attributes through the following three different methods:

[0075] The first method is to use open source detection tools to automatically determine the category of the voice environment, sound quality level and language. The voice environment includes recording studios, conference rooms, lecture halls, and outdoors; the sound quality level is divided into high, medium, and low; the language is divided according to the actual data coverage, including Chinese, English, Russian, Japanese, Sichuanese, Northeastern dialect or Cantonese, etc.

[0076] The second method is to automatically judge the voice style based on the data acquisition channel. For example, the podcast data source can obtain tag information based on the website: finance, history, news, health, etc. The audio novel data source can be divided into fantasy, history, martial arts, suspense, horror, etc. according to the novel type.

[0077] The third method is to use manual annotation + self-supervised pre-training to obtain fine-grained emotions. For details, please refer to the following description of the emotion encoding model trained, and the emotion features can be output through the emotion encoding model. Figure 4 Schematic diagram of the public speech training set provided by the embodiment of the present invention, such as Figure 4As shown, the public speech training set includes multiple sample voices, the speech environment, sound quality level, language, voice style and fine-grained emotion corresponding to each sample voice.

[0078] It should be noted that finetune data for high-quality speech training sets can be obtained in the following ways: The speech samples in the high-quality speech training set are high-quality speech captured by professional speakers in a recording studio using professional equipment. Compared to the diverse speech environments, languages, voice styles, and emotional coverage of the pretrain data, the high-quality speech training set can only cover a subset of emotions and voice styles. The speech samples in the high-quality speech training set are annotated with emotional description text. The high-quality speech training set is used to build a speech synthesis model for the target speaker.

[0079] Step 302: Input the preset learnable emotion space and the semantic features into the emotion coding network of the initial emotion coding model to obtain the second sample emotion features output by the emotion coding network.

[0080] For example, due to the limited types of emotions, a preset learnable emotion space of fixed dimension is pre-designed. The preset learnable emotion space is composed of multiple emotion vectors, and each emotion vector represents a different emotion. The preset learnable emotion space is input into the emotion coding network of the initial emotion coding model. Normally, the emotion coding network can combine the emotions represented by the preset learnable emotion space to obtain more emotions. However, since more emotions may include emotions that are irrelevant to the input second sample speech, it is necessary to interact information between the emotion coding network and the speech coding network so that the emotion coding network obtains the semantic features output by the speech coding network, and then based on the semantic features and the emotions represented by the preset learnable emotion space, finally outputs the second sample emotion features that are consistent with the semantic features, also called emotion tokens.

[0081] In addition, the second sample speech is input into the timbre coding network of the initial emotion coding model to obtain the timbre coding output by the timbre coding network, and the second sample text corresponding to the second sample speech is input into the text coding network of the initial emotion coding model to obtain the text content coding output by the text coding network.

[0082] Step 303: Determine a first minimum mutual information loss between the second sample emotion feature and the timbre encoding, and determine a second minimum mutual information loss between the second sample emotion feature and the text content encoding.

[0083] For example, when obtaining the second sample emotion feature corresponding to the second sample speech, the timbre code corresponding to the second sample speech, and the text content code corresponding to the second sample speech, the first minimum mutual information loss of the second sample emotion feature and the timbre code is calculated based on the following formula (1): , the first minimum mutual information loss The purpose is to remove the timbre-related information from the emotional features of the second sample. Specifically, vCLUB estimation can be used:

[0084]

[0085] in, represents the second sample emotion feature corresponding to the i-th frame of speech in the second sample speech, represents the timbre code corresponding to the i-th frame of speech in the second sample speech, Indicates the first The timbre coding corresponding to the frame speech, Indicates Known The probability distribution of Indicates Known The probability distribution of Indicates the total number of frame speech in the second sample speech, also known as batch size. The physical meaning of expression is given the second sample emotion feature When the second sample sentiment feature The possibility of the timbre code x belonging to any other frame of speech is as consistent as possible. In other words, the emotional feature of the second sample It has nothing to do with the tone.

[0086] And the second minimum mutual information loss of the second sample sentiment feature and text content encoding is calculated based on the following formula (2): , the second minimum mutual information loss The purpose is to remove the text content related information from the second sample sentiment feature, which can be estimated using vCLUB:

[0087]

[0088] in, Indicates the text content encoding corresponding to the i-th frame of speech in the second sample speech, Indicates the first The text content encoding corresponding to the frame speech, Indicates Known The probability distribution of Indicates Known The second minimum mutual information loss The physical meaning of expression is given the second sample emotion feature When the second sample sentiment feature The possibility of the text content code t belonging to any other frame of speech is as consistent as possible. In other words, the emotional feature of the second sample It has nothing to do with the text content.

[0089] Step 304: Based on the first minimum mutual information loss and the second minimum mutual information loss, update the network parameters of the emotion coding network to obtain the emotion coding model.

[0090] For example, when the first minimum mutual information loss and the second minimum mutual information loss are obtained, the network parameters of the emotion coding network can be updated by the sum of the first minimum mutual information loss and the second minimum mutual information loss until the convergence condition is reached, and finally the emotion coding model is obtained.

[0091] In this embodiment, the network parameters of the emotion coding network can be updated based on the first minimum mutual information loss of the second sample emotion feature and timbre encoding, and the second minimum mutual information loss of the second sample emotion feature and text content encoding, so as to extract emotion-related information from the sample speech, and remove information irrelevant to emotion such as text content and timbre through the minimum mutual information decoupling training criterion, thereby improving the purity of the emotion information, thereby improving the purity of the emotion feature output by the trained emotion coding model, and improving the decoupling control capability of different speech attributes.

[0092] In one embodiment, Figure 5 This is the second diagram of the training process of the emotion coding model provided by the embodiment of the present invention. Figure 5 As shown, the above step 304 updates the network parameters of the emotion coding network based on the first minimum mutual information loss and the second minimum mutual information loss to obtain the emotion coding model, which can be specifically achieved by the following steps:

[0093] Step 3041: Input the emotion description text corresponding to the second sample speech into the description text encoding network of the initial emotion encoding model to obtain the emotion description code output by the description text encoding network.

[0094] For example, the annotator needs to annotate the emotion of the second sample speech. Based on his or her own perception of the second sample speech, the annotator uses descriptive words or texts that he or she feels are appropriate as the annotation result. The descriptive words can be one or more. For example, the annotation result is <a little bit of joyful feeling, lively tone, and relatively fast speaking speed>, so as to make a fine-grained description of the emotion of the second sample speech, and the annotation result is called the emotion description text corresponding to the second sample speech. The emotion description text corresponding to the second sample speech is input into the description text encoding network of the initial emotion encoding model, and the emotion description text is encoded by the description text encoding network to obtain the emotion description code output by the description text encoding network.

[0095] Step 3042: Determine the contrastive learning loss based on the second sample emotion feature and the emotion description code.

[0096] For example, when obtaining the second sample emotion feature and emotion description code corresponding to the second sample speech, the contrastive learning loss is calculated based on the following formula (3). The purpose of the contrastive learning loss is to align the second sample emotion feature and the corresponding emotion description code to the same space. Specifically, the second sample emotion feature is used as the anchor point, and the corresponding emotion description code is used as a positive and negative example. InfoNCE loss is used for training. The contrastive learning loss is:

[0097]

[0098] in, represents the emotion description code corresponding to the i-th frame of speech in the second sample speech, Indicates the first The emotion description encoding corresponding to the speech frame. f represents the distance metric, which can be a dot product. E represents the expectation, indicating that the contrastive learning loss is calculated based on the average result of the training set. It can be approximated using Monte Carlo estimation, typically implemented through batch-based (a set of training samples) gradient descent training. The physical meaning of contrastive learning loss is that among all emotion description encodings, the distance between the emotion feature of the second sample and the emotion description encoding that matches the emotion feature of the second sample is the closest.

[0099] Step 3043: Based on the first minimum mutual information loss, the second minimum mutual information loss, and the contrastive learning loss, the network parameters of the emotion encoding network and the network parameters of the description text encoding network are updated to obtain the emotion encoding model.

[0100] For example, when the first minimum mutual information loss, the second minimum mutual information loss, and the contrastive learning loss are obtained, the total loss can be determined based on the following formula (4): , based on the total loss The network parameters of the emotion encoding network and the network parameters of the description text encoding network are updated until the convergence conditions are reached, and finally the emotion encoding model is obtained.

[0101]

[0102] Figure 6 is a schematic diagram of the framework of the initial emotion coding model provided by the embodiment of the present invention, such as Figure 6 As shown, the initial emotion coding model includes a speech coding network (also called a speech information coding network), an emotion coding network (also called an emotion and mood coding network), a text coding network, a timbre coding network and a description text coding network, wherein the speech is input into the speech coding network to obtain the semantic features output by the speech coding network, the preset learnable emotion space is input into the emotion coding network, and information interaction is performed with the semantic features of the speech coding network to finally obtain the emotion and mood token output by the emotion coding network; the emotion description text of the speech is input into the description text coding network to obtain the emotion description code (also called description code) output by the description text coding network; the text content corresponding to the speech is input into the text coding network to obtain the text content code output by the text coding network; and the speech is input into the timbre coding network to obtain the timbre code output by the timbre coding network.

[0103] It should be noted that the network parameters of the emotion encoding network and the description text encoding network in the initial emotion encoding model require training, while the parameters of other modules can use pre-trained models. The emotion encoding network and the description text encoding network can use traditional text pre-trained models, such as the Bidirectional Encoder Representations from Transformers (BERT). The speech encoding network can use traditional speech pre-trained models, such as HuBERT and wav2vec; the text encoding network can use traditional text pre-trained models, such as BERT; and the timbre encoding network can use traditional voiceprint models, such as xvector.

[0104] In this embodiment, based on the first minimum mutual information loss of the second sample emotion feature and timbre coding, the second minimum mutual information loss of the second sample emotion feature and text content coding, and the contrastive learning loss determined based on the second sample emotion feature and emotion description coding, the network parameters of the emotion coding network and the network parameters of the description text coding network are updated by combining contrastive learning loss and minimum mutual information loss. This not only improves the purity of the emotion feature output by the trained emotion coding model, but also improves the accuracy of the emotion description coding output by the emotion coding model. In addition, for unlabeled sample speech, the trained emotion coding model can output accurate emotion features, thereby reducing the workload of labeling.

[0105] In one embodiment, Figure 7 This is one of the training flow diagrams of the speech synthesis model provided by the embodiment of the present invention, such as Figure 7 As shown, the speech synthesis model is trained based on the following method:

[0106] Step 701: Input the first sample text and the first sample emotion feature into the speech feature extraction network of the initial speech synthesis model to obtain the sample speech feature output by the speech feature extraction network.

[0107] The initial speech synthesis model may use VALL-E, Fish-speech, or chatTTS, etc., and the present invention does not limit this.

[0108] For example, Figure 8 : is a schematic diagram of the framework of the initial speech synthesis model provided by the embodiment of the present invention, such as Figure 8 As shown, the initial speech synthesis model includes a speech feature extraction network that converts text to semantic tokens and a speech decoding network. For example, the speech feature extraction network can be a GPT (Generative Pre-trained Transformer) model. The speech feature extraction network determines the content and rhythmic style of the synthesized speech. The speech attribute control module primarily operates on the speech feature extraction network. Specifically, the speech attribute control module inputs sample text and its corresponding emotional features. Furthermore, the synthesized speech of the previous frame can be input into the speech feature extraction network. This allows the speech feature extraction network to use the previous frame's synthesized speech as a reference to generate the synthesized speech for the current frame, thereby improving the accuracy of the synthesized speech.

[0109] For example, each sample speech in the public speech training set can be used as the first sample speech, and the first sample text corresponding to the first sample speech and the first sample emotional feature corresponding to the first sample speech can be input into the speech feature extraction network of the initial speech synthesis model to obtain the sample speech features output by the speech feature extraction network.

[0110] Step 702: Input the sample speech features into the speech decoding network of the initial speech synthesis model to obtain a first predicted speech output by the speech decoding network.

[0111] For example, when the sample speech features are obtained from the speech feature extraction network, the sample speech features are input into the speech decoding network of the initial speech synthesis model, and the sample speech features are decoded by the speech decoding network to obtain the first predicted speech corresponding to the first sample text.

[0112] Step 703: Based on the first predicted speech and the first sample speech, update the network parameters of the speech feature extraction network and the network parameters of the speech decoding network to obtain the speech synthesis model.

[0113] For example, when the first predicted speech corresponding to each first sample text is obtained, a first loss function is constructed based on the first predicted speech corresponding to each first sample text and the first sample speech (speech label) corresponding to each first sample text, and the network parameters of the speech feature extraction network and the network parameters of the speech decoding network are updated based on the first loss function until the convergence condition is reached, and finally a speech synthesis model is obtained.

[0114] In this embodiment, the initial speech synthesis model can be trained based on the sample text and sample emotion features corresponding to each sample speech in the public speech training set, so that the final speech synthesis model can achieve control of fine-grained emotions, thereby improving the controllability of speech synthesis.

[0115] In one embodiment, step 102 inputs the text to be synthesized and the emotional attribute into a speech synthesis model to obtain the target speech output by the speech synthesis model, which can be specifically achieved by:

[0116] The text to be synthesized, the emotional attributes and other attributes are input into the speech synthesis model to obtain the target speech output by the speech synthesis model; the other attributes include at least one of the following: target speech environment, target sound quality level, target language and target speech style, and the sample speech feature is obtained after the first sample text, the first sample emotional feature and other sample attributes corresponding to the first sample speech are input into the speech feature extraction network.

[0117] Among them, the other sample attributes include at least one of the following: a sample voice environment corresponding to the first sample voice, a sample sound quality level corresponding to the first sample voice, a sample language corresponding to the first sample voice, and a sample voice style corresponding to the first sample voice.

[0118] For example, in the stage of training the initial speech synthesis model based on the public speech training set, the obtained first sample speech corresponding to the sample speech environment, sample sound quality level, sample language and sample speech style, first sample emotional characteristics, first sample speech corresponding to the first sample text, and the identifier of the speaker corresponding to the first sample speech can be spliced ​​to obtain splicing information, and the splicing information can be input into the initial speech synthesis model for training.

[0119] Similarly, the text to be synthesized, emotional attributes, target speech environment, target sound quality level, target language, target speech style and the identification of the target speaker can be spliced ​​together to obtain target splicing information, and the target splicing information can be input into the speech synthesis model together to obtain the target speech of the target speaker output by the speech synthesis model.

[0120] It should be noted that the first sample emotion feature corresponding to the first sample speech can be obtained in different ways: for the sample speech with emotion description text annotation, the emotion description code of the annotated emotion description text is directly used as the first sample emotion feature (emotion token) corresponding to the first sample speech; for the sample speech without emotion description text annotation, the sample speech is input into the trained emotion coding model to obtain the emotion token output by the emotion coding model, and the emotion token is used as the first sample emotion feature corresponding to the first sample speech.

[0121] In this embodiment, the text to be synthesized, emotional attributes, target speech environment, target sound quality level, target language, target speech style and the identification of the target speaker can be input into the speech synthesis model together to obtain the target speech of the target speaker output by the speech synthesis model, thereby realizing the control of the speech synthesis model over different speech attributes and further improving the controllability of speech synthesis; in addition, the present invention uniformly divides speech attributes into emotion, speech environment, sound quality level, language and speech style, and obtains a more comprehensive speech attribute division system, so that the speech synthesis model can control different speech attributes, that is, control speech attributes of each dimension.

[0122] In one embodiment, Figure 9 This is the second diagram of the training process of the speech synthesis model provided by the embodiment of the present invention. Figure 9As shown, the above step 703 updates the network parameters of the speech feature extraction network and the network parameters of the speech decoding network based on the first predicted speech and the first sample speech to obtain the speech synthesis model, which can be specifically implemented in the following manner:

[0123] Step 7031: Based on the first predicted speech and the first sample speech, update the network parameters of the speech feature extraction network and the network parameters of the speech decoding network to obtain a reference speech synthesis model.

[0124] For example, when the first predicted speech corresponding to each first sample text is obtained, a first loss function is constructed based on the first predicted speech corresponding to each first sample text and the first sample speech (speech label) corresponding to each first sample text, and the network parameters of the speech feature extraction network and the network parameters of the speech decoding network are updated based on the first loss function until the convergence conditions are reached, and finally the speech synthesis model is referenced.

[0125] Step 7032: Obtain a third sample text corresponding to a third sample voice and a third sample emotion feature corresponding to the third sample voice, where the third sample voice is a voice recorded by a sample subject at a sample location.

[0126] For example, each sample speech in a high-quality speech training set can be used as a third sample speech, and a third sample text corresponding to the third sample speech can be obtained from the high-quality speech training set. The third sample emotion feature corresponding to the third sample speech can be obtained in different ways: for the third sample speech with emotion description text annotation, the emotion description encoding of the annotated emotion description text is directly used as the third sample emotion feature (emotion token) corresponding to the third sample speech; for the third sample speech without emotion description text annotation, the third sample speech is input into a trained emotion coding model to obtain the emotion token output by the emotion coding model, and the emotion token is used as the third sample emotion feature corresponding to the third sample speech.

[0127] Step 7033: Input the third sample text, the third sample emotional feature, and the identifier of the sample object into the reference speech synthesis model to obtain a second predicted speech output by the reference speech synthesis model.

[0128] For example, when the third sample text corresponding to the third sample speech and the third sample emotion feature corresponding to the third sample speech are obtained, the third sample text, the third sample emotion feature and the identifier of the sample object are input into the reference speech synthesis model to obtain the second predicted speech corresponding to the third sample text output by the reference speech synthesis model.

[0129] Step 7034: Based on the second predicted speech and the third sample speech, update the network parameters of the speech feature extraction network and the network parameters of the speech decoding network in the reference speech synthesis model to obtain the speech synthesis model.

[0130] For example, when the second predicted speech corresponding to each third sample text is obtained, a second loss function is constructed based on the second predicted speech corresponding to each third sample text and the third sample speech (speech label) corresponding to each third sample text, and the network parameters of the speech feature extraction network and the network parameters of the speech decoding network in the reference speech synthesis model are updated based on the second loss function until the convergence conditions are reached, and finally the speech synthesis model is obtained.

[0131] It should be noted that since the public speech training set contains a variety of different speech attributes, the reference speech synthesis model trained based on the public speech training set can cover the control of a variety of different speech attributes. In other words, the reference speech synthesis model has the performance of controlling a variety of different speech attributes. Therefore, even if the high-quality speech training set corresponds to a smaller number of speech attributes, the speech synthesis model finally trained can also have the control of other uncovered speech attributes, thereby realizing speech synthesis controlled by attribute transfer.

[0132] It should be noted that the speech synthesis model can also be trained based on a first proportion of sample speech in a high-quality speech training set and a second proportion of sample speech in a public speech training set, wherein the first proportion is greater than the second proportion, for example, the first proportion is 80% and the second proportion is 20%. The present invention does not limit this.

[0133] It should be noted that the present invention also subdivides the way of adding speech attributes. For speech attributes such as speech environment, language and speech style, the user controls the granularity coarsely and adopts coarse-grained discrete category representation; for emotional attributes, users need fine-grained control, and adopt the contrastive learning loss pre-training method to pull the emotional features and emotional description encoding into the same space, and combine with the constraint of minimum mutual information loss to further reduce the information in the emotional features that is not related to the emotion, thereby improving the purity of the emotional features.

[0134] In summary, the speech synthesis method provided by the present invention divides speech information other than text into speech attributes such as speech environment, sound quality level, language, voice style, emotion, and speaker's timbre. These speech attributes are annotated through automatic annotation or self-supervised decoupled pre-training and input into a speech synthesis model for control, thereby achieving control over the attributes of the synthesized speech. Specifically, for speech environment, sound quality level, language, and voice style, since user-required control precision is limited, annotation can be performed using designed categories and automatically identified through information sources or tools. For fine-grained emotion, a large amount of publicly available speech and emotion description text is pre-collected to extract fixed-length fine-grained emotion features. By comparing pre-training with the minimum mutual information constraints of text content and timbre, the fine-grained emotion encoding space is aligned with the text sentence-level semantic encoding space to extract attribute-decoupled fine-grained emotion features. After the different speech attribute annotations are obtained, they are input into the speech synthesis model along with the text to train a speech synthesis model base with controllable speech attributes. Finally, fine-tuning is performed with the target speaker's data to obtain a speech synthesis model with controllable speech attributes for the target speaker. During the inference phase, the text to be synthesized and the specified attributes input by the user are input into the trained speech synthesis model to achieve control over different attributes of speech synthesis, thereby improving the controllability and flexibility of the speech synthesis model.

[0135] The speech synthesis device provided by the present invention is described below. The speech synthesis device described below and the speech synthesis method described above can be referenced to each other.

[0136] Figure 10 FIG. 1 is a structural diagram of a speech synthesis device provided by an embodiment of the present invention. Figure 10 As shown, the speech synthesis device 1000 includes an acquisition unit 1001 and a synthesis unit 1002; wherein:

[0137] An acquisition unit 1001 is used to acquire the text to be synthesized and the emotional attributes;

[0138] A synthesis unit 1002 is configured to input the text to be synthesized and the emotional attribute into a speech synthesis model to obtain a target speech output by the speech synthesis model;

[0139] In which, the speech synthesis model is trained based on the first sample text corresponding to the first sample speech and the first sample emotional feature corresponding to the first sample speech, the first sample emotional feature is obtained after the first sample speech is input into the emotional coding model, the emotional coding model is trained based on the minimum mutual information loss of the target coding and the second sample emotional feature, the target coding includes the timbre coding of the second sample speech and / or the text content coding of the second sample text of the second sample speech, and the second sample emotional feature is obtained after the second sample speech is input into the initial emotional coding model.

[0140] The speech synthesis device provided by the present invention inputs the acquired text to be synthesized and the emotional attributes into a trained speech synthesis model to obtain the target speech output by the speech synthesis model. The speech synthesis model is trained based on the first sample text corresponding to the first sample speech and the first sample emotional feature corresponding to the first sample speech. The first sample emotional feature is obtained by inputting the first sample speech into an emotional coding model. The emotional coding model is trained based on the minimum mutual information loss between the target coding and the second sample emotional feature. The target coding includes the timbre coding of the second sample speech and / or the text content coding of the second sample text of the second sample speech. The second sample emotional feature is obtained by inputting the second sample speech into the initial emotional coding model. It can be seen that the present invention can train the emotional coding model based on the minimum mutual information loss between the target coding and the second sample emotional feature, so that the emotional features output by the emotional coding model do not include irrelevant information such as timbre and text content. Furthermore, the speech synthesis model trained based on the emotional features output by the emotional coding model and the sample text can achieve fine-grained emotional control, thereby improving the controllability of speech synthesis.

[0141] Based on any of the above embodiments, the acquiring unit 1001 is specifically configured to:

[0142] When receiving the attribute description text input by the user, inputting the attribute description text into the text macro model to obtain the emotional attribute output by the text macro model;

[0143] When the emotion template voice input by the user is received, the emotion template voice is input into the emotion coding model to obtain the emotion feature output by the emotion coding model, and the emotion feature is determined as the emotion attribute.

[0144] Based on any of the above embodiments, the emotion coding model is trained based on the following method:

[0145] Inputting the second sample speech into the speech coding network of the initial emotion coding model to obtain semantic features output by the speech coding network;

[0146] Inputting the preset learnable emotion space and the semantic features into the speech coding network of the initial emotion coding model to obtain the second sample emotion feature output by the emotion coding network;

[0147] Determining a first minimum mutual information loss between the second sample emotion feature and the timbre encoding, and determining a second minimum mutual information loss between the second sample emotion feature and the text content encoding;

[0148] Based on the first minimum mutual information loss and the second minimum mutual information loss, network parameters of the emotion coding network are updated to obtain the emotion coding model.

[0149] Based on any of the foregoing embodiments, updating the network parameters of the emotion coding network based on the first minimum mutual information loss and the second minimum mutual information loss to obtain the emotion coding model includes:

[0150] Inputting the emotion description text corresponding to the second sample speech into the description text encoding network of the initial emotion encoding model to obtain the emotion description code output by the description text encoding network;

[0151] determining a contrastive learning loss based on the second sample emotion feature and the emotion description code;

[0152] Based on the first minimum mutual information loss, the second minimum mutual information loss, and the contrastive learning loss, network parameters of the emotion encoding network and network parameters of the description text encoding network are updated to obtain the emotion encoding model.

[0153] Based on any of the above embodiments, the speech synthesis model is trained based on the following method:

[0154] Inputting the first sample text and the first sample emotion feature into a speech feature extraction network of an initial speech synthesis model to obtain sample speech features output by the speech feature extraction network;

[0155] Inputting the sample speech features into the speech decoding network of the initial speech synthesis model to obtain a first predicted speech output by the speech decoding network;

[0156] Based on the first predicted speech and the first sample speech, the network parameters of the speech feature extraction network and the network parameters of the speech decoding network are updated to obtain the speech synthesis model.

[0157] Based on any of the above embodiments, the synthesis unit 1002 is specifically configured to:

[0158] Inputting the text to be synthesized, the emotional attribute, and other attributes into the speech synthesis model to obtain the target speech output by the speech synthesis model; the other attributes include at least one of the following: target speech environment, target sound quality level, target language, and target speech style; the sample speech feature is obtained by inputting the first sample text, the first sample emotional feature, and other sample attributes corresponding to the first sample speech into the speech feature extraction network;

[0159] Among them, the other sample attributes include at least one of the following: a sample voice environment corresponding to the first sample voice, a sample sound quality level corresponding to the first sample voice, a sample language corresponding to the first sample voice, and a sample voice style corresponding to the first sample voice.

[0160] Based on any of the foregoing embodiments, the updating of network parameters of the speech feature extraction network and network parameters of the speech decoding network based on the first predicted speech and the first sample speech to obtain the speech synthesis model includes:

[0161] Based on the first predicted speech and the first sample speech, updating network parameters of the speech feature extraction network and network parameters of the speech decoding network to obtain a reference speech synthesis model;

[0162] Obtaining a third sample text corresponding to a third sample voice and a third sample emotion feature corresponding to the third sample voice, wherein the third sample voice is a voice recorded by a sample subject at a sample location;

[0163] Inputting the third sample text, the third sample emotional feature, and the identifier of the sample object into the reference speech synthesis model to obtain a second predicted speech output by the reference speech synthesis model;

[0164] Based on the second predicted speech and the third sample speech, the network parameters of the speech feature extraction network and the network parameters of the speech decoding network in the reference speech synthesis model are updated to obtain the speech synthesis model.

[0165] Figure 11 FIG is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention, such as Figure 11As shown, the electronic device may include: a processor 1110, a communications interface 1120, a memory 1130, and a communication bus 1140, wherein the processor 1110, the communications interface 1120, and the memory 1130 communicate with each other via the communication bus 1140. The processor 1110 may call logic instructions in the memory 1130 to execute a speech synthesis method, which includes: obtaining a text to be synthesized and an emotional attribute;

[0166] Inputting the text to be synthesized and the emotional attribute into a speech synthesis model to obtain a target speech output by the speech synthesis model;

[0167] In which, the speech synthesis model is trained based on the first sample text corresponding to the first sample speech and the first sample emotional feature corresponding to the first sample speech, the first sample emotional feature is obtained after the first sample speech is input into the emotional coding model, the emotional coding model is trained based on the minimum mutual information loss of the target coding and the second sample emotional feature, the target coding includes the timbre coding of the second sample speech and / or the text content coding of the second sample text of the second sample speech, and the second sample emotional feature is obtained after the second sample speech is input into the initial emotional coding model.

[0168] Furthermore, the logic instructions in the aforementioned memory 1130 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0169] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program, the computer program being storable on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer is capable of performing the speech synthesis method provided by the above methods, the method including: obtaining a text to be synthesized and an emotional attribute;

[0170] Inputting the text to be synthesized and the emotional attribute into a speech synthesis model to obtain a target speech output by the speech synthesis model;

[0171] In which, the speech synthesis model is trained based on the first sample text corresponding to the first sample speech and the first sample emotional feature corresponding to the first sample speech, the first sample emotional feature is obtained after the first sample speech is input into the emotional coding model, the emotional coding model is trained based on the minimum mutual information loss of the target coding and the second sample emotional feature, the target coding includes the timbre coding of the second sample speech and / or the text content coding of the second sample text of the second sample speech, and the second sample emotional feature is obtained after the second sample speech is input into the initial emotional coding model.

[0172] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the computer program is implemented to perform the speech synthesis method provided by the above methods, the method comprising: obtaining a text to be synthesized and an emotional attribute;

[0173] Inputting the text to be synthesized and the emotional attribute into a speech synthesis model to obtain a target speech output by the speech synthesis model;

[0174] In which, the speech synthesis model is trained based on the first sample text corresponding to the first sample speech and the first sample emotional feature corresponding to the first sample speech, the first sample emotional feature is obtained after the first sample speech is input into the emotional coding model, the emotional coding model is trained based on the minimum mutual information loss of the target coding and the second sample emotional feature, the target coding includes the timbre coding of the second sample speech and / or the text content coding of the second sample text of the second sample speech, and the second sample emotional feature is obtained after the second sample speech is input into the initial emotional coding model.

[0175] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0176] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A speech synthesis method, characterized in that: include: Obtain the text to be synthesized and its emotional attributes; Inputting the text to be synthesized and the emotional attribute into a speech synthesis model to obtain a target speech output by the speech synthesis model; The speech synthesis model is trained based on a first sample text corresponding to a first sample speech and a first sample emotion feature corresponding to the first sample speech, the first sample emotion feature is obtained by inputting the first sample speech into an emotion coding model, the emotion coding model is trained based on a target code and a minimum mutual information loss of a second sample emotion feature, the target code includes a timbre code of a second sample speech and / or a text content code of a second sample text of the second sample speech, and the second sample emotion feature is obtained by inputting the second sample speech into an initial emotion coding model; The emotion coding model is trained based on the following method: Inputting the second sample speech into the speech coding network of the initial emotion coding model to obtain semantic features output by the speech coding network; Inputting the preset learnable emotion space and the semantic features into the speech coding network of the initial emotion coding model to obtain the second sample emotion feature output by the emotion coding network; Determining a first minimum mutual information loss between the second sample emotion feature and the timbre encoding, and determining a second minimum mutual information loss between the second sample emotion feature and the text content encoding; Based on the first minimum mutual information loss and the second minimum mutual information loss, the network parameters of the emotion coding network are updated to obtain the emotion coding model; The updating of the network parameters of the emotion coding network based on the first minimum mutual information loss and the second minimum mutual information loss to obtain the emotion coding model includes: Inputting the emotion description text corresponding to the second sample speech into the description text encoding network of the initial emotion encoding model to obtain the emotion description code output by the description text encoding network; determining a contrastive learning loss based on the second sample emotion feature and the emotion description code; Based on the first minimum mutual information loss, the second minimum mutual information loss, and the contrastive learning loss, network parameters of the emotion encoding network and network parameters of the description text encoding network are updated to obtain the emotion encoding model.

2. The speech synthesis method according to claim 1, wherein: Obtaining sentiment attributes includes: When receiving the attribute description text input by the user, inputting the attribute description text into the text macro model to obtain the emotional attribute output by the text macro model; When the emotion template voice input by the user is received, the emotion template voice is input into the emotion coding model to obtain the emotion feature output by the emotion coding model, and the emotion feature is determined as the emotion attribute.

3. The speech synthesis method according to any one of claims 1 to 2, characterized in that: The speech synthesis model is trained based on the following method: Inputting the first sample text and the first sample emotion feature into a speech feature extraction network of an initial speech synthesis model to obtain sample speech features output by the speech feature extraction network; Inputting the sample speech features into the speech decoding network of the initial speech synthesis model to obtain a first predicted speech output by the speech decoding network; Based on the first predicted speech and the first sample speech, the network parameters of the speech feature extraction network and the network parameters of the speech decoding network are updated to obtain the speech synthesis model.

4. The speech synthesis method according to claim 3, wherein: The step of inputting the text to be synthesized and the emotional attribute into a speech synthesis model to obtain a target speech output by the speech synthesis model includes: Inputting the text to be synthesized, the emotional attribute, and other attributes into the speech synthesis model to obtain the target speech output by the speech synthesis model; the other attributes include at least one of the following: target speech environment, target sound quality level, target language, and target speech style; the sample speech feature is obtained by inputting the first sample text, the first sample emotional feature, and other sample attributes corresponding to the first sample speech into the speech feature extraction network; Among them, the other sample attributes include at least one of the following: a sample voice environment corresponding to the first sample voice, a sample sound quality level corresponding to the first sample voice, a sample language corresponding to the first sample voice, and a sample voice style corresponding to the first sample voice.

5. The speech synthesis method according to claim 3, wherein: The updating of network parameters of the speech feature extraction network and network parameters of the speech decoding network based on the first predicted speech and the first sample speech to obtain the speech synthesis model includes: Based on the first predicted speech and the first sample speech, updating network parameters of the speech feature extraction network and network parameters of the speech decoding network to obtain a reference speech synthesis model; Obtaining a third sample text corresponding to a third sample voice and a third sample emotion feature corresponding to the third sample voice, wherein the third sample voice is a voice recorded by a sample subject at a sample location; Inputting the third sample text, the third sample emotional feature, and the identifier of the sample object into the reference speech synthesis model to obtain a second predicted speech output by the reference speech synthesis model; Based on the second predicted speech and the third sample speech, the network parameters of the speech feature extraction network and the network parameters of the speech decoding network in the reference speech synthesis model are updated to obtain the speech synthesis model.

6. A speech synthesis device, characterized in that: include: An acquisition unit, used to acquire the text to be synthesized and the emotional attributes; A synthesis unit, configured to input the text to be synthesized and the emotional attribute into a speech synthesis model to obtain a target speech output by the speech synthesis model; The speech synthesis model is trained based on a first sample text corresponding to a first sample speech and a first sample emotion feature corresponding to the first sample speech, the first sample emotion feature is obtained by inputting the first sample speech into an emotion coding model, the emotion coding model is trained based on a target code and a minimum mutual information loss of a second sample emotion feature, the target code includes a timbre code of a second sample speech and / or a text content code of a second sample text of the second sample speech, and the second sample emotion feature is obtained by inputting the second sample speech into an initial emotion coding model; The emotion coding model is trained based on the following method: Inputting the second sample speech into the speech coding network of the initial emotion coding model to obtain semantic features output by the speech coding network; Inputting the preset learnable emotion space and the semantic features into the speech coding network of the initial emotion coding model to obtain the second sample emotion feature output by the emotion coding network; Determining a first minimum mutual information loss between the second sample emotion feature and the timbre encoding, and determining a second minimum mutual information loss between the second sample emotion feature and the text content encoding; Based on the first minimum mutual information loss and the second minimum mutual information loss, the network parameters of the emotion coding network are updated to obtain the emotion coding model; The updating of the network parameters of the emotion coding network based on the first minimum mutual information loss and the second minimum mutual information loss to obtain the emotion coding model includes: Inputting the emotion description text corresponding to the second sample speech into the description text encoding network of the initial emotion encoding model to obtain the emotion description code output by the description text encoding network; determining a contrastive learning loss based on the second sample emotion feature and the emotion description code; Based on the first minimum mutual information loss, the second minimum mutual information loss, and the contrastive learning loss, network parameters of the emotion encoding network and network parameters of the description text encoding network are updated to obtain the emotion encoding model.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the speech synthesis method according to any one of claims 1 to 5 is implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the speech synthesis method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Speech synthesis method and device, readable medium and electronic equipment

    CN119028313A

  • Multi-style audio synthesis method, apparatus and device, and storage medium

    WO2022116432A1