Method, apparatus, and electronic device for determining audio datasets

By selecting voice fragments based on emotion types and intensities using acoustic and semantic understanding models, the method enriches voice datasets, improving the emotional expressiveness of synthesized speech.

JP2026053485APending Publication Date: 2026-03-25BAIDU INT TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Voice synthesis models struggle to express a wide range of emotions due to voice fragments in training datasets originating from monotonous scenes like readings and news broadcasts, leading to limited emotional expressiveness.

Method used

A method and apparatus for determining a voice dataset by selecting voice fragments based on emotion types and intensities, using acoustic and semantic understanding models to ensure emotional consistency and richness, followed by training a speech synthesis model on this enhanced dataset.

Benefits of technology

The approach enhances the emotional expressiveness of synthesized speech by ensuring emotional truthfulness and consistency, enabling more genuine emotional expression in synthesized voices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026053485000001_ABST
    Figure 2026053485000001_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, and electronic device for determining a speech dataset for training a speech synthesis model that can synthesize speech that expresses genuine emotions. [Solution] The method involves obtaining a set of audio fragments and text corresponding to the audio fragments in the set, determining the first emotion type of the audio fragment, the emotion intensity of the first emotion type, the second emotion type of the text, and the emotion intensity of the second emotion type, further determining whether or not to retain the audio fragments, determining an audio dataset based on the retained audio fragments in the set and the text corresponding to the retained audio fragments, and determining whether or not to retain the audio fragments by combining the emotion intensities of the emotion types, thereby ensuring the emotion intensity of the emotion types of the audio fragments in the audio dataset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to deep learning, natural language processing, voice technology, large-scale models, and particularly to a method, apparatus, and electronic device for determining a voice dataset.

Background Art

[0002] Currently, the voice fragments in the voice dataset on which a voice synthesis model depends during training mainly originate from scenes such as reading aloud and news broadcasts. The voice fragments obtained from the above scenes have relatively single emotions and little variation. Therefore, it is difficult for the emotional expressiveness of the voice synthesized by the voice synthesis model to meet the needs.

Summary of the Invention

[0003] The present disclosure provides a method, apparatus, and electronic device for determining a voice dataset.

Means for Solving the Problems

[0004] According to one aspect of the present disclosure, a method for determining a voice dataset is provided. The method includes: obtaining a voice fragment set and text corresponding to the voice fragments in the voice fragment set; determining a first emotion type of the voice fragment, an emotion intensity of the first emotion type, a second emotion type of the text, and an emotion intensity of the second emotion type; determining whether to retain the voice fragment based on the first emotion type, the emotion intensity of the first emotion type, the second emotion type, and the emotion intensity of the second emotion type; and determining a voice dataset based on the retained voice fragments in the voice fragment set and the text corresponding to the retained voice fragments.

[0005] Another aspect of the present disclosure provides a method for training a speech synthesis model, the method comprising: acquiring a speech dataset, the speech dataset including speech fragments, emotion types of the speech fragments, and text corresponding to the speech fragments, wherein the speech dataset is determined based on the method for determining the speech dataset described above; acquiring an initial speech synthesis model; and performing a training process on the initial speech synthesis model based on the speech fragments, emotion types of the speech fragments, and text to acquire a trained speech synthesis model.

[0006] According to another aspect of the present disclosure, a determination device for an audio dataset is provided, the device including: an acquisition module for acquiring an audio fragment set and text corresponding to an audio fragment in the audio fragment set; a first determination module for determining a first emotion type of the audio fragment, the emotion intensity of the first emotion type, a second emotion type of the text, and the emotion intensity of the second emotion type; a second determination module for determining whether or not to retain the audio fragment based on the first emotion type, the emotion intensity of the first emotion type, the second emotion type, and the emotion intensity of the second emotion type; and a third determination module for determining an audio dataset based on retained audio fragments in the audio fragment set and text corresponding to the retained audio fragment.

[0007] According to another aspect of the present disclosure, a training apparatus for a speech synthesis model is provided, the apparatus comprising: a first acquisition module for acquiring a speech dataset, wherein the speech dataset includes speech fragments, emotion types of the speech fragments, and text corresponding to the speech fragments, and the speech dataset is determined based on the above method for determining the speech dataset; a second acquisition module for acquiring an initial speech synthesis model; and a training processing module for performing a training process on the initial speech synthesis model based on the speech fragments, the emotion types of the speech fragments, and the text to acquire a trained speech synthesis model.

[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising at least one processor and a memory communicably connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform a method for determining a speech dataset provided by the present disclosure or a method for training a speech synthesis model provided by the present disclosure.

[0009] According to another aspect of the present disclosure, a non-temporary computer-readable storage medium is provided which stores computer instructions, the computer instructions causing a computer to perform a method for training a speech synthesis model provided in the present disclosure, or the method for training a speech synthesis model provided in the present disclosure.

[0010] According to another aspect of the present disclosure, a computer program is provided which, when executed by a processor, implements steps of a method for training a speech synthesis model provided by the present disclosure, or implements steps of a method for training a speech synthesis model provided by the present disclosure.

[0011] Please understand that the content described in this section is not intended to identify any essential or important features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will be readily apparent from the following description. [Brief explanation of the drawing]

[0012] The drawings are provided for the purpose of better understanding this technical proposal and do not limit the scope of this disclosure. [Figure 1] This is a schematic diagram relating to the first embodiment of the present disclosure. [Figure 2] This is a schematic diagram relating to a second embodiment of the present disclosure. [Figure 3] This is a schematic diagram relating to a third embodiment of the present disclosure. [Figure 4] This is a schematic diagram for determining the audio dataset. [Figure 5] This is a schematic diagram relating to a fourth embodiment of the present disclosure. [Figure 6] This is a schematic diagram relating to the fifth embodiment of the present disclosure. [Figure 7] This is a block diagram of an electronic device for implementing a method for determining an audio dataset or a method for training a speech synthesis model according to an embodiment of the present disclosure. [Modes for carrying out the invention]

[0013] Hereinafter, exemplary embodiments of the present disclosure will be described in conjunction with the drawings, and various details of the embodiments of the present disclosure will be included for the sake of clarity, and these should be considered as illustrative only. Accordingly, those skilled in the art should be aware that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, well-known functions and structures will be omitted in the following description.

[0014] Currently, the speech fragments in speech datasets used to train speech synthesis models primarily originate from scenes such as readings and news broadcasts. Because the speech segments obtained from these scenes exhibit relatively monolithic and limited emotional variation, the emotional expressiveness of the synthesized speech by the speech synthesis model often fails to meet the needs.

[0015] Figure 1 is a schematic diagram relating to a first embodiment of the present disclosure, the method for determining an audio dataset of the embodiment of the present disclosure can be applied to an audio dataset determination device, which can be configured as an electronic device, enabling the electronic device to perform the audio dataset determination function.

[0016] Here, the electronic device may be a device having computing capabilities, such as a personal computer (PC), a mobile terminal, or a server. The mobile terminal may be a hardware device having various operating systems, touch panels, and / or displays, such as an in-car device, a mobile phone, a tablet, a personal digital assistant, a wearable device, a smart speaker, a server, or a server cluster.

[0017] The device for determining the audio dataset may be software on an electronic device, such as software for determining the audio dataset. In the following embodiment, the execution entity is an electronic device as an example.

[0018] As shown in Figure 1, the method for determining this audio dataset may include the following steps.

[0019] In Step 101, the audio fragment set and the text corresponding to the audio fragments in the audio fragment set are obtained.

[0020] In an embodiment of the present disclosure, the voice fragment may be a voice fragment of a single speaker with interference such as background and noise removed. The source of the voice fragment may be, for example, the voice of any object, the voice of any video, etc., which is not specifically limited here and can be set according to actual needs.

[0021] The language type of the voice fragment may be any language type or the target language type. The language type may be, for example, Chinese, English, etc., which is not specifically limited here.

[0022] In step 102, determine the first emotion type of the voice fragment, the emotion intensity of the first emotion type, the second emotion type of the text, and the emotion intensity of the second emotion type.

[0023] In an embodiment of the present disclosure, the process by which the electronic device executes step 102 may be, for example, as follows. Input the voice fragment into an acoustic feature model, and obtain at least one first candidate emotion type, the first probability of the first candidate emotion type, and the emotion intensity output from the acoustic feature model. Select the first emotion type from at least one first candidate emotion type based on the first probability. Input the voice fragment into a semantic understanding model, and obtain at least one second candidate emotion type, the second probability of the second candidate emotion type, and the emotion intensity output from the semantic understanding model. Select the second emotion type from at least one second candidate emotion type based on the second probability.

[0024] In an embodiment of the present disclosure, the number of acoustic feature models may be plural. The model structures of the plural acoustic feature models may be different. The acoustic feature model can evaluate the emotion tension at the expression level of the voice fragment, that is, the emotion intensity of the emotion type, by capturing the rhythm variation, tone dynamics, and energy change in the voice fragment.

[0025] The emotional intensity of an emotion type can be determined and obtained based on rhythmic variations, tone dynamics, and energy changes in a speech fragment. Rhythmic variations can be represented, for example, using the maximum difference in fundamental frequencies within the fundamental frequency dynamic range of the speech fragment. A larger maximum difference indicates greater rhythmic variation in the speech fragment. Tone dynamics can be represented, for example, using the standard deviation of the fundamental frequencies of the speech fragment. A larger standard deviation of fundamental frequencies indicates greater pitch variation in the speech fragment. Energy changes can be represented using the standard deviation of the energy of the speech fragment. A larger standard deviation of energy indicates greater amplitude variation in the speech fragment.

[0026] The emotional intensity of an emotion type can be positively correlated with the numerical values ​​of three parameters in a speech fragment: the maximum difference in fundamental frequencies within the fundamental frequency dynamic range, the standard deviation of fundamental frequencies, and the standard deviation of energy. The emotional intensity of an emotion type can be obtained by weighting and summing the numerical values ​​of at least two of the above three parameters of the speech fragment.

[0027] Examples of acoustic feature models include the emotion vector model (Emotion2Vec), the extended channel attention time delay neural network (ECAPA-TDNN), and the ResNet-based speech emotion recognition model (ResNet-SER). These are not specifically limited here and can be configured according to actual needs.

[0028] By combining the first emotion type and the emotion intensity of the first emotion type output from the acoustic feature model, it is possible to determine whether or not to retain the audio fragment, select audio fragments with high emotion intensity, and extend the emotional truthfulness of audio fragments in the audio dataset.

[0029] There may be multiple semantic understanding models. The model structures of these multiple semantic understanding models may differ. The semantic understanding models can extract emotional intentions contained in the text and evaluate the emotional types and the emotional intensity of those types in the text.

[0030] Examples of semantic understanding models include the BERT+Text-Classifier, which combines BERT, and the SenseVoice model. We won't specifically limit ourselves here, as these can be configured according to actual needs.

[0031] By combining the second emotion type and the emotion intensity of the second emotion type output from the semantic understanding model, it is possible to determine whether or not to retain the audio fragment, and by selecting audio fragments with high emotion intensity, i.e., audio fragments with strong emotional tension, to determine the audio dataset, the emotional truthfulness of audio fragments in the audio dataset can be extended.

[0032] In step 103, a decision is made whether or not to retain the audio fragment based on the first emotion type, the emotional intensity of the first emotion type, the second emotion type, and the emotional intensity of the second emotion type.

[0033] In embodiments of the present disclosure, for example, the process by which the electronic device performs step 103 may be as follows: determine whether the first emotion type and the second emotion type match; determine whether the emotion intensity of the first emotion type and the emotion intensity of the second emotion type are equal to or greater than an emotion intensity threshold; if the first emotion type and the second emotion type match and the emotion intensity of both the first emotion type and the emotion intensity of the second emotion type are equal to or greater than an emotion intensity threshold, decide to retain the audio fragment; if the first emotion type and the second emotion type do not match, or if the emotion intensity of the first emotion type is less than an emotion intensity threshold, or if the emotion intensity of the second emotion type is less than an emotion intensity threshold, decide not to retain the audio fragment.

[0034] If the first emotion type and the second emotion type do not match, it indicates that the audio fragment is not using the correct emotion type to express the text. If the emotion intensity of the first emotion type is lower than the emotion intensity threshold, it indicates that the audio fragment has poor emotional expressiveness.

[0035] The system determines whether the first emotion type and the second emotion type match, whether the emotional intensity of the first emotion type and the emotional intensity of the second emotion type are above an emotional intensity threshold, and further determines whether to retain the audio fragment. This ensures emotional consistency between the retained audio fragment and the corresponding text.

[0036] As a candidate for the above technical proposal, the electronic device may decide to retain the audio fragment if the emotional intensity of both the first emotion type and the second emotion type are equal to or greater than the emotional intensity threshold, and decide not to retain the audio fragment if the emotional intensity of the first emotion type is less than the emotional intensity threshold, or if the emotional intensity of the second emotion type is less than the emotional intensity threshold.

[0037] In embodiments of this disclosure, in other embodiments, there may be multiple first emotion types and multiple second emotion types. Correspondingly, the process by which the electronic device performs step 103 may be as follows, for example: if multiple first emotion types match, the total emotion intensity of the first emotion types is determined based on the emotion intensity of the first emotion types and the weights of the acoustic feature model for determining the first emotion types; if multiple second emotion types match, the total emotion intensity of the second emotion types is determined based on the emotion intensity of the second emotion types and the weights of the semantic understanding model for determining the second emotion types; and based on the first emotion types, the total emotion intensity of the first emotion types, the second emotion types, and the total emotion intensity of the second emotion types, the device decides whether or not to retain the audio fragment; and if multiple first emotion types do not match, or if multiple second emotion types do not match, the device decides not to retain the audio fragment.

[0038] The weights of an acoustic feature model can be determined based on the decision accuracy of the acoustic feature model in terms of emotion type, and the decision accuracy and weights can be positively correlated. The weights of a semantic understanding model can be determined based on the decision accuracy of the semantic understanding model in terms of emotion type, and the decision accuracy and weights can be positively correlated.

[0039] For example, the process by which an electronic device determines whether or not to retain a voice fragment based on a first emotion type, the total emotion intensity of the first emotion type, a second emotion type, and the total emotion intensity of the second emotion type may be as follows: if the first emotion type and the second emotion type match, and both the total emotion intensity of the first emotion type and the total emotion intensity of the second emotion type are above the emotion intensity threshold, the device decides to retain the voice fragment; if the first emotion type and the second emotion type do not match, or if the total emotion intensity of the first emotion type is below the emotion intensity threshold, or if the total emotion intensity of the second emotion type is below the emotion intensity threshold, the device decides not to retain the voice fragment.

[0040] In another example, the process by which an electronic device determines whether or not to retain a voice fragment based on a first emotion type, the total emotion intensity of the first emotion type, a second emotion type, and the total emotion intensity of the second emotion type may be as follows: if the total emotion intensity of both the first emotion type and the second emotion type are equal to or greater than the emotion intensity threshold, it is decided to retain the voice fragment; if the total emotion intensity of the first emotion type is less than the emotion intensity threshold, or if the total emotion intensity of the second emotion type is less than the emotion intensity threshold, it is decided not to retain the voice fragment.

[0041] The total emotional intensity of a first emotion type is determined by combining the weights of an acoustic feature model, the total emotional intensity of a second emotion type is determined by combining the weights of a semantic understanding model, it is determined whether the first emotion type and the second emotion type match, it is determined whether the total emotional intensity of the first emotion type and the total emotional intensity of the second emotion type are above an emotional intensity threshold, and it is further determined whether or not to retain the audio fragment. This ensures emotional consistency between the retained audio fragment and the corresponding text.

[0042] In step 104, the audio dataset is determined based on the retained audio fragments in the audio fragment set and the text corresponding to those fragments.

[0043] In embodiments of the present disclosure, to facilitate training for a controllable speech synthesis model of subsequent emotion types, the electronic device can determine a speech dataset based on the retained speech fragments in the speech fragment set, the text corresponding to the retained speech fragments, and one of a first emotion type and a second emotion type.

[0044] In embodiments of this disclosure, after step 104, the electronic device may also perform the following process: fine-tune training on an acoustic feature model and a semantic understanding model based on speech fragments of multiple languages ​​and the text corresponding to the speech fragments in a speech dataset.

[0045] By performing fine-tuning training on acoustic feature models and semantic understanding models based on speech fragments in multiple languages ​​and the text corresponding to those speech fragments in a speech dataset, the accuracy of the acoustic feature models and semantic understanding models can be further improved, thereby further improving the accuracy of the subsequent determined speech dataset.

[0046] In the method for determining the audio dataset of the embodiments of this disclosure, an audio fragment set and text corresponding to the audio fragments in the audio fragment set are obtained, a first emotion type and emotion intensity of the first emotion type of the audio fragments are determined, a second emotion type and emotion intensity of the second emotion type of the text are determined, it is determined whether or not to retain the audio fragments based on the first emotion type, emotion intensity of the first emotion type, second emotion type and emotion intensity of the second emotion type, an audio dataset is determined based on the retained audio fragments and the text corresponding to the retained audio fragments in the audio fragment set, and it is determined whether or not to retain the audio fragments by combining the emotion intensities of the emotion types, thereby ensuring the emotion intensity of the emotion types of the audio fragments in the audio dataset, and a speech synthesis model trained on this audio dataset can synthesize speech that expresses genuine emotions.

[0047] To further enhance the richness of the emotional types of the audio fragments, arbitrary audio can be processed to obtain a set of audio fragments. As shown in Figure 2, Figure 2 is a schematic diagram relating to a second embodiment of the present disclosure, and the embodiment shown in Figure 2 may include the following steps.

[0048] In step 201, multiple audio files to be processed are acquired.

[0049] In step 202, interference cancellation and speaker separation are performed on each audio to obtain the speaker voice of at least one single speaker.

[0050] In embodiments of this disclosure, the process by which the electronic device performs step 202 may be as follows, for example: inputting speech into a sound source separation model, obtaining the separated speech output from the sound source separation model, inputting the separated speech into a speaker separation model, and obtaining each speaker speech output from the speaker separation model.

[0051] The sound source separation model is used to separate background noise and other interfering sounds from the speech, thereby obtaining the separated speech, i.e., the speech from which the interfering sounds have been removed. Each speaker's voice is the voice of a single speaker.

[0052] In embodiments of this disclosure, a format unification process can be performed on each audio before step 202 in order to improve the interference rejection efficiency and speaker separation efficiency of each audio. For example, a unification process can be performed on at least one of the following for each audio: sampling rate, channel information, volume information, etc. The unified sampling rate may be, for example, 24 kHz. The unified volume information may be, for example, -20 dBFS.

[0053] In step 203, speech activity detection and segmentation processing are performed on the speaker voice of at least one single speaker to obtain speaker fragments of at least one single speaker.

[0054] In embodiments of the present disclosure, the process by which the electronic device performs step 203 may be as follows, for example, that for each single speaker's speaker voice, the speaker voice is input to a speech activity detection model, a segmentation result is obtained from the speech activity detection model, and the segmentation result includes each speaker fragment of the single speaker. A speaker fragment is a fragment of speech that includes a single speaker.

[0055] In step 204, a set of speech fragments is determined based on speaker fragments from at least one single speaker.

[0056] In embodiments of this disclosure, in order to further ensure that the speaker voice is the voice of a single speaker, the electronic device may perform the following steps prior to step 204: perform a duplicate speaker detection process on the speaker voice of a single speaker to determine whether or not there are duplicate fragments in the speaker voice; and if there are duplicate fragments in the speaker voice, perform a duplicate fragment removal process on the speaker voice.

[0057] An electronic device can input speaker speech into a duplicate detection model, retrieve duplicate fragments output from the model, determine the starting point of the duplicate fragments, and remove the duplicate fragments from the speaker speech based on the starting point.

[0058] In embodiments of this disclosure, in order to further improve the quality of the audio fragments in the audio fragment set, the electronic device may also perform the following steps before step 204: perform sound quality scoring on the speaker fragments of a single speaker to obtain a sound quality score value for the speaker fragments, and perform filtering control processing on the speaker fragments based on the sound quality score value and the duration of the speaker fragments.

[0059] The process by which an electronic device performs filtering control processing on speaker fragments based on a sound quality score and the duration of the speaker fragment may be as follows: if the sound quality score is below a score threshold, or if the duration is outside the duration range, filtering processing is performed on the speaker fragment.

[0060] If the duration is outside the specified duration range, it may be either too short or too long. If the duration of the speaker fragment is too short, there may be problems in detecting the emotion type, or the detected emotion type may be inaccurate. If the duration of the speaker fragment is too long, there may be problems in detecting the emotion type because it is not fixed. Sound quality evaluation is used to assess scores such as the intelligibility of the speaker fragment.

[0061] If the sound quality score is below the score threshold, or if the duration is outside the duration range, filtering can be applied to the speaker fragment to further improve the quality of the retained speaker fragment and thus the overall quality of the speech fragments in the speech fragment set.

[0062] In Step 205, speech recognition processing is performed on the audio fragments in the audio fragment set to obtain the text corresponding to the audio fragments.

[0063] In embodiments of the present disclosure, the process by which the electronic device performs step 205 may be as follows: determine the language type to which the speech fragment belongs; obtain a target speech recognition model by selecting from a plurality of candidate speech recognition models based on the language type; input the speech fragment into the target speech recognition model; and obtain the text output from the target speech recognition model.

[0064] Different candidate speech recognition models have different accuracy in processing speech fragments of different language types; in other words, different candidate speech recognition models are better suited to processing speech fragments of different language types. Therefore, by selecting a target speech recognition model from among multiple candidate models based on the language type and performing speech recognition processing on the speech fragment, the accuracy of the determined and acquired text can be further improved.

[0065] In step 206, the first emotion type of the audio fragment, the emotion intensity of the first emotion type, the second emotion type of the text, and the emotion intensity of the second emotion type are determined.

[0066] In step 207, a decision is made whether or not to retain the audio fragment based on the first emotion type, the emotional intensity of the first emotion type, the second emotion type, and the emotional intensity of the second emotion type.

[0067] In step 208, the audio dataset is determined based on the retained audio fragments in the audio fragment set and the text corresponding to those fragments.

[0068] For detailed information on steps 206 to 208, please refer to steps 102 to 104 in the embodiment shown in Figure 1; a detailed explanation is omitted here.

[0069] In the method for determining the audio dataset of the embodiments of this disclosure, multiple audio files to be processed are acquired, interference removal processing and speaker separation processing are performed on each audio file to acquire speaker audio of at least one single speaker, audio activity detection and segmentation processing are performed on speaker audio of at least one single speaker to acquire speaker fragments of at least one single speaker, an audio fragment set is determined based on the speaker fragments of at least one single speaker, speech recognition processing is performed on the audio fragments in the audio fragment set to acquire text corresponding to the audio fragments, and the first emotion type of the audio fragment, the emotion intensity of the first emotion type, the second emotion type of the text, and the second emotion are determined. The system determines the emotional intensity of each type, decides whether or not to retain an audio fragment based on the first emotional type, the emotional intensity of the first emotional type, the second emotional type, and the emotional intensity of the second emotional type, determines the audio dataset based on the retained audio fragments in the audio fragment set and the text corresponding to the retained audio fragments, performs interference removal, speaker separation, audio activity detection, and segmentation on the audio to be processed to obtain audio fragments, and the audio to be processed can be any audio, thereby improving the richness of the emotional types of the audio fragments and improving the richness of the emotional types and the truthfulness of the emotions of the audio fragments in the audio dataset.

[0070] Figure 3 is a schematic diagram relating to a third embodiment of the present disclosure. The speech synthesis model training method of the embodiment of the present disclosure can be applied to a speech synthesis model training device, which can be configured as an electronic device, enabling this electronic device to perform the speech synthesis model training function.

[0071] Here, the electronic device may be any device having computing capabilities, such as a personal computer (PC), a mobile terminal, or a server. The mobile terminal may be a hardware device having various operating systems, touch panels, and / or displays, such as an in-car device, a mobile phone, a tablet, a personal digital assistant, a wearable device, a smart speaker, a server, or a server cluster.

[0072] The training device for a speech synthesis model may be software on an electronic device, such as training software for a speech synthesis model. In the following embodiment, the execution entity is an electronic device as an example.

[0073] As shown in Figure 3, the training method for this speech synthesis model can include the following steps.

[0074] In step 301, an audio dataset is acquired, which includes audio fragments, the emotion type of the audio fragments, and the text corresponding to the audio fragments, and the audio dataset is determined based on the method for determining the audio dataset shown in Figure 1 or Figure 2.

[0075] The audio fragments in the audio dataset may be audio fragments obtained by performing interference removal, speaker separation, audio activity detection, and segmentation on arbitrary audio.

[0076] In step 302, an initial speech synthesis model is obtained.

[0077] The initial speech synthesis model may be a pre-trained speech synthesis model.

[0078] In step 303, the initial speech synthesis model is trained based on the speech fragment, the emotion type of the speech fragment, and the text to obtain the trained speech synthesis model.

[0079] In embodiments of this disclosure, the process by which the electronic device performs step 303 may be, for example, as follows: inputting text and emotion type into a speech synthesis model; obtaining predicted speech fragments output from the speech synthesis model; determining loss function values ​​based on the predicted speech fragments, speech fragments corresponding to the text, and the loss function of the speech synthesis model; performing parameter tuning on the speech synthesis model based on the loss function values ​​to obtain a trained speech synthesis model.

[0080] The trained speech synthesis model can be used for speech synthesis processing in at least one scene, such as a human-machine interaction scene, a dubbing scene, a script interpretation scene, and an anchor scene, but is not specifically limited thereto. Human-machine interaction scenes include, but are not specifically limited thereto, virtual human interaction, digital character interaction, customer service interaction, educational platform interaction, psychological counseling, emotional companion, and elderly assistant interaction scenes.

[0081] In the speech synthesis model training method of the embodiments of this disclosure, an audio dataset is acquired, the audio dataset includes audio fragments, the emotion type of the audio fragments, and the text corresponding to the audio fragments, the audio dataset is determined based on the method for determining the audio dataset shown in Figure 1 or Figure 2, an initial speech synthesis model is acquired, the initial speech synthesis model is trained based on the audio fragments, the emotion type of the audio fragments, and the text, and a trained speech synthesis model is acquired, the emotion intensity of the emotion type of the audio fragments in the audio dataset is high, and this increases the truthfulness of the emotion of the audio fragments synthesized and acquired by the speech synthesis model.

[0082] An example is given below. Figure 4 is a schematic diagram for determining the audio dataset. Figure 4 includes the following steps 401 to 406.

[0083] In step 401, the audio to be processed (i.e., the original audio data in Figure 4) is obtained.

[0084] In step 402, a format unification process is performed on the audio (i.e., standardization as shown in Figure 4).

[0085] In step 403, interference cancellation and audio activity detection are performed on the audio to obtain multiple audio fragments.

[0086] In Step 404, speaker separation is performed on the audio fragments to obtain the audio fragments (i.e., clean audio) from the audio fragment set.

[0087] In step 405, determine at least one emotion type (i.e., emotion label) and emotion intensity of the emotion type for each audio fragment in the audio fragment set.

[0088] In step 406, a decision is made whether or not to retain the voice fragment based on at least one emotion type and the emotion intensity of the emotion type (i.e., whether or not to retain the voice fragment by fusing the intensity), and the retained voice fragment (i.e., the emotion voice) is then obtained.

[0089] To realize the above embodiments, the present disclosure further provides a device for determining an audio dataset. As shown in Figure 5, Figure 5 is a schematic diagram relating to a fourth embodiment of the present disclosure. This audio dataset determination device 50 may include an acquisition module 501, a first determination module 502, a second determination module 503, and a third determination module 504.

[0090] Acquisition module 501 is used to acquire a set of audio fragments and the text corresponding to the audio fragments in the set of audio fragments; first decision module 502 is used to determine the first emotion type of the audio fragments, the emotion intensity of the first emotion type, the second emotion type of the text, and the emotion intensity of the second emotion type; second decision module 503 is used to determine whether or not to retain the audio fragments based on the first emotion type, the emotion intensity of the first emotion type, the second emotion type, and the emotion intensity of the second emotion type; and third decision module 504 is used to determine the audio dataset based on the retained audio fragments in the set of audio fragments and the text corresponding to the retained audio fragments.

[0091] In one possible embodiment of the embodiments of the present disclosure, the first decision module 502 is used to input the speech fragment into an acoustic feature model to obtain at least one first candidate emotion type, a first probability of the first candidate emotion type, and an emotion intensity output from the acoustic feature model, and to select the first emotion type from the at least one first candidate emotion type based on the first probability; to input the speech fragment into a semantic understanding model to obtain at least one second candidate emotion type, a second probability of the second candidate emotion type, and an emotion intensity output from the semantic understanding model, and to select the second emotion type from the at least one second candidate emotion type based on the second probability.

[0092] In one possible embodiment of the embodiments of the present disclosure, the second determination module 503 is used to determine whether the first emotion type and the second emotion type match, whether the emotion intensity of the first emotion type and the emotion intensity of the second emotion type are equal to or greater than an emotion intensity threshold, and to determine to retain the audio fragment if the first emotion type and the second emotion type match and the emotion intensity of both the first emotion type and the emotion intensity of the second emotion type are equal to or greater than the emotion intensity threshold.

[0093] In one possible embodiment of the embodiments of the present disclosure, the second decision module 503 is specifically used to determine not to retain the audio fragment if the first emotion type and the second emotion type do not match, or if the emotion intensity of the first emotion type is less than the emotion intensity threshold, or if the emotion intensity of the second emotion type is less than the emotion intensity threshold.

[0094] One possible embodiment of the embodiments of the present disclosure is a plurality of first emotion types and a plurality of second emotion types, wherein the second decision module 503 specifically determines the total emotion intensity of the first emotion types based on the emotion intensity of the first emotion types and the weights of the acoustic feature model for determining the first emotion types when a plurality of the first emotion types match, and determines the total emotion intensity of the second emotion types based on the emotion intensity of the second emotion types and the weights of the semantic understanding model for determining the second emotion types when a plurality of the second emotion types match, and is used to determine whether or not to retain the audio fragment based on the first emotion types, the total emotion intensity of the first emotion types, the second emotion types, and the total emotion intensity of the second emotion types.

[0095] In one possible embodiment of the embodiments of the present disclosure, the second decision module 503 is specifically used to determine not to retain the voice fragment if a further number of the first emotion types do not match, or if a further number of the second emotion types do not match.

[0096] In one possible embodiment of the embodiments of the present disclosure, the acquisition module 501 includes an acquisition unit, a first processing unit, a second processing unit, a determination unit, and a recognition processing unit, wherein the acquisition unit is used to acquire a plurality of audio to be processed; the first processing unit is used to acquire speaker audio of at least one single speaker by performing interference removal processing and speaker separation processing on each audio; the second processing unit is used to acquire speaker fragments of at least one single speaker by performing speech activity detection and segmentation processing on the speaker audio of at least one single speaker; the determination unit is used to determine the set of audio fragments based on the speaker fragments of at least one single speaker; and the recognition processing unit is used to acquire text corresponding to the audio fragments by performing speech recognition processing on the audio fragments in the set of audio fragments.

[0097] In one possible embodiment of the embodiments of the present disclosure, the acquisition module 501 further includes a detection processing unit and a removal processing unit, wherein the detection processing unit is used to perform duplicate speaker detection processing on a single speaker's voice and to determine whether or not duplicate fragments exist in the speaker's voice, and the removal processing unit is used to perform a removal processing of the duplicate fragments on the speaker's voice if duplicate fragments exist in the speaker's voice.

[0098] In one possible embodiment of the embodiments of the present disclosure, the acquisition module 501 further includes a scoring processing unit and a control processing unit, wherein the scoring processing unit is used to perform sound quality scoring processing on a speaker fragment of a single speaker to obtain a sound quality score value for the speaker fragment, and the control processing unit is used to perform filtering control processing on the speaker fragment based on the sound quality score value and the duration of the speaker fragment.

[0099] In one possible embodiment of the embodiments of the present disclosure, the control processing unit is specifically used to perform filtering on the speaker fragment when the sound quality score value is less than or equal to a score threshold, or when the time length is not within a time length range.

[0100] In one possible embodiment of the embodiments of the present disclosure, the recognition processing unit is used to determine the language type to which the speech fragment belongs, to select a target speech recognition model from among a plurality of candidate speech recognition models based on the language type, to input the speech fragment into the target speech recognition model, and to obtain the text output from the target speech recognition model.

[0101] In one possible embodiment of the embodiments of the present disclosure, the audio dataset includes audio fragments in multiple languages, and the device further includes a training processing module for performing fine-tuning training on the acoustic feature model and the semantic understanding model based on the audio fragments in the multiple languages ​​in the audio dataset and the text corresponding to the audio fragments.

[0102] In the speech dataset determination device of the embodiment of this disclosure, a speech fragment set and text corresponding to the speech fragments in the speech fragment set are acquired, a first emotion type and emotion intensity of the first emotion type of the speech fragments are determined, a second emotion type and emotion intensity of the second emotion type of the text are determined, a decision is made whether or not to retain the speech fragments based on the first emotion type, emotion intensity of the first emotion type, second emotion type and emotion intensity of the second emotion type is determined, a speech dataset is determined based on the retained speech fragments in the speech fragment set and the text corresponding to the retained speech fragments, and a decision is made whether or not to retain the speech fragments by combining the emotion intensities of the emotion types, thereby ensuring the emotion intensity of the emotion types of the speech fragments in the speech dataset, and speech that expresses true emotions can be synthesized by a speech synthesis model trained on this speech dataset.

[0103] To realize the above embodiments, the present disclosure further provides a training device for a speech synthesis model. As shown in Figure 6, Figure 6 is a schematic diagram relating to a fifth embodiment of the present disclosure. This speech synthesis model training device 60 may include a first acquisition module 601, a second acquisition module 602, and a training processing module 603.

[0104] The first acquisition module 601 is used to acquire an audio dataset, the audio dataset including an audio fragment, the emotion type of the audio fragment, and the text corresponding to the audio fragment, and the audio dataset is determined based on the method for determining the audio dataset shown in Figure 1 or Figure 2.

[0105] The second acquisition module 602 is used to acquire the initial speech synthesis model.

[0106] The training processing module 603 is used to train the initial speech synthesis model based on the speech fragment, the emotion type of the speech fragment, and the text, in order to obtain a trained speech synthesis model.

[0107] In the speech synthesis model training apparatus of the embodiments of this disclosure, a speech dataset is acquired, the speech dataset includes speech fragments, the emotion type of the speech fragments, and the text corresponding to the speech fragments, the speech dataset is determined based on the method for determining the speech dataset shown in Figure 1 or Figure 2, an initial speech synthesis model is acquired, the initial speech synthesis model is trained based on the speech fragments, the emotion type of the speech fragments, and the text, and a trained speech synthesis model is acquired, the emotion intensity of the emotion type of the speech fragments in the speech dataset is high, and this increases the truthfulness of the emotion of the speech fragments synthesized and acquired by the speech synthesis model.

[0108] Furthermore, in this proposed technology disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of relevant user personal information will all be carried out with the user's consent, in accordance with the provisions of relevant laws and regulations, and will not violate public order and morals.

[0109] According to embodiments of the present disclosure, the present disclosure further provides electronic devices, readable storage media, and computer programs.

[0110] Figure 7 is a schematic block diagram of an exemplary electronic device 700 for carrying out embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the description herein and / or the implementation of the present disclosure as requested.

[0111] As shown in Figure 7, the electronic device 700 includes a computing unit 701 capable of performing various appropriate operations and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 may also store various programs and data necessary for the operation of the electronic device 700. The computing unit 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0112] Multiple components of the electronic device 700 are connected to an I / O interface 708, which includes an input unit 706 such as a keyboard and mouse, an output unit 707 such as various types of displays and speakers, a storage unit 705 such as a magnetic disk and an optical disk, and a communication unit 709 such as a network card, modem, and wireless communication transceiver. The communication unit 709 enables the electronic device 700 to exchange information / data with other devices via computer networks such as the Internet and / or various telecom networks.

[0113] The computing unit 701 may be a variety of general-purpose and / or dedicated processing components having processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, computing units that execute various machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs each of the methods and processes described in the preamble, for example, a method for determining a speech dataset or a method for training a speech synthesis model. For example, in some embodiments, a method for determining a speech dataset or a method for training a speech synthesis model can be implemented as a computer software program tangibly contained in a machine-readable medium such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed into the electronic device 700 via ROM 702 and / or a communication unit 709. When the computer program is loaded into RAM 703 and executed by the computing unit 701, one or more steps of the method for determining a speech dataset or a method for training a speech synthesis model described in the preamble are performed. Selectively, in other embodiments, the computing unit 701 may be configured by any other suitable means (e.g., via firmware) to perform a method for determining a speech dataset or a method for training a speech synthesis model.

[0114] Various embodiments of the systems and technologies described above in this specification can be implemented as digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include being implemented by one or more computer programs, which may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be an application-specific or general-purpose programmable processor, which may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0115] Program code for performing the methods of this disclosure can be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing device, so that when executed by the processor or controller, the functions / operations defined in the flowcharts and / or block diagrams are performed. The program code may run entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine, partially on a remote machine, or entirely on a remote machine or server.

[0116] In the context of this disclosure, a machine-readable medium may be a tangible medium that contains or can store a program used by an instruction execution system, apparatus, or device, or a program for use in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples of machine-readable storage media include one or more line-based electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0117] To provide user interaction, the systems and technologies described herein can be implemented on a computer, which may have a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball), and the user may provide input to the computer using the keyboard and pointing device. Other types of devices may also provide user interaction, for example, the feedback provided to the user may be any form of sensing feedback (e.g., vision feedback, auditory feedback, or haptic feedback), and may receive input from the user in any form (including acoustic input, voice input, or haptic input).

[0118] The systems and technologies described herein can be run on computing systems including backend components (e.g., data servers), computing systems including middleware components (e.g., application servers), computing systems including frontend components (e.g., user computers having a graphical user interface or web browser, through which users can interact with embodiments of the systems and technologies described herein), or any combination of such backend components, middleware components, and frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0119] A computer system can include clients and servers. Clients and servers are generally geographically separated and typically interact via a communication network. The client-server relationship is generated by computer programs running on corresponding computers that have a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server incorporating blockchain technology.

[0120] It should be understood that the steps can be rearranged, added, or deleted using the various forms of flows shown above. For example, each step described in this disclosure may be performed in parallel, sequentially, or in a different order, as long as the technical invention disclosed herein can achieve the desired results.

[0121] The specific embodiments described above do not limit the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, subcombinations, and substitutions can be made depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure must be within the scope of protection of this disclosure.

Claims

1. A method for determining the audio dataset, A step of obtaining an audio fragment set and the text corresponding to the audio fragment in the audio fragment set, The steps include determining a first emotion type of the audio fragment, the emotion intensity of the first emotion type, a second emotion type of the text, and the emotion intensity of the second emotion type. A step of determining whether or not to retain the audio fragment based on the first emotion type, the emotional intensity of the first emotion type, the second emotion type, and the emotional intensity of the second emotion type, The steps include determining an audio dataset based on the retained audio fragments in the audio fragment set and the text corresponding to the retained audio fragments, A method for determining an audio dataset, including [specific data points].

2. The steps of determining the first emotion type of the audio fragment, the emotion intensity of the first emotion type, the second emotion type of the text, and the emotion intensity of the second emotion type are: The steps include inputting the aforementioned audio fragment into an acoustic feature model and obtaining at least one first candidate emotion type, a first probability of the first candidate emotion type, and an emotion intensity output from the acoustic feature model, A step of selecting the first emotion type from the at least one first candidate emotion type based on the first probability, The steps include inputting the aforementioned audio fragment into a semantic understanding model and obtaining at least one second candidate emotion type, a second probability of the second candidate emotion type, and an emotion intensity output from the semantic understanding model, A step of selecting the second emotion type from the at least one second candidate emotion type based on the second probability, A method for determining an audio dataset according to claim 1, including the following:

3. The step of determining whether or not to retain the audio fragment based on the first emotion type, the emotional intensity of the first emotion type, the second emotion type, and the emotional intensity of the second emotion type is: The steps include determining whether the first emotion type and the second emotion type match, and determining whether the emotional intensity of the first emotion type and the emotional intensity of the second emotion type are equal to or greater than an emotional intensity threshold, The steps include determining to retain the audio fragment if the first emotion type and the second emotion type match, and the emotional intensity of both the first emotion type and the emotional intensity of the second emotion type are equal to or greater than the emotional intensity threshold, A method for determining an audio dataset according to claim 1, including the following:

4. The step of determining whether or not to retain the audio fragment based on the first emotion type, the emotional intensity of the first emotion type, the second emotion type, and the emotional intensity of the second emotion type is: A method for determining an audio dataset according to claim 3, further comprising the step of determining not to retain the audio fragment if the first emotion type and the second emotion type do not match, or if the emotion intensity of the first emotion type is less than the emotion intensity threshold, or if the emotion intensity of the second emotion type is less than the emotion intensity threshold.

5. The number of the first emotion type is multiple, and the number of the second emotion type is multiple, The step of determining whether or not to retain the audio fragment based on the first emotion type, the emotional intensity of the first emotion type, the second emotion type, and the emotional intensity of the second emotion type is: If multiple first emotion types match, the total emotion intensity of the first emotion type is determined based on the emotion intensity of the first emotion type and the weights of the acoustic feature model for determining the first emotion type. If multiple second emotion types match, the total emotion intensity of the second emotion type is determined based on the emotion intensity of the second emotion type and the weights of the semantic understanding model for determining the second emotion type. A step of determining whether or not to retain the audio fragment based on the first emotion type, the total emotion intensity of the first emotion type, the second emotion type, and the total emotion intensity of the second emotion type, A method for determining an audio dataset according to claim 1, including the following:

6. The step of determining whether or not to retain the audio fragment based on the first emotion type, the emotional intensity of the first emotion type, the second emotion type, and the emotional intensity of the second emotion type is: A method for determining an audio dataset according to claim 5, further comprising the step of determining not to retain the audio fragment if a plurality of the first emotion types do not match, or if a plurality of the second emotion types do not match.

7. The step of obtaining the audio fragment set and the text corresponding to the audio fragments in the audio fragment set is: Steps include acquiring audio from multiple sources, The steps include: for each audio, performing interference removal processing and speaker separation processing on the audio to obtain speaker audio from at least one single speaker; The steps include: performing speech activity detection and segmentation processing on the speaker voice of the at least one single speaker to obtain speaker fragments of the at least one single speaker; The steps include determining the set of speech fragments based on the speaker fragments of at least one single speaker, The steps include: performing speech recognition processing on the speech fragments in the speech fragment set to obtain the text corresponding to the speech fragments; A method for determining an audio dataset according to claim 1, including the following:

8. Before performing speech activity detection and segmentation processing on the speaker voice of the at least one single speaker to obtain the speaker fragment of the at least one single speaker, the method: The steps include: performing duplicate speaker detection processing on a single speaker's voice to determine whether or not duplicate fragments exist in the speaker's voice; If duplicate fragments exist in the speaker's voice, the speaker's voice is subjected to a process to remove the duplicate fragments. A method for determining an audio dataset according to claim 7, further comprising:

9. Before determining the set of speech fragments based on the speaker fragment of at least one single speaker, the method The steps include: performing sound quality scoring processing on a speaker fragment of a single speaker to obtain a sound quality score value for the speaker fragment; The steps include: performing a filtering control process on the speaker fragment based on the sound quality score value and the duration of the speaker fragment; A method for determining an audio dataset according to claim 7, further comprising:

10. The step of performing filtering control processing on the speaker fragment based on the sound quality score value and the duration of the speaker fragment is: A method for determining an audio dataset according to claim 9, further comprising the step of performing a filtering process on the speaker fragment if the sound quality score value is less than or equal to a score value threshold, or if the time length is not within the time length range.

11. The step of performing speech recognition processing on the speech fragments in the aforementioned speech fragment set to obtain the text corresponding to the speech fragments is: The steps include determining the language type to which the aforementioned audio fragment belongs, The steps include: selecting a target speech recognition model from among multiple candidate speech recognition models based on the aforementioned language type to obtain a target speech recognition model; The steps include inputting the audio fragment into the target speech recognition model and obtaining the text output from the target speech recognition model, A method for determining an audio dataset according to claim 7, including the following:

12. The audio dataset includes audio fragments in multiple languages, and the method is as follows: A method for determining an audio dataset according to claim 2, further comprising the step of performing fine-tuning training on the acoustic feature model and the semantic understanding model based on audio fragments of multiple languages ​​in the audio dataset and the text corresponding to the audio fragments.

13. A method for training speech synthesis models, A step of acquiring an audio dataset, wherein the audio dataset includes audio fragments, emotion types of the audio fragments, and text corresponding to the audio fragments, and the audio dataset is determined based on the method for determining an audio dataset according to any one of claims 1 to 12. Steps to obtain an initial speech synthesis model, The steps include: training the initial speech synthesis model based on the speech fragment, the emotion type of the speech fragment, and the text to obtain a trained speech synthesis model; Training methods for speech synthesis models, including [specific details omitted].

14. A device for determining audio datasets, A set of audio fragments and an acquisition module for obtaining the text corresponding to the audio fragments in the set of audio fragments, A first determination module for determining the first emotion type of the audio fragment, the emotion intensity of the first emotion type, the second emotion type of the text, and the emotion intensity of the second emotion type, A second decision module for determining whether or not to retain the audio fragment based on the first emotion type, the emotional intensity of the first emotion type, the second emotion type, and the emotional intensity of the second emotion type, A third decision module for determining an audio dataset based on the retained audio fragments in the audio fragment set and the text corresponding to the retained audio fragments, A decision-making device for audio datasets, including [specific component].

15. A training device for speech synthesis models, A first acquisition module for acquiring an audio dataset, wherein the audio dataset includes an audio fragment, an emotion type of the audio fragment, and text corresponding to the audio fragment, and the audio dataset is determined based on the method for determining an audio dataset described in any one of claims 1 to 12. A second acquisition module for obtaining the initial speech synthesis model, A training processing module for performing training on the initial speech synthesis model based on the aforementioned speech fragment, the emotion type of the speech fragment, and the text, in order to obtain a trained speech synthesis model. A training device for speech synthesis models, including a speech synthesis model.

16. It is an electronic device, At least one processor, A memory that is communicably connected to at least one processor, Includes, An electronic device wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 12 or the method according to claim 13.

17. A non-temporary, computer-readable storage medium in which computer instructions are stored, The computer instruction is a non-temporary computer-readable storage medium that causes a computer to perform the method according to any one of claims 1 to 12 or the method according to claim 13.

18. A computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 12 or the method according to claim 13.