A voice generation method and apparatus
By acquiring the text to be processed and the emotion type, and adjusting the speech feature parameters using emotion units and text-to-speech units, speech with multiple emotions is generated. This solves the problem that traditional speech generation technology cannot generate speech with multiple emotions, and achieves the effect of quickly adjusting the generated emotion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI XIYU JIZHI TECH CO LTD
- Filing Date
- 2025-04-22
- Publication Date
- 2026-04-28
AI Technical Summary
Traditional speech generation techniques cannot generate speech with a mixture of emotions, and when it is necessary to generate a non-predetermined emotion, they cannot quickly adjust the generated emotion of the speech generation model without changing the model structure.
By acquiring the text to be processed and the emotion type, using emotion units and text-to-speech units, and adjusting the speech feature parameters according to the adjustment parameters, speech with a variety of mixed emotions is generated, including exact matching and fuzzy matching principles. The target speech is generated by using a weighted or average fusion strategy.
It enables the rapid adjustment of generated emotions and the generation of speech with mixed emotions without changing the structure of the speech generation model, thereby improving the user experience.
Smart Images

Figure CN120808748B_ABST
Abstract
Description
Technical Field
[0001] This case is a divisional application of invention patent application number 202510503104.2, filed on April 22, 2025, entitled "A Speech Generation Method and Apparatus". The present invention relates to the field of speech generation, and more specifically, to a speech generation method and apparatus. Background Technology
[0002] Speech generation is a technology that converts text into speech. This technology is widely used in various scenarios, such as AI (Artificial Intelligence) companion chat scenarios. In these applications, the goal is to generate richer and more immersive voices for users, thereby enhancing the user experience.
[0003] However, traditional speech generation technology can only provide users with speech of a few predetermined emotions, and cannot provide speech with a mixture of multiple emotions. It also cannot maintain the structure of the speech generation model when it is necessary to generate speech with non-predetermined emotions, and cannot quickly adjust the generated emotions of the speech generation model. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a speech generation method and apparatus that can generate speech with a mixture of emotions, and can quickly adjust the generated emotions of the speech generation model without changing the speech generation model structure when it is necessary to generate speech with a non-predetermined emotion.
[0005] In a first aspect, embodiments of this application provide a speech generation method, the speech generation method comprising:
[0006] Obtain the text to be processed, the initial speech corresponding to the text to be processed, and at least one emotion type;
[0007] At least one adjustment parameter is obtained based on at least one emotion type; the adjustment parameter is used to adjust the speech feature parameters.
[0008] The text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters, are input into the text-to-speech unit. The speech feature parameters of the intermediate speech after the text to be processed or the speech feature parameters of the initial speech are transformed to obtain the target speech.
[0009] In one possible implementation, at least one emotion type corresponding to the text to be processed is obtained through the following steps:
[0010] Based on the matching principle, at least one emotion type is matched according to the text content to be processed;
[0011] The matching principle refers to matching one or more emotion types based on emotion words, emotion phrases, semantic information, and contextual information in the text to be processed.
[0012] In one possible implementation, the matching principle includes an exact matching principle and a fuzzy matching principle. The matching of at least one emotion type based on the text content to be processed, according to the matching principle, includes:
[0013] Based on the exact matching principle, the text to be processed is analyzed to obtain the most matching mandatory emotion type;
[0014] And / or, based on the fuzzy matching principle, analyze the text to be processed to obtain at least two candidate emotion types with the highest matching degree;
[0015] The required emotion type and / or at least one alternative emotion type are determined as the final emotion type.
[0016] In one possible implementation, obtaining at least one adjustment parameter corresponding to at least one emotion type includes:
[0017] Based on at least one emotion type, the initial voice is input into at least one emotion unit corresponding to the emotion type to obtain at least one initial adjustment parameter corresponding to at least one adjustment parameter type;
[0018] Based on the fusion strategy, initial adjustment parameters of the same type are fused to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type.
[0019] In one possible implementation, the step of fusing initial adjustment parameters of the same type based on a fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type includes:
[0020] Based on the degree of matching between the text to be processed and each emotion type, at least one initial fusion weight is determined for each emotion type.
[0021] Based on at least one fusion weight corresponding to each emotion type, the target fusion weight corresponding to each emotion type is obtained;
[0022] For any adjustment parameter type, based on the target fusion weight corresponding to each emotion type, at least one initial adjustment parameter of the adjustment parameter type corresponding to each emotion type is weighted and fused to obtain the target adjustment parameter corresponding to the adjustment parameter type.
[0023] In one possible implementation, the step of fusing initial adjustment parameters of the same type based on a fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type further includes:
[0024] For any adjustment parameter type, the target adjustment parameter corresponding to that adjustment parameter type is obtained by averaging and fusing at least one initial adjustment parameter corresponding to that adjustment parameter type.
[0025] In one possible implementation, the speech generation model includes at least two distinct emotion units and the text-to-speech unit, and the target emotion unit corresponding to the target emotion type in the speech generation model is trained through the following steps:
[0026] Obtain at least one corresponding speech training sample, wherein the speech training sample includes at least a first speech sample and a second speech sample, wherein the first speech sample is neutral speech obtained by a text-to-speech model, and the second speech sample is speech corresponding to the target emotion type, and the text content of the second speech sample is the same as the text content of the first speech sample.
[0027] The target emotion unit is trained using at least one speech training sample, and the verification parameters of the target emotion unit are evaluated based on at least one test speech sample. If the evaluation result of the target emotion unit meets the preset verification conditions, the training of the target emotion unit is completed.
[0028] In one possible implementation, the training processes for each emotion unit are independent of each other, and the training processes for the emotion units are independent of the training processes for the text-to-speech units.
[0029] In one possible implementation, the voice training sample further includes a third voice sample, the text content of which is the same as the text content of the first voice sample and the text content of the second voice sample, and the similarity between the emotion type of the third voice sample and the emotion type of the second voice sample is greater than a similarity threshold.
[0030] Secondly, embodiments of this application also provide a speech generation device, the speech generation device comprising:
[0031] The acquisition module is used to acquire the text to be processed, the initial speech corresponding to the text to be processed, and at least one emotion type;
[0032] The input module is used to obtain at least one adjustment parameter corresponding to at least one emotion type; the adjustment parameter is used to adjust the speech feature parameters.
[0033] The input module is further configured to input the text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters into the text-to-speech unit, and transform the speech feature parameters of the intermediate speech after the text to be processed or the speech feature parameters of the initial speech to obtain the target speech.
[0034] This application provides a speech generation method and apparatus. The method includes: acquiring a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; acquiring at least one adjustment parameter corresponding to the at least one emotion type; inputting the text to be processed and the adjustment parameter, or the initial speech and the adjustment parameter, into a text-to-speech unit; transforming the speech feature parameters of the intermediate speech after the text to be processed is converted, or the speech feature parameters of the initial speech, to obtain the target speech. This application can generate speech with a mixture of emotions, and does not change the speech generation model structure when it is necessary to generate speech with a non-predetermined emotion, quickly adjusting the generated emotion of the speech generation model. Attached Figure Description
[0035] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 A flowchart of a speech generation method provided in an embodiment of this application is shown;
[0037] Figure 2 A flowchart illustrating the training process of an emotion unit provided in an embodiment of this application is shown.
[0038] Figure 3 This paper shows a schematic diagram of the structure of a speech generation device provided in an embodiment of this application;
[0039] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0040] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.
[0041] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0042] To enable those skilled in the art to utilize the content of this application, and in conjunction with the specific application scenario of "speech generation," the following implementation methods are provided. For those skilled in the art, the general principles defined herein can be applied to other embodiments and application scenarios without departing from the spirit and scope of this application. Although this application is primarily described within the "speech generation field," it should be understood that this is merely an exemplary embodiment.
[0043] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.
[0044] The following is a detailed description of a speech generation method provided in the embodiments of this application.
[0045] Reference Figure 1 The diagram shown is a flowchart illustrating a speech generation method provided in an embodiment of this application. The exemplary steps of this embodiment are described below:
[0046] S101. Obtain the text to be processed, the initial speech corresponding to the text to be processed, and at least one emotion type.
[0047] In this application's implementation, the text to be processed refers to the text for which speech needs to be generated; this text can be text input by the user or text used to reply to the user in a dialogue scenario. The initial speech is neutral speech obtained by inputting the text to be processed into any trained text-to-speech model. Neutral speech refers to speech without any emotion, i.e., emotionless speech. Therefore, the initial speech does not contain any emotion and is emotionless speech. Emotion types include, but are not limited to, happiness, anger, indifference, anxiety, surprise, etc., and can be any emotion that changes the speech feature parameters. Speech feature parameters can include tone, intonation, and speech rate. Tone refers to the attitude and emotional tendency expressed by the speaker through speech. Intonation refers to the rise and fall of the pitch of the voice during speech. Speech rate refers to the speed at which speech is delivered.
[0048] The text to be processed can correspond to one or more emotion types, which can be obtained using one or more methods among matching principles, user specification, and emotion analysis models. User specification refers to one or more emotion types selected by the user from several candidate emotion type options.
[0049] The matching principle is used to obtain at least one emotion type corresponding to the text to be processed, including: matching at least one emotion type based on the content of the text to be processed according to the matching principle; wherein, the matching principle refers to matching one or more emotion types based on emotion words, emotion phrases, semantic information, and contextual information in the text to be processed.
[0050] In this application embodiment, the matching principle includes the exact matching principle and the fuzzy matching principle. Specifically, the text to be processed is analyzed based on the exact matching principle to obtain the most matching mandatory emotion type; and / or the text to be processed is analyzed based on the fuzzy matching principle to obtain at least two candidate emotion types with the highest matching degree; and the mandatory emotion type and / or at least one candidate emotion type are determined as the final emotion type.
[0051] Specifically, based on the principle of exact matching, the text to be processed is analyzed for emotional words and / or emotional phrases to obtain the most matching mandatory emotional type. Based on the principle of fuzzy matching, the semantic information and / or contextual information in the text to be processed are analyzed to obtain at least two candidate emotional types with the highest degree of matching; the candidate emotional types are sent to the user so that the user can select at least one candidate emotional type; the mandatory emotional type and the at least one candidate emotional type selected by the user are determined as the final emotional type.
[0052] Among these methods, artificial intelligence models are used to match at least one emotion type based on the content of the text to be processed. For example, GPT4 (Generative Pre-trained Transformer 4) or BERT (Bidirectional Encoder Representations from Transformers) both support precise matching of an emotion type and belong to the emotion analysis module in natural language processing. BERT also supports fuzzy matching of at least one emotion based on the text content, meaning that a text to be processed can correspond to multiple emotion types, or several emotion types with high matching degrees can be calculated simultaneously as candidates.
[0053] The process employs a sentiment analysis model to obtain at least one sentiment type corresponding to the text to be processed, including: in a dialogue scenario, determining the content and sentiment type of the user's voice message or the content and sentiment type of the user's text message; the text to be processed refers to the text used to reply to the user; combining the content of the user's voice message or the content of the text message with the text to be processed to form the current dialogue content; identifying historical dialogue content with a similarity greater than a preset similarity; and determining the sentiment type corresponding to the text to be processed based on a preset number of historical sentiment types that replied to the user in the historical dialogue content.
[0054] S102. Obtain at least one adjustment parameter corresponding to at least one emotion type; the adjustment parameter is used to adjust the speech feature parameters.
[0055] In this embodiment of the application, based on at least one emotion type, the initial speech is input into at least one emotion unit corresponding to the emotion type to obtain at least one initial adjustment parameter corresponding to at least one adjustment parameter type; the initial adjustment parameters with the same adjustment parameter type are fused based on the fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type.
[0056] Each emotion type corresponds to an emotion unit; the emotion unit corresponding to an emotion type is used to determine the adjustment parameter corresponding to at least one adjustment parameter type under that emotion type. Each emotion type corresponds to at least one adjustment parameter type; the adjustment parameter types corresponding to different emotion types can be the same or different. Adjustment parameter type refers to the type of adjustment parameter. Speech feature parameters refer to speech parameters that can express emotions.
[0057] For example, if the emotion types include happiness and surprise, the initial voice input is entered into the emotion unit corresponding to happiness to obtain the adjustment parameter of at least one adjustment parameter type under the emotion type of happiness; the initial voice input is entered into the emotion unit corresponding to surprise to obtain the adjustment parameter of at least one adjustment parameter type under the emotion type of surprise.
[0058] In the embodiments of this application, two fusion strategies are provided: a weighted fusion method and an average fusion method. The fusion strategy is used to fuse initial adjustment parameters of the same type to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type. This includes fusing initial adjustment parameters of the same type using either a weighted fusion method or an average fusion method to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type.
[0059] Furthermore, based on a weighted fusion method, initial adjustment parameters of the same type are fused to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type, including:
[0060] Step 1: Based on the degree of matching between the text to be processed and each emotion type, determine at least one initial fusion weight corresponding to each emotion type.
[0061] In this application's embodiments, the higher the degree of matching with the text to be processed, the greater the fusion weight corresponding to the emotion type. This application provides three methods for determining the initial fusion weight corresponding to the emotion type. The greater the fusion weight corresponding to the emotion type, the greater the adjustment effect of the initial adjustment parameters corresponding to that emotion type on the speech.
[0062] The first method for determining the initial fusion weights corresponding to emotion types is as follows: based on the degree of matching between the text to be processed and each emotion type, the initial fusion weights corresponding to each emotion type are determined according to the preset weight setting rules. The preset weight setting rules explain how to determine the fusion weights corresponding to each emotion type based on the degree of matching between the text to be processed and each emotion type.
[0063] For example, the preset weighting rule can be that the fusion weight corresponding to the emotion type with the highest matching degree is 0.5, and the fusion weight corresponding to the second to fifth highest matching degree emotion types is 0.1. Assuming that the emotion type with the highest matching degree is happiness, then the initial fusion weight corresponding to the emotion type of happiness is 0.5, and the initial fusion weight corresponding to the other emotion types with the second to fifth highest matching degree is 0.1.
[0064] The second method for determining the initial fusion weights corresponding to emotion types is to add the matching degree of the text to be processed with all emotion types to obtain the sum; and to obtain the initial fusion weights corresponding to each emotion type by taking the ratio of the matching degree of the text to be processed with each emotion type to the sum.
[0065] For example, the emotion types include happy, excited, and surprised. The matching degree of the text to be processed with the happy emotion type is 0.8, the matching degree of the text to be processed with the excited emotion type is 0.7, and the matching degree of the text to be processed with the surprised emotion type is 0.6. The matching degree of the text to be processed with all emotion types is added together to obtain the sum, i.e., sum = 0.8 + 0.7 + 0.6 = 2.1. Based on the sum, the weights are normalized and calculated: the fusion weight corresponding to the happy emotion type is 0.8 / 2.1 = 8 / 21; the fusion weight corresponding to the excited emotion type is 0.7 / 2.1 = 1 / 3; and the fusion weight corresponding to the surprised emotion type is 0.6 / 2.1 = 2 / 7.
[0066] The third way to determine the fusion weight corresponding to the emotion type is to determine the preset fusion weight corresponding to each emotion type set by the user as the initial fusion weight corresponding to each emotion type.
[0067] Step 2: Based on at least one fusion weight corresponding to each emotion type, obtain the target fusion weight corresponding to each emotion type.
[0068] In this embodiment of the application, the weighted sum or average of all fusion weights corresponding to each emotion type is calculated to determine the target fusion weight corresponding to each emotion type.
[0069] Step 3: For any adjustment parameter type, based on the target fusion weight corresponding to each emotion type, perform weighted fusion on at least one initial adjustment parameter of that adjustment parameter type corresponding to each emotion type to obtain the target adjustment parameter value corresponding to that adjustment parameter type.
[0070] Furthermore, based on the average fusion method, initial adjustment parameters of the same adjustment parameter type are fused to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type, including: for any adjustment parameter type, the at least one initial adjustment parameter corresponding to that adjustment parameter type is averaged and fused to obtain the target adjustment parameter corresponding to that adjustment parameter type.
[0071] In this application implementation, if the user does not use the aforementioned three methods for determining the weights corresponding to emotion types, then the multiple emotion parameters will be averaged and fused together.
[0072] S103. Input the text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters, into the text-to-speech unit, and transform the speech feature parameters of the intermediate speech after the text to be processed or the speech feature parameters of the initial speech to obtain the target speech.
[0073] In this embodiment, the text to be processed and adjustment parameters are input into the text-to-speech unit, and the speech feature parameters of the intermediate speech after the text to be processed are transformed to obtain the target speech. Alternatively, the initial speech and adjustment parameters are input into the text-to-speech unit, and the speech feature parameters of the initial speech are transformed to obtain the target speech.
[0074] Here, the text-to-speech unit includes at least one intermediate layer. Different intermediate layers transform the speech feature parameters of the intermediate speech or the initial speech based on adjustment parameters of different adjustment parameter types. An intermediate layer transforms the speech feature parameters of the intermediate speech or the initial speech based on adjustment parameters of one adjustment parameter type. Specifically, transforming the speech feature parameters of the intermediate speech or the initial speech includes: the i-th intermediate layer transforming the speech feature parameters of the intermediate speech or the initial speech based on its corresponding adjustment parameters; i = i + 1; jump to "the i-th intermediate layer transforming the speech feature parameters of the intermediate speech or the initial speech based on its corresponding adjustment parameters" to continue execution; where the initial value of i is 1, and the maximum value of i is the number of intermediate layers.
[0075] Furthermore, the speech generation model in this embodiment includes at least two different emotion units and a text-to-speech unit. The text-to-speech unit and all emotion units are connected respectively.
[0076] In this embodiment, the emotion unit transmits several adjustment parameters to the speech generation model, so that the emotional information is integrated with the final generated speech content to generate speech content that reflects the emotion. The training processes of each emotion unit are independent of each other, and the training processes of the emotion units are independent of the training processes of the text-to-speech units.
[0077] The text-to-speech unit can be any common TTS (text-to-speech) model, such as the Tacotron series, DeepSpeech series, FastSpeech series, TransformerTTS, VALL-E series, or TTS models based on generative streams or adversarial networks that can convert text into speech. There are no restrictions on this.
[0078] Each emotion unit consists of one or more multilayer perceptron (MLP) modules or convolutional neural network (CNN) modules. Preferably, the emotion unit consists of one multilayer perceptron module or convolutional neural network module.
[0079] One or more selected emotion units output different adaptive adjustment parameters to the text-to-speech unit to embed emotional information into the speech, rather than inputting the same parameters to each layer in the text-to-speech unit.
[0080] The text-to-speech model and each emotion unit are connected through a structure based on AdaLN (Adaptive Layer Normalization) or LoRA (Low-Rank Adaptation).
[0081] The emotion unit passes several adjustment parameters to the text-to-speech unit, integrating emotional information with the final generated speech content. Specifically, each emotion unit outputs multiple adaptive adjustment parameters corresponding to its respective emotion type. These fine-tuning parameters are then injected into different intermediate layers within the text-to-speech unit, performing a linear transformation on the results generated at each layer to modify parameters such as intonation, rate, and tone of the generated speech. It's important to note that this step does not alter the text content; it only changes the speech parameters.
[0082] Further, the text-to-speech unit is trained through the following steps: obtaining at least one text training sample and a corresponding speech conversion label for each text training sample, wherein the text training samples refer to the text that needs to be converted into speech; the speech conversion label refers to the speech obtained by converting the text training samples through the trained text-to-speech model; training the text-to-speech unit based on the text training samples and the corresponding speech conversion labels for each text training sample; and evaluating the verification parameters of the target emotion unit based at least on test text samples. If the evaluation result of the target emotion unit meets the preset verification conditions, the training of the target emotion unit is completed.
[0083] Specifically, the validation parameters of the target emotion unit are determined based on the test text samples, or the set of test text samples and validation text samples. If the evaluation result of the target emotion unit meets the preset validation conditions, the training of the target emotion unit is completed.
[0084] Here, when training the text-to-speech unit, the text-to-speech unit is isolated from all emotion units. Only the text-to-speech unit is trained, and no parameters of the emotion units are changed due to the training of the text-to-speech unit.
[0085] Furthermore, referring to Figure 2 The diagram shown is a flowchart of the training process for the emotion unit provided in an embodiment of this application.
[0086] S201. Obtain at least one corresponding speech training sample.
[0087] The speech training samples include at least a first speech sample and a second speech sample. The first speech sample is neutral speech obtained through a text-to-speech model, and the second speech sample is the speech corresponding to the target emotion type, with the text content of the second speech sample being the same as that of the first speech sample. Neutral speech refers to speech without any emotion. The target emotion type refers to the emotion type corresponding to the target emotion unit that needs to be trained.
[0088] In addition, the voice training samples may also include a third voice sample, the text content of which is the same as the text content of the first voice sample and the text content of the second voice sample, and the similarity between the emotion type of the third voice sample and the emotion type of the second voice sample is greater than the similarity threshold.
[0089] In the embodiments of this application, the similarity between two emotion types refers to the similarity between the speech features of two voices with the same text content of the two emotion types.
[0090] Here, because human emotions are continuous and mixed, for example, the speech feature parameters of the emotional types of "happy" and "excited" partially overlap. Therefore, it is necessary to introduce some similar speech training samples as negative training samples to strengthen the understanding of "happy" by the emotion unit corresponding to "happy".
[0091] S202. The target emotion unit is trained using at least one speech training sample, and the verification parameters of the target emotion unit are evaluated based on at least one test speech sample. If the evaluation result of the target emotion unit meets the preset verification conditions, the training of the target emotion unit is completed.
[0092] In this embodiment, a first speech sample is input into a target emotion unit to obtain at least one adjustment parameter corresponding to the target emotion type; the at least one adjustment parameter corresponding to the target emotion type is input into a trained text-to-speech unit to transform the first speech sample to obtain generated speech; the target emotion unit is updated based on the generated speech and a second speech sample so that the generated speech can move closer to the second speech sample; or the target emotion unit is updated based on the generated speech, the second speech sample, and a third speech sample so that the generated speech can move closer to the second speech sample and move away from the third speech sample; and the verification parameters of the target emotion unit are evaluated at least based on a test speech sample. If the evaluation result of the target emotion unit meets the preset verification conditions, the training of the target emotion unit is completed.
[0093] Specifically, the verification parameters of the target emotion unit are determined based on the test speech samples, or the set of test speech samples and verification speech samples. If the evaluation result of the target emotion unit meets the preset verification conditions, the training of the target emotion unit is completed.
[0094] The verification parameters include at least one of consistency verification, confusion rate verification, and error rate verification.
[0095] For example, to train an emotion unit corresponding to the happy emotion type, the happy emotion unit is trained based on the first speech sample, the second speech sample of the happy emotion type, and the third speech sample, so that it can output adjustment parameters that make the emotion of the speech "speaking happily".
[0096] Here, several pre-trained emotion units are acquired to ensure that each emotion unit can accurately convert speech without emotion into speech with the corresponding emotion. Isolating the text-to-speech units from the emotion units, and from each emotion unit itself, during training simplifies the training process and reduces its complexity.
[0097] In addition, if the already trained emotion units do not cover enough emotion types and it is necessary to add more emotion units, the text-to-speech unit and other emotion units can be isolated, and the speech training samples corresponding to the added emotion types can be used to train the newly added emotion units in accordance with steps S201~S202.
[0098] This application provides a speech generation method and apparatus. The method includes: acquiring a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; acquiring at least one adjustment parameter corresponding to the at least one emotion type; inputting the text to be processed and the adjustment parameter, or the initial speech and the adjustment parameter, into a text-to-speech unit; transforming the speech feature parameters of the intermediate speech after the text to be processed is converted, or the speech feature parameters of the initial speech, to obtain the target speech. This application can generate speech with a mixture of emotions, and does not change the speech generation model structure when it is necessary to generate speech with a non-predetermined emotion, quickly adjusting the generated emotion of the speech generation model.
[0099] Based on the same inventive concept, this application also provides a speech generation device corresponding to the speech generation method. Since the principle of the device in this application is similar to that of the speech generation method described above in this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0100] Reference Figure 3 The diagram shown is a schematic representation of a speech generation device provided in an embodiment of this application. The speech generation device includes:
[0101] The acquisition module 301 is used to acquire the text to be processed, the initial speech corresponding to the text to be processed, and at least one emotion type;
[0102] Input module 302 is used to obtain at least one adjustment parameter corresponding to at least one emotion type; the adjustment parameter is used to adjust speech feature parameters;
[0103] The input module 302 is further configured to input the text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters, into the text-to-speech unit to transform the speech feature parameters of the intermediate speech after the text to be processed or the speech feature parameters of the initial speech to obtain the target speech. In one possible implementation, the acquisition module 301 is specifically configured to precisely match and / or fuzzily match at least one emotion type of the text to be processed; wherein, the precise matching is based on matching emotion types based on emotion words or emotion phrases in the text to be processed; and the fuzzy matching is based on matching emotion types based on semantic information and / or contextual information of the text to be processed.
[0104] This application provides a speech generation device that can generate speech with a mixture of emotions. When it is necessary to generate speech with a non-predetermined emotion, the speech generation model structure is not changed, and the generated emotion of the speech generation model is quickly adjusted.
[0105] like Figure 4 As shown in the embodiment of this application, an electronic device 400 includes a processor 401, a memory 402, and a bus. The memory 402 stores machine-readable instructions that can be executed by the processor 401. When the electronic device is running, the processor 401 communicates with the memory 402 via the bus, and the processor 401 executes the machine-readable instructions to perform the steps of the speech generation method described above.
[0106] Specifically, the memory 402 and processor 401 mentioned above can be general-purpose memory and processor, without any specific limitations. When the processor 401 runs the computer program stored in the memory 402, it can execute the above-mentioned speech generation method.
[0107] Corresponding to the above-described speech generation method, this application embodiment also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the above-described speech generation method.
[0108] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces; the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.
[0109] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0110] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0111] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the information processing methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0112] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A speech generation method, characterized in that, The method includes: Obtain the text to be processed, the initial speech corresponding to the text to be processed, and at least one emotion type; At least one adjustment parameter is obtained based on at least one emotion type; the adjustment parameter is used to adjust the speech feature parameters; wherein, each emotion type corresponds to an emotion unit, and the emotion unit corresponding to the emotion type is used to determine the adjustment parameter corresponding to at least one adjustment parameter type under the emotion type; The text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters, are input into the text-to-speech unit. The speech feature parameters of the intermediate speech after the text to be processed or the speech feature parameters of the initial speech are transformed to obtain the target speech. The speech generation model includes at least two different emotion units and text-to-speech units. The target emotion unit corresponding to the target emotion type in the speech generation model is trained through the following steps: At least one corresponding speech training sample is obtained, including a first speech sample, a second speech sample, and a third speech sample. The first speech sample is neutral speech obtained through a text-to-speech model. The second speech sample is speech corresponding to the target emotion type, and the text content of the second speech sample is the same as the text content of the first speech sample. The text content of the third speech sample is the same as the text content of both the first and second speech samples, and the similarity between the emotion type of the third speech sample and the emotion type of the second speech sample is greater than a similarity threshold. The target emotion unit is trained using the at least one speech training sample. Then, the validation parameters of the target emotion unit are evaluated based on at least one test speech sample. If the evaluation result of the target emotion unit meets a preset validation condition, the target emotion unit training is complete. The step of obtaining at least one adjustment parameter corresponding to at least one emotion type includes: inputting the initial speech into at least one emotion unit corresponding to the emotion type according to at least one emotion type to obtain at least one initial adjustment parameter corresponding to at least one adjustment parameter type; and fusing the initial adjustment parameters with the same adjustment parameter type based on a fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type. The step of training the target emotion unit using at least one speech training sample, and then evaluating the verification parameters of the target emotion unit based at least on a test speech sample, wherein the target emotion unit training is complete if the evaluation result of the target emotion unit meets a preset verification condition, includes: inputting the first speech sample into the target emotion unit to obtain at least one adjustment parameter corresponding to the target emotion type; inputting the at least one adjustment parameter corresponding to the target emotion type into a trained text-to-speech unit to transform the first speech sample to obtain generated speech; updating the target emotion unit according to the generated speech, the second speech sample, and the third speech sample, so that the generated speech can move closer to the second speech sample and further away from the third speech sample; and evaluating the verification parameters of the target emotion unit based at least on a test speech sample, wherein the target emotion unit training is complete if the evaluation result of the target emotion unit meets a preset verification condition.
2. The speech generation method according to claim 1, characterized in that, At least one emotion type corresponding to the text to be processed is obtained through the following steps: Based on the matching principle, at least one emotion type is matched according to the text content to be processed; The matching principle refers to matching one or more emotion types based on emotion words, emotion phrases, semantic information, and contextual information in the text to be processed.
3. The speech generation method according to claim 2, characterized in that, The matching principles include exact matching and fuzzy matching. Based on the matching principles, matching at least one emotion type according to the text content to be processed includes: Based on the exact matching principle, the text to be processed is analyzed to obtain the most matching mandatory emotion type; And / or, based on the fuzzy matching principle, analyze the text to be processed to obtain at least two candidate emotion types with the highest matching degree; The required emotion type and / or at least one alternative emotion type are determined as the final emotion type.
4. The speech generation method according to claim 1, characterized in that, The method of fusing initial adjustment parameters of the same type based on a fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type includes: Based on the degree of matching between the text to be processed and each emotion type, at least one initial fusion weight is determined for each emotion type. Based on at least one fusion weight corresponding to each emotion type, the target fusion weight corresponding to each emotion type is obtained; For any adjustment parameter type, based on the target fusion weight corresponding to each emotion type, at least one initial adjustment parameter of the adjustment parameter type corresponding to each emotion type is weighted and fused to obtain the target adjustment parameter corresponding to the adjustment parameter type.
5. The speech generation method according to claim 1, characterized in that, The step of fusing initial adjustment parameters of the same type based on a fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type further includes: For any adjustment parameter type, the target adjustment parameter corresponding to that adjustment parameter type is obtained by averaging and fusing at least one initial adjustment parameter corresponding to that adjustment parameter type.
6. The speech generation method according to claim 1, characterized in that, The training processes for each emotion unit are independent of each other, and the training processes for the emotion units are independent of the training processes for the text-to-speech units.
7. A speech generation device, characterized in that, The device includes: The acquisition module is used to acquire the text to be processed, the initial speech corresponding to the text to be processed, and at least one emotion type; An input module is used to obtain at least one adjustment parameter corresponding to at least one emotion type; the adjustment parameter is used to adjust speech feature parameters; wherein, each emotion type corresponds to an emotion unit, and the emotion unit corresponding to the emotion type is used to determine the adjustment parameter corresponding to at least one adjustment parameter type under the emotion type; The input module is further configured to input the text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters into the text-to-speech unit, and transform the speech feature parameters of the intermediate speech after the text to be processed is converted or the speech feature parameters of the initial speech to obtain the target speech; The speech generation model includes at least two different emotion units and a text-to-speech unit. The input module trains the target emotion unit corresponding to the target emotion type in the speech generation model through the following steps: obtaining at least one corresponding speech training sample, the speech training sample including a first speech sample, a second speech sample, and a third speech sample, the first speech sample being neutral speech obtained through text-to-speech model processing, the second speech sample being speech corresponding to the target emotion type, and the text content of the second speech sample being the same as the text content of the first speech sample, the text content of the third speech sample being the same as the text content of the first speech sample and the text content of the second speech sample, and the similarity between the emotion type of the third speech sample and the emotion type of the second speech sample being greater than a similarity threshold; training the target emotion unit through the at least one speech training sample, and then evaluating the verification parameters of the target emotion unit based at least on test speech samples; if the evaluation result of the target emotion unit meets the preset verification conditions, the target emotion unit training is complete. Specifically, the input module is used to input the initial speech into at least one emotion unit corresponding to at least one emotion type according to at least one emotion type, to obtain at least one initial adjustment parameter corresponding to at least one adjustment parameter type; and to fuse the initial adjustment parameters with the same adjustment parameter type based on a fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type. The input module is specifically configured to input the first speech sample into the target emotion unit to obtain at least one adjustment parameter corresponding to the target emotion type; input the at least one adjustment parameter corresponding to the target emotion type into the trained text-to-speech unit to transform the first speech sample to obtain generated speech; update the target emotion unit according to the generated speech, the second speech sample, and the third speech sample, so that the generated speech can move closer to the second speech sample and further away from the third speech sample; and evaluate the verification parameters of the target emotion unit based at least on the test speech sample. If the evaluation result of the target emotion unit meets the preset verification conditions, the training of the target emotion unit is completed.
Citation Information
Patent Citations
Speech synthesis method and device, electronic equipment and storage medium
CN114678003A
Voice emotion adjustment method, device, equipment and product
CN118197331A