Voice generation method and device

By obtaining the pending text, initial speech and emotion type, obtaining adjustment parameters according to the emotion type, inputting the text to voice unit to transform the speech feature parameters to generate the target speech. It solves the problem that traditional speech generation technology cannot generate mixed speech with multiple emotions, realizes the rapid adjustment of the generated emotions of the speech generation model, and improves the flexibility and user experience of speech generation.

CN120015012AActive Publication Date: 2025-05-16SHANGHAI XIYU JIZHI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510503104.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-16
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Traditional speech generation technology cannot generate speech with mixed emotions, and when it is necessary to generate speech with non-predetermined emotions, it is difficult to quickly adjust the generated emotions of the speech generation model.

Method used

By obtaining the pending text, initial speech and emotion type, obtaining adjustment parameters according to the emotion type, inputting the text to voice unit to transform the speech feature parameters to generate the target speech. The method includes the principles of exact match and fuzzy match, used to determine sentiment types and generate target adjustment parameters through fusion strategies.

Benefits of technology

It realizes the generation of voices with mixed emotions, and does not change the voice generation model structure when it is necessary to generate voices with non-predetermined emotions, and quickly adjusts the generated emotions of the voice generation model, improving the flexibility and user experience of voice generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015012A_ABST
    Figure CN120015012A_ABST
Patent Text Reader

Abstract

The invention relates to the field of voice generation, in particular to a voice generation method and device, and the method comprises the steps: obtaining a to-be-processed text, an initial voice corresponding to the to-be-processed text, and at least one emotion type; obtaining at least one corresponding adjustment parameter according to the at least one emotion type; and inputting the to-be-processed text and the adjustment parameter, or the initial voice and the adjustment parameter into a character-to-voice unit, and converting the voice feature parameter of the intermediate voice converted from the to-be-processed text or the voice feature parameter of the initial voice to obtain a target voice. According to the invention, the voice with various mixed emotions can be generated, the structure of the voice generation model is not changed when the voice with non-predetermined emotions needs to be generated, and the generated emotions of the voice generation model are rapidly adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech generation, and in particular to a speech generation method and device. Background Art

[0002] Speech generation is a technology that converts text into speech. This technology is widely used in various scenarios, such as AI (Artificial Intelligence) companion chat scenarios. In various application scenarios, it is hoped that richer and more immersive speech can be generated for users to improve the user experience.

[0003] However, traditional speech generation technology can only provide users with speech with a few predetermined emotions, but cannot provide speech with multiple mixed emotions. It is also unable to generate speech with non-predetermined emotions without changing the structure of the speech generation model, and is unable to quickly adjust the generated emotions of the speech generation model. Summary of the invention

[0004] In view of this, the purpose of the present application is to provide a speech generation method and device, which can generate speech with multiple mixed emotions, and when it is necessary to generate speech with non-predetermined emotions, the speech generation model structure does not need to be changed, and the generated emotions of the speech generation model can be quickly adjusted.

[0005] In a first aspect, an embodiment of the present application provides a speech generation method, the speech generation method comprising: Acquire a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; Acquire at least one corresponding adjustment parameter according to at least one emotion type; the adjustment parameter is used to adjust the speech feature parameter; The text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters are input into a text-to-speech unit, and the speech feature parameters of the intermediate speech converted from the text to be processed or the speech feature parameters of the initial speech are transformed to obtain a target speech.

[0006] In a possible implementation, at least one emotion type corresponding to the text to be processed is obtained through the following steps: Based on the matching principle, matching at least one emotion type according to the content of the text to be processed; The matching principle refers to matching one or more emotion types based on emotion words, emotion phrases, semantic information, and context information in the text to be processed.

[0007] In a possible implementation, the matching principle includes an exact matching principle and a fuzzy matching principle, and the matching principle-based method matches at least one emotion type according to the content of the text to be processed, including: Analyze the text to be processed based on the exact matching principle to obtain a most matching mandatory emotion type; and / or, analyzing the text to be processed based on the fuzzy matching principle to obtain at least two candidate emotion types with the highest matching degree; The mandatory emotion type and / or at least one candidate emotion type are determined as final emotion types.

[0008] In a possible implementation, acquiring at least one corresponding adjustment parameter according to at least one emotion type includes: According to at least one emotion type, inputting the initial speech into at least one emotion unit corresponding to the emotion type, and obtaining at least one initial adjustment parameter corresponding to at least one adjustment parameter type; Based on the fusion strategy, initial adjustment parameters of the same adjustment parameter type are fused to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type.

[0009] In a possible implementation manner, the fusing of initial adjustment parameters of the same adjustment parameter type based on a fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type includes: Determining at least one initial fusion weight corresponding to each emotion type according to the matching degree between the text to be processed and each emotion type; According to at least one fusion weight corresponding to each emotion type, a target fusion weight corresponding to each emotion type is obtained; For any adjustment parameter type, based on the target fusion weight corresponding to each emotion type, at least one initial adjustment parameter of the adjustment parameter type corresponding to each emotion type is weightedly fused to obtain the target adjustment parameter corresponding to the adjustment parameter type.

[0010] In a possible implementation manner, the fusing initial adjustment parameters of the same adjustment parameter type based on the fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type further includes: For any adjustment parameter type, at least one initial adjustment parameter corresponding to the adjustment parameter type is averaged and fused to obtain a target adjustment parameter corresponding to the adjustment parameter type.

[0011] In a possible implementation, the speech generation model includes at least two different emotion units and the text-to-speech unit, and the target emotion unit corresponding to the target emotion type in the speech generation model is trained by the following steps: Acquire at least one corresponding speech training sample, wherein the speech training sample includes at least a first speech sample and a second speech sample, wherein the first speech sample is a neutral speech obtained by processing a text-to-speech model, and the second speech sample is a speech corresponding to the target emotion type, and the text content of the second speech sample is the same as the text content of the first speech sample; The target emotion unit is trained by using the at least one speech training sample, and a verification parameter of the target emotion unit is evaluated based on at least a test speech sample. If the evaluation result of the target emotion unit meets a preset verification condition, the training of the target emotion unit is completed.

[0012] In a possible implementation, the training processes of the emotion units are independent of each other, and the training process of the emotion unit is independent of the training process of the text-to-speech unit.

[0013] In a possible implementation, the speech training sample also includes a third speech sample, the text content of the third speech sample is the same as the text content of the first speech sample and the text content of the second speech sample, and the similarity between the emotion type of the third speech sample and the emotion type of the second speech sample is greater than a similarity threshold.

[0014] In a second aspect, an embodiment of the present application further provides a speech generating device, the speech generating device comprising: An acquisition module, used for acquiring a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; An input module, used for acquiring at least one corresponding adjustment parameter according to at least one emotion type; the adjustment parameter is used for adjusting the speech feature parameter; The input module is also used to input the text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters into a text-to-speech unit, transform the speech feature parameters of the intermediate speech converted from the text to be processed or the speech feature parameters of the initial speech to obtain the target speech.

[0015] The embodiment of the present application provides a speech generation method and device, the method comprising: obtaining a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; obtaining at least one corresponding adjustment parameter according to at least one emotion type; inputting the text to be processed and the adjustment parameter, or the initial speech and the adjustment parameter into a text-to-speech unit, transforming the speech feature parameters of the intermediate speech after the text to be processed is converted or the speech feature parameters of the initial speech, and obtaining the target speech. The present application can generate speech with multiple mixed emotions, and when it is necessary to generate speech with non-predetermined emotions, the speech generation model structure is not changed, and the generated emotion of the speech generation model is quickly adjusted. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 A flow chart of a speech generation method provided by an embodiment of the present application is shown; Figure 2 A training flow chart of an emotion unit provided in an embodiment of the present application is shown; Figure 3 A schematic diagram of the structure of a speech generating device provided in an embodiment of the present application is shown; Figure 4 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0018] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of explanation and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn in real proportion. The flowchart used in this application shows the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowchart can be implemented out of sequence, and the steps without logical context can be reversed in order or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart under the guidance of the content of the present application, or remove one or more operations from the flowchart.

[0019] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.

[0020] In order to enable those skilled in the art to use the content of this application, the following implementation is provided in conjunction with the specific application scenario "speech generation field". For those skilled in the art, the general principles defined herein can be applied to other embodiments and application scenarios without departing from the spirit and scope of this application. Although this application is mainly described around the "speech generation field", it should be understood that this is only an exemplary embodiment.

[0021] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.

[0022] The following is a detailed description of a speech generation method provided in an embodiment of the present application.

[0023] Reference Figure 1 As shown, it is a flowchart of a speech generation method provided in an embodiment of the present application. The exemplary steps of the embodiment of the present application are described below: S101, obtaining a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type.

[0024] In the implementation mode of the present application, the text to be processed refers to the text for which speech needs to be generated; the text to be processed can be the text input by the user, or it can be the text used to reply to the user in the dialogue scenario. The initial speech is the neutral speech obtained by inputting the text to be processed into any trained text-to-speech model. Neutral speech refers to speech without any emotion, that is, emotionless speech. Therefore, the initial speech does not contain any emotion and is emotionless speech. Emotion types include but are not limited to happy, angry, indifferent, anxious, surprised, etc., which can be any emotion that changes the speech feature parameters. Among them, the speech feature parameters may include tone, intonation, and speech speed. Tone refers to the attitude and emotional tendency expressed by the speaker through speech. Intonation refers to the rise and fall of the voice when speaking. Speech speed refers to the speed of language when speaking.

[0025] The text to be processed may correspond to one or more emotion types, which may be obtained by using one or more methods of matching principle, user specification and emotion analysis model. User specification refers to one or more emotion types clicked by the user from a number of emotion type options to be selected.

[0026] Adopting a matching principle to obtain at least one emotion type corresponding to the text to be processed, including: matching at least one emotion type according to the content of the text to be processed based on the matching principle; wherein the matching principle refers to matching one or more emotion types based on emotion vocabulary, emotion phrases, semantic information, and context information in the text to be processed.

[0027] In an embodiment of the present application, the matching principle includes an exact matching principle and a fuzzy matching principle, wherein the text to be processed is analyzed based on the exact matching principle to obtain the most matching required emotion type; and / or the text to be processed is analyzed based on the fuzzy matching principle to obtain at least two alternative emotion types with the highest degree of matching; and the required emotion type and / or at least one alternative emotion type are determined as the final emotion type.

[0028] Specifically, the emotional words and / or emotional phrases in the text to be processed are analyzed based on the exact matching principle to obtain a most matching mandatory emotional type. The semantic information and / or context information in the text to be processed are analyzed based on the fuzzy matching principle to obtain at least two candidate emotional types with the highest matching degree; the candidate emotional types are sent to the user so that the user selects at least one candidate emotional type; and the mandatory emotional type and at least one candidate emotional type selected by the user are determined as the final emotional type.

[0029] Among them, the artificial intelligence model is used based on the matching principle to match at least one emotion type according to the content of the text to be processed. For example, GPT4 (Generative Pre-trained Transformer 4) or BERT (Bidirectional Encoder Representations from Transformers) both support accurate matching of one emotion type and belong to the emotion analysis module in natural language processing. BERT also supports fuzzy matching of at least one emotion according to the text content, that is, one text to be processed corresponds to multiple emotion types, or several emotion types with high matching degrees are calculated at the same time as alternatives.

[0030] A sentiment analysis model is used to obtain at least one sentiment type corresponding to a text to be processed, including: in a conversation scenario, determining the content and sentiment type of a voice sent by a user, or the content and sentiment type of a text sent by a user; the text to be processed refers to text used to reply to the user; the content of the voice or text sent by the user and the text to be processed are combined to form the current conversation content; historical conversation content whose similarity to the current conversation content is greater than a preset similarity; and historical sentiment types that reply to a preset number of users in the historical conversation content are determined as the sentiment type corresponding to the text to be processed.

[0031] S102. Acquire at least one corresponding adjustment parameter according to at least one emotion type; the adjustment parameter is used to adjust the speech feature parameter.

[0032] In an embodiment of the present application, according to at least one emotion type, the initial speech is input into at least one emotion unit corresponding to the emotion type to obtain at least one initial adjustment parameter corresponding to at least one adjustment parameter type; based on the fusion strategy, the initial adjustment parameters of the same adjustment parameter type are fused to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type.

[0033] Among them, each emotion type corresponds to an emotion unit; the emotion unit corresponding to the emotion type is used to determine the adjustment parameter corresponding to at least one adjustment parameter type under the emotion type. Each emotion type corresponds to at least one adjustment parameter type; the adjustment parameter types corresponding to different emotion types can be the same or different. The adjustment parameter type refers to the type of the adjustment parameter. The voice feature parameter refers to the voice parameter that can express the emotion.

[0034] For example, the emotion types include happiness and surprise, then the initial speech is input into the emotion unit corresponding to happiness, and an adjustment parameter of at least one adjustment parameter type under the emotion type of happiness is obtained; the initial speech is input into the emotion unit corresponding to surprise, and an adjustment parameter of at least one adjustment parameter type under the emotion type of surprise is obtained.

[0035] In the implementation of the present application, two fusion strategies are provided, namely, a weighted fusion method and an average fusion method. Based on the fusion strategy, initial adjustment parameters of the same adjustment parameter type are fused to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type, including: based on the weighted fusion method or the average fusion method, initial adjustment parameters of the same adjustment parameter type are fused to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type.

[0036] Further, initial adjustment parameters of the same adjustment parameter type are fused based on a weighted fusion method to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type, including: Step 1: Determine at least one initial fusion weight corresponding to each emotion type according to the matching degree between the text to be processed and each emotion type.

[0037] In the implementation mode of the present application, the higher the degree of matching with the text to be processed, the greater the fusion weight corresponding to the emotion type. The embodiment of the present application provides three ways to determine the initial fusion weight corresponding to the emotion type. The greater the fusion weight corresponding to the emotion type, the greater the adjustment strength of the initial adjustment parameter corresponding to the emotion type on the speech.

[0038] The first method of determining the initial fusion weight corresponding to the emotion type is: according to the preset weight setting rules, based on the matching degree between the text to be processed and each emotion type, the initial fusion weight corresponding to each emotion type is determined; wherein, the preset weight setting rules explain how to determine the fusion weight corresponding to each emotion type based on the matching degree between the text to be processed and each emotion type.

[0039] For example, the preset weight setting rule may be that the fusion weight corresponding to the emotion type with the highest matching degree is 0.5, and the fusion weights corresponding to the emotion types with the second highest matching degree to the fifth highest matching degree are all 0.1. Assuming that the emotion type with the highest matching degree is happy, the initial fusion weight corresponding to the emotion type corresponding to happy is 0.5, and the initial fusion weights corresponding to the other emotion types with the second to fifth matching degrees are 0.1.

[0040] The second method to determine the initial fusion weight corresponding to the emotion type is: add the matching degree between the text to be processed and all emotion types to obtain the sum value; calculate the ratio of the matching degree between the text to be processed and each emotion type to the sum value to obtain the initial fusion weight corresponding to each emotion type.

[0041] For example, the emotion types include happy, excited, and surprised. The matching degree between the text to be processed and the emotion type of happy is 0.8, the matching degree between the text to be processed and the emotion type of excited is 0.7, and the matching degree between the text to be processed and the emotion type of surprised is 0.6. The matching degree between the text to be processed and all emotion types is added to get the sum value, that is, the sum value = 0.8 + 0.7 + 0.6 = 2.1. The weights are normalized based on the sum value, and the fusion weight corresponding to the emotion type of happy is 0.8 / 2.1 = 8 / 21. The fusion weight corresponding to the emotion type of excited is 0.7 / 2.1 = 1 / 3. The fusion weight corresponding to the emotion type of surprised is 0.6 / 2.1 = 2 / 7.

[0042] The third method of determining the fusion weight corresponding to the emotion type is to determine the preset fusion weight corresponding to each emotion type set by the user as the initial fusion weight corresponding to each emotion type.

[0043] Step 2: According to at least one fusion weight corresponding to each emotion type, a target fusion weight corresponding to each emotion type is obtained.

[0044] In the implementation manner of the present application, all fusion weights corresponding to each emotion type are weighted summed or averaged to determine the target fusion weight corresponding to each emotion type.

[0045] Step three: for any adjustment parameter type, based on the target fusion weight corresponding to each emotion type, weighted fusion is performed on at least one initial adjustment parameter of the adjustment parameter type corresponding to each emotion type to obtain the corresponding target adjustment parameter value of the adjustment parameter type.

[0046] Furthermore, initial adjustment parameters of the same adjustment parameter type are fused based on an average fusion method to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type, including: for any adjustment parameter type, at least one initial adjustment parameter corresponding to the adjustment parameter type is averaged and fused to obtain a target adjustment parameter corresponding to the adjustment parameter type.

[0047] In the implementation manner of the present application, if the user does not adopt the above three methods of determining the weights corresponding to the emotion types, the multiple emotion parameters are averaged and merged.

[0048] S103: Input the text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters, into a text-to-speech unit, transform the speech feature parameters of the intermediate speech converted from the text to be processed or the speech feature parameters of the initial speech to obtain the target speech.

[0049] In the implementation mode of the present application, the text to be processed and the adjustment parameters are input into the text-to-speech unit, and the speech feature parameters of the intermediate speech after the text to be processed are transformed to obtain the target speech. Alternatively, the initial speech and the adjustment parameters are input into the text-to-speech unit, and the speech feature parameters of the initial speech are transformed to obtain the target speech.

[0050] Here, the text-to-speech unit includes at least one intermediate layer, and different intermediate layers transform the speech feature parameters of the intermediate speech after the text is converted or the speech feature parameters of the initial speech based on adjustment parameters of different adjustment parameter types; an intermediate layer transforms the speech feature parameters of the intermediate speech after the text is converted or the speech feature parameters of the initial speech based on adjustment parameters of one adjustment parameter type. Specifically, the speech feature parameters of the intermediate speech after the text is converted or the speech feature parameters of the initial speech are transformed, including: the i-th intermediate layer transforms the speech feature parameters of the intermediate speech after the text is converted or the speech feature parameters of the initial speech based on the corresponding adjustment parameters; i=i+1; jump to "the i-th intermediate layer transforms the speech feature parameters of the intermediate speech after the text is converted or the speech feature parameters of the initial speech based on the corresponding adjustment parameters" to continue execution; wherein the initial value of i is 1, and the maximum value of i is the number of intermediate layers.

[0051] Furthermore, the speech generation model in the embodiment of the present application includes at least two different emotion units and a text-to-speech unit. The text-to-speech unit and all emotion units are connected respectively.

[0052] In the implementation mode of the present application, the emotion unit transmits several adjustment parameters to the speech generation model so that the emotion information is integrated with the final generated speech content to generate speech content that can reflect the emotion. The training process between each emotion unit is independent of each other, and the training process of the emotion unit is independent of the training process of the text-to-speech unit.

[0053] The text-to-speech unit can be a common TTS (text-to-speech) model, such as the Tacotron series, DeepSpeech series, FastSpeech series, TransformerTTS, VALL-E series, TTS models based on generative streams or adversarial networks, and other models that can convert text into speech. There is no restriction on this.

[0054] Each emotion unit is composed of one or more multilayer perceptron modules (Multilayer Perceptron, MLP) or convolutional neural network (CNN) modules. Preferably, the emotion unit is composed of one multilayer perceptron module or convolutional neural network.

[0055] The selected one or more emotion units will output different adaptive adjustment parameters to the text-to-speech unit to embed the emotion information into the speech, instead of inputting the same parameters to each layer in the text-to-speech unit.

[0056] The text-to-speech model and each emotion unit are implemented through a structure based on AdaLN (Adaptive Layer Normalization) or LoRA (Low-Rank Adaptation).

[0057] The emotion unit transmits several adjustment parameters to the text-to-speech unit so that the emotion information is integrated with the final generated speech content. Specifically, each emotion unit outputs adaptive adjustment parameters corresponding to multiple adjustment parameter types according to the corresponding emotion type, and each fine-tuning parameter is injected into different intermediate layers in the text-to-speech unit, and the results generated by each layer in the text-to-speech unit are linearly transformed to transform the intonation, speech speed, tone and other parameters of the generated speech. It should be noted that this step will not change the text content, but only the parameters of the speech.

[0058] Furthermore, the text-to-speech unit is trained through the following steps: obtaining at least one text training sample and a converted speech label corresponding to each text training sample, wherein the several text training samples refer to texts that need to be converted into speech; the converted speech label refers to the speech obtained by converting the text training samples through the trained text-to-speech model; the text-to-speech unit is trained based on the text training samples and the converted speech labels corresponding to each text training sample; and the verification parameters of the target emotion unit are evaluated at least based on the test text sample. If the evaluation result of the target emotion unit meets the preset verification conditions, the training of the target emotion unit is completed.

[0059] Specifically, the verification parameter of the target emotion unit is determined based on the test text sample, or a set of the test text sample and the verification text sample. If the evaluation result of the target emotion unit meets the preset verification condition, the training of the target emotion unit is completed.

[0060] Here, when training the text-to-speech unit, the text-to-speech unit is isolated from all emotion units, and only the text-to-speech unit is trained, and any parameters of the emotion unit are not changed due to the training of the text-to-speech unit.

[0061] Further, refer to Figure 2 As shown, it is a training flow chart of the emotion unit provided in the embodiment of the present application.

[0062] S201. Obtain at least one corresponding speech training sample.

[0063] The speech training samples include at least a first speech sample and a second speech sample, wherein the first speech sample is a neutral speech obtained by processing a text-to-speech model, and the second speech sample is a speech corresponding to a target emotion type, and the text content of the second speech sample is the same as the text content of the first speech sample. Neutral speech refers to speech without any emotion. The target emotion type refers to the emotion type corresponding to the target emotion unit to be trained.

[0064] In addition, the speech training sample may also include a third speech sample, the text content of the third speech sample is the same as the text content of the first speech sample and the text content of the second speech sample, and the similarity between the emotion type of the third speech sample and the emotion type of the second speech sample is greater than a similarity threshold.

[0065] In the implementation manner of the present application, the similarity between two emotion types refers to the similarity between the speech features of two voices with the same text content of the two emotion types.

[0066] Here, since human emotions are continuous and mixed, for example, the speech feature parameters of speech with the emotion types of "happy" and "excited" partially overlap, some similar speech training samples need to be introduced here as negative training samples to strengthen the understanding of "happy" by the emotion unit corresponding to "happy".

[0067] S202: Train the target emotion unit through at least one speech training sample, and evaluate the verification parameter of the target emotion unit based on at least a test speech sample. If the evaluation result of the target emotion unit meets a preset verification condition, the training of the target emotion unit is completed.

[0068] In an embodiment of the present application, a first speech sample is input into a target emotion unit to obtain at least one adjustment parameter corresponding to the target emotion type; at least one adjustment parameter corresponding to the target emotion type is input into a trained text-to-speech unit to transform the first speech sample to obtain a generated speech; the target emotion unit is updated according to the generated speech and the second speech sample so that the generated speech can approach the second speech sample; or the target emotion unit is updated according to the generated speech, the second speech sample and the third speech sample so that the generated speech can approach the second speech sample and move away from the third speech sample; and the verification parameters of the target emotion unit are evaluated at least based on the test speech sample. If the evaluation result of the target emotion unit meets the preset verification conditions, the training of the target emotion unit is completed.

[0069] Specifically, the verification parameter of the target emotion unit is determined based on the test speech sample, or the set of the test speech sample and the verification speech sample. If the evaluation result of the target emotion unit meets the preset verification condition, the training of the target emotion unit is completed.

[0070] The verification parameters include at least one of consistency verification, confusion rate verification, and error rate verification.

[0071] For example, to train an emotion unit corresponding to a happy emotion type, the happy emotion unit is trained based on a first speech sample, a second speech sample of a happy emotion type, and a third speech sample, so that it can output adjustment parameters that make the emotion of the speech "speaking happily".

[0072] Here, several trained emotion units are obtained to ensure that each emotion unit can accurately convert speech without emotion into speech with corresponding emotion. The isolation training between the text-to-speech unit and the emotion unit and between each emotion unit can also make the training process simpler and reduce the complexity of training.

[0073] In addition, if the emotion types covered by the trained emotion units are insufficient and additional emotion units need to be added, the text-to-speech units and other emotion units can be isolated, and the speech training samples corresponding to the added emotion types can be used to train the newly added emotion units according to steps S201~S202.

[0074] The embodiment of the present application provides a speech generation method and device, the method comprising: obtaining a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; obtaining at least one corresponding adjustment parameter according to at least one emotion type; inputting the text to be processed and the adjustment parameter, or the initial speech and the adjustment parameter into a text-to-speech unit, transforming the speech feature parameters of the intermediate speech after the text to be processed is converted or the speech feature parameters of the initial speech, and obtaining the target speech. The present application can generate speech with multiple mixed emotions, and when it is necessary to generate speech with non-predetermined emotions, the speech generation model structure is not changed, and the generated emotion of the speech generation model is quickly adjusted.

[0075] Based on the same inventive concept, a speech generation device corresponding to the speech generation method is also provided in the embodiment of the present application. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the above-mentioned speech generation method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0076] Reference Figure 3 FIG. 1 is a schematic diagram of a speech generating device provided in an embodiment of the present application, wherein the speech generating device comprises: An acquisition module 301 is used to acquire a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; An input module 302, configured to obtain at least one corresponding adjustment parameter according to at least one emotion type; the adjustment parameter is used to adjust the speech feature parameter; The input module 302 is also used to input the text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters into a text-to-speech unit, transform the speech feature parameters of the intermediate speech after the text to be processed is converted or the speech feature parameters of the initial speech, and obtain the target speech. In a possible implementation, the acquisition module 301 is specifically used to accurately match and / or fuzzily match at least one emotion type of the text to be processed; wherein the accurate matching is based on matching the emotion type with the emotion words or emotion phrases in the text to be processed; and the fuzzy matching is based on matching the emotion type with the semantic information and / or context information of the text to be processed.

[0077] An embodiment of the present application provides a speech generation device that can generate speech with a variety of mixed emotions. When it is necessary to generate speech with non-predetermined emotions, the speech generation model structure is not changed, and the generated emotions of the speech generation model are quickly adjusted.

[0078] like Figure 4 As shown, an electronic device 400 provided in an embodiment of the present application includes: a processor 401, a memory 402 and a bus, wherein the memory 402 stores machine-readable instructions executable by the processor 401. When the electronic device is running, the processor 401 communicates with the memory 402 through the bus, and the processor 401 executes the machine-readable instructions to perform the steps of the above-mentioned speech generation method.

[0079] Specifically, the above-mentioned memory 402 and processor 401 can be general-purpose memories and processors, which are not specifically limited here. When the processor 401 runs the computer program stored in the memory 402, the above-mentioned speech generation method can be executed.

[0080] Corresponding to the above-mentioned speech generation method, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned speech generation method are executed.

[0081] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working process of the system and device described above can refer to the corresponding process in the method embodiment, and will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0082] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0083] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0084] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the information processing method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks, or optical disks.

[0085] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A speech generation method, characterized in that: The method comprises: Acquire a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; Acquire at least one corresponding adjustment parameter according to at least one emotion type; the adjustment parameter is used to adjust the speech feature parameter; The text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters are input into a text-to-speech unit, and the speech feature parameters of the intermediate speech converted from the text to be processed or the speech feature parameters of the initial speech are transformed to obtain a target speech.

2. The speech generation method according to claim 1, characterized in that: Obtain at least one emotion type corresponding to the text to be processed by the following steps: Based on the matching principle, matching at least one emotion type according to the content of the text to be processed; The matching principle refers to matching one or more emotion types based on emotion words, emotion phrases, semantic information, and context information in the text to be processed.

3. The speech generation method according to claim 2, characterized in that: The matching principle includes an exact matching principle and a fuzzy matching principle. The matching principle is based on the content of the text to be processed, and at least one emotion type is matched, including: Analyze the text to be processed based on the exact matching principle to obtain a most matching mandatory emotion type; and / or, analyzing the text to be processed based on the fuzzy matching principle to obtain at least two candidate emotion types with the highest matching degree; The mandatory emotion type and / or at least one candidate emotion type are determined as final emotion types.

4. The speech generation method according to claim 2, characterized in that: The acquiring at least one corresponding adjustment parameter according to at least one emotion type includes: According to at least one emotion type, inputting the initial speech into at least one emotion unit corresponding to the emotion type, and obtaining at least one initial adjustment parameter corresponding to at least one adjustment parameter type; Based on the fusion strategy, initial adjustment parameters of the same adjustment parameter type are fused to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type.

5. The speech generation method according to claim 4, characterized in that: The fusing of initial adjustment parameters of the same adjustment parameter type based on the fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type includes: Determining at least one initial fusion weight corresponding to each emotion type according to the matching degree between the text to be processed and each emotion type; According to at least one fusion weight corresponding to each emotion type, a target fusion weight corresponding to each emotion type is obtained; For any adjustment parameter type, based on the target fusion weight corresponding to each emotion type, at least one initial adjustment parameter of the adjustment parameter type corresponding to each emotion type is weightedly fused to obtain the target adjustment parameter corresponding to the adjustment parameter type.

6. The speech generation method according to claim 4, characterized in that: The method of fusing the initial adjustment parameters of the same adjustment parameter type based on the fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type further includes: For any adjustment parameter type, at least one initial adjustment parameter corresponding to the adjustment parameter type is averaged and fused to obtain a target adjustment parameter corresponding to the adjustment parameter type.

7. The speech generation method according to any one of claims 4 to 6, characterized in that: The speech generation model includes at least two different emotion units and the text-to-speech unit, and the target emotion unit corresponding to the target emotion type in the speech generation model is trained by the following steps: Acquire at least one corresponding speech training sample, wherein the speech training sample includes at least a first speech sample and a second speech sample, wherein the first speech sample is a neutral speech obtained by processing a text-to-speech model, and the second speech sample is a speech corresponding to the target emotion type, and the text content of the second speech sample is the same as the text content of the first speech sample; The target emotion unit is trained by using the at least one speech training sample, and then the verification parameter of the target emotion unit is evaluated based on at least a test speech sample. If the evaluation result of the target emotion unit meets a preset verification condition, the training of the target emotion unit is completed.

8. The speech generation method according to claim 7, characterized in that: The training processes of the emotion units are independent of each other, and the training process of the emotion unit is independent of the training process of the text-to-speech unit.

9. The speech generation method according to claim 7, characterized in that: The speech training sample also includes a third speech sample, the text content of the third speech sample is the same as the text content of the first speech sample and the text content of the second speech sample, and the similarity between the emotion type of the third speech sample and the emotion type of the second speech sample is greater than a similarity threshold.

10. A speech generating device, characterized in that: The device comprises: An acquisition module, used for acquiring a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; An input module, used for acquiring at least one corresponding adjustment parameter according to at least one emotion type; the adjustment parameter is used for adjusting the speech feature parameter; The input module is also used to input the text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters into a text-to-speech unit, transform the speech feature parameters of the intermediate speech converted from the text to be processed or the speech feature parameters of the initial speech to obtain the target speech.

Citation Information

Patent Citations

  • Speech synthesis method and system and terminal equipment

    CN108615524A

  • Intelligent robot-oriented voice synthesis method and device

    CN109461435A

  • Emotional speech synthesis method and device, electronic equipment and storage medium

    CN119649795A

  • Text-to-speech with emotional content

    US20160078859A1