Voice generation method and device

By obtaining the text to be processed and the emotion type, and using the matching principle and fusion strategy to generate adjustment parameters, the problem that traditional speech generation technology cannot generate mixed speech with multiple emotions is solved, and the generated emotions of the speech generation model can be quickly adjusted, thereby improving the user experience.

CN120808748AActive Publication Date: 2025-10-17SHANGHAI XIYU JIZHI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511060135.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-10-17
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Traditional speech generation technology cannot generate speech with mixed emotions, and when non-predetermined emotions need to be generated, it is impossible to quickly adjust the generated emotions of the speech generation model without changing the model structure.

Method used

By obtaining the text to be processed and the emotion type, the matching principle and fusion strategy are used to generate adjustment parameters, adjust the speech feature parameters, generate speech with multiple mixed emotions, including exact matching and fuzzy matching, and use the emotion unit and text-to-speech unit for independent training to generate the target speech.

Benefits of technology

It achieves the rapid adjustment of generated emotions without changing the structure of the speech generation model, and generates speech with multiple mixed emotions, thereby improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808748A_ABST
    Figure CN120808748A_ABST
Patent Text Reader

Abstract

The invention relates to the field of voice generation, in particular to a voice generation method and device, and the method comprises the steps: obtaining a to-be-processed text, an initial voice corresponding to the to-be-processed text, and at least one emotion type; obtaining at least one corresponding adjustment parameter according to the at least one emotion type; and inputting the to-be-processed text and the adjustment parameter, or the initial voice and the adjustment parameter into a character-to-voice unit, and converting the voice feature parameter of the intermediate voice converted from the to-be-processed text or the voice feature parameter of the initial voice to obtain a target voice. According to the invention, the voice with various mixed emotions can be generated, the structure of the voice generation model is not changed when the voice with non-predetermined emotions needs to be generated, and the generated emotions of the voice generation model are rapidly adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This case is a divisional application of the invention patent application number 202510503104.2, filed on April 22, 2025, entitled "A Method and Apparatus for Speech Generation." The present invention relates to the field of speech generation, and more specifically, to a method and apparatus for speech generation. Background Art

[0002] Speech generation is a technology that converts text into speech. This technology is widely used in various scenarios, such as AI (artificial intelligence) companion chat. In each of these scenarios, we aim to generate richer, more immersive speech for users, enhancing the user experience.

[0003] However, traditional speech generation technology can only provide users with speech with a few predetermined emotions, and cannot provide speech with multiple mixed emotions. It is also unable to generate speech with non-predetermined emotions without changing the speech generation model structure, and cannot quickly adjust the generated emotions of the speech generation model. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a speech generation method and device that can generate speech with multiple mixed emotions, and when it is necessary to generate speech with non-predetermined emotions, the speech generation model structure does not need to be changed, and the generated emotions of the speech generation model can be quickly adjusted.

[0005] In a first aspect, an embodiment of the present application provides a speech generation method, the speech generation method comprising: Obtaining a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; Acquire at least one corresponding adjustment parameter according to at least one emotion type; the adjustment parameter is used to adjust the speech feature parameter; The text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters are input into a text-to-speech unit, and the speech feature parameters of the intermediate speech converted from the text to be processed or the speech feature parameters of the initial speech are transformed to obtain the target speech.

[0006] In a possible implementation, at least one emotion type corresponding to the text to be processed is obtained through the following steps: Based on a matching principle, matching at least one emotion type according to the content of the text to be processed; The matching principle refers to matching one or more emotion types based on emotion words, emotion phrases, semantic information, and context information in the text to be processed.

[0007] In a possible implementation, the matching principles include an exact matching principle and a fuzzy matching principle, and the matching at least one emotion type based on the matching principles and according to the text content to be processed includes: analyzing the text content to be processed based on the exact matching principle to obtain one mandatory emotion type with the highest matching degree; and / or, analyzing the text content to be processed based on the fuzzy matching principle to obtain at least two optional emotion types with the highest matching degrees; determining the mandatory emotion type and / or the at least one optional emotion type as the final emotion type.

[0008] In a possible implementation, the obtaining the at least one adjustment parameter corresponding to the at least one emotion type includes: inputting the initial voice into at least one emotion unit corresponding to the at least one emotion type according to the at least one emotion type to obtain at least one initial adjustment parameter of at least one adjustment parameter type; fusing the initial adjustment parameters of the same adjustment parameter type based on a fusion strategy to obtain at least one target adjustment parameter of the at least one adjustment parameter type.

[0009] In a possible implementation, the fusing the initial adjustment parameters of the same adjustment parameter type based on the fusion strategy to obtain the at least one target adjustment parameter of the at least one adjustment parameter type includes: determining at least one initial fusion weight corresponding to each emotion type according to the matching degree between the text content to be processed and each emotion type; obtaining a target fusion weight corresponding to each emotion type according to the at least one initial fusion weight corresponding to each emotion type; for any adjustment parameter type, performing weighted fusion on at least one initial adjustment parameter of the adjustment parameter type corresponding to each emotion type based on the target fusion weight corresponding to each emotion type to obtain a target adjustment parameter of the adjustment parameter type.

[0010] In a possible implementation, the fusing the initial adjustment parameters of the same adjustment parameter type based on the fusion strategy to obtain the at least one target adjustment parameter of the at least one adjustment parameter type further includes: for any adjustment parameter type, performing average fusion on at least one initial adjustment parameter of the adjustment parameter type to obtain a target adjustment parameter of the adjustment parameter type.

[0011] In a possible implementation, the speech generation model comprises at least two different emotion units and the text-to-speech unit, and a target emotion unit corresponding to a target emotion type in the speech generation model is trained through the following steps: obtaining at least one corresponding speech training sample, the speech training sample comprising at least a first speech sample and a second speech sample, the first speech sample being neutral speech obtained by processing a text-to-speech model, and the second speech sample being speech corresponding to the target emotion type and having the same text content as the first speech sample; training the target emotion unit through the at least one speech training sample, evaluating a verification parameter of the target emotion unit based at least on a test speech sample, and if the evaluation result of the target emotion unit meets a preset verification condition, the training of the target emotion unit is completed.

[0012] In a possible implementation, the training processes of the emotion units are independent of each other, and the training process of each emotion unit is independent of the training process of the text-to-speech unit.

[0013] In a possible implementation, the speech training sample further comprises a third speech sample, the text content of the third speech sample being the same as that of the first speech sample and the second speech sample, and the similarity between the emotion type of the third speech sample and the emotion type of the second speech sample being greater than a similarity threshold.

[0014] In a second aspect, the embodiments of the present application further provide a speech generation device, comprising: a obtaining module configured to obtain a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; an input module configured to obtain at least one adjustment parameter corresponding to the at least one emotion type, the adjustment parameter being used to adjust a speech feature parameter; the input module is further configured to input the text to be processed and the adjustment parameter, or the initial speech and the adjustment parameter into a text-to-speech unit, transform a speech feature parameter of an intermediate speech converted from the text to be processed or a speech feature parameter of the initial speech, and obtain a target speech.

[0015] The embodiment of the present application provides a speech generation method and device, the method comprises the following steps: obtaining a to-be-processed text, an initial speech corresponding to the to-be-processed text and at least one emotion type; obtaining at least one adjustment parameter corresponding to the at least one emotion type according to the at least one emotion type; inputting the to-be-processed text and the adjustment parameter, or the initial speech and the adjustment parameter into a text-to-speech unit; and transforming speech feature parameters of an intermediate speech converted from the to-be-processed text or speech feature parameters of the initial speech, to obtain a target speech. The present application can generate a speech with mixed emotions, and the speech generation model structure is not changed when a speech with a non-predefined emotion needs to be generated, so that the generation emotion of the speech generation model can be quickly adjusted. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor under the premise of these drawings.

[0017] Figure 1 A flow chart of a speech generation method provided by the embodiment of the present application is shown; Figure 2 A training flow chart of an emotion unit provided by the embodiment of the present application is shown; Figure 3 A structural schematic diagram of a speech generation device provided by the embodiment of the present application is shown; Figure 4 A structural schematic diagram of an electronic device provided by the embodiment of the present application is shown. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solutions and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only play the purpose of illustration and description, and do not limit the protection scope of the present application. In addition, it should be understood that the schematic drawings are not drawn according to the actual proportion. The flow chart shows the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flow chart can not be implemented in sequence, and the steps without logical context relationship can be reversed in sequence or implemented simultaneously. In addition, one or more other operations can be added to the flow chart or removed from the flow chart under the guidance of the content of the present application by those skilled in the art.

[0019] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.

[0020] To enable those skilled in the art to use the present disclosure, the following embodiments are provided in conjunction with a specific application scenario, the field of speech generation. Those skilled in the art will appreciate that the general principles defined herein may be applied to other embodiments and application scenarios without departing from the spirit and scope of this disclosure. While this disclosure is primarily described in the field of speech generation, it should be understood that this is merely an exemplary embodiment.

[0021] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.

[0022] The following is a detailed description of a speech generation method provided in an embodiment of the present application.

[0023] Reference Figure 1 FIG. 1 is a flow chart of a speech generation method provided in an embodiment of the present application. The exemplary steps of the embodiment of the present application are described below: S101: Acquire a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type.

[0024] In the embodiments of the present application, the text to be processed refers to the text for which speech needs to be generated; the text to be processed can be text input by the user, or it can be text used to reply to the user in a conversation scenario. The initial speech is a neutral speech obtained by inputting the text to be processed into any trained text-to-speech model. Neutral speech refers to speech without any emotion, that is, emotionless speech. Therefore, the initial speech does not contain any emotion and is emotionless speech. Emotion types include but are not limited to happiness, anger, indifference, anxiety, surprise, etc., and can be any emotion that changes the speech feature parameters. Among them, the speech feature parameters may include tone, intonation, and speaking speed. Tone refers to the attitude and emotional tendency expressed by the speaker through speech. Intonation refers to the rise and fall of the voice when speaking. Speaking speed refers to the speed of the language when speaking.

[0025] The text to be processed may correspond to one or more emotion types, which may be obtained by using one or more methods selected from a matching principle, user specification, and an emotion analysis model. User specification refers to one or more emotion types selected by the user from among several emotion type options.

[0026] A matching principle is adopted to obtain at least one emotion type corresponding to the text to be processed, including: matching at least one emotion type based on the content of the text to be processed based on the matching principle; wherein the matching principle refers to matching one or more emotion types based on emotion vocabulary, emotion phrases, semantic information, and contextual information in the text to be processed.

[0027] In an embodiment of the present application, the matching principles include an exact matching principle and a fuzzy matching principle, wherein the text to be processed is analyzed based on the exact matching principle to obtain the most matching required emotion type; and / or the text to be processed is analyzed based on the fuzzy matching principle to obtain at least two alternative emotion types with the highest degree of matching; the required emotion type and / or at least one alternative emotion type are determined as the final emotion type.

[0028] Specifically, the emotional words and / or emotional phrases in the text to be processed are analyzed based on an exact matching principle to obtain a single mandatory emotion type that best matches the text. The semantic information and / or contextual information in the text to be processed are analyzed based on a fuzzy matching principle to obtain at least two candidate emotion types with the highest degree of matching. The candidate emotion types are sent to the user so that the user can select at least one candidate emotion type. The mandatory emotion type and the at least one candidate emotion type selected by the user are determined as the final emotion type.

[0029] This involves using AI models based on matching principles to match at least one emotion type to the content of the text being processed. For example, GPT4 (Generative Pre-trained Transformer 4) or BERT (Bidirectional Encoder Representations from Transformers) both support precise matching of a single emotion type and are part of the sentiment analysis module in natural language processing. BERT also supports fuzzy matching of at least one emotion based on the text content, meaning that a single text can be mapped to multiple emotion types, or several highly matched emotion types can be calculated simultaneously as alternatives.

[0030] The emotion analysis model is used to obtain at least one emotion type corresponding to the to-be-processed text, including: in a dialogue scenario, determining the content and emotion type of the voice sent by the user, or the content and emotion type of the text sent by the user; the to-be-processed text refers to the text used to reply to the user; the content of the voice or the content of the text sent by the user and the to-be-processed text form current dialogue content; historical dialogue content with a similarity greater than a preset similarity to the current dialogue content is determined; and the emotion type corresponding to the to-be-processed text is determined according to a preset number of historical emotion types in the historical dialogue content that reply to the user.

[0031] In S102, at least one adjustment parameter corresponding to the at least one emotion type is obtained; the adjustment parameter is used to adjust the voice feature parameter.

[0032] In the embodiments of the present application, the initial voice is input into at least one emotion unit corresponding to the emotion type according to the at least one emotion type, to obtain at least one initial adjustment parameter corresponding to at least one adjustment parameter type; and the initial adjustment parameters of the same adjustment parameter type are fused based on a fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type.

[0033] Each emotion type corresponds to an emotion unit; the emotion unit corresponding to the emotion type is used to determine the adjustment parameter corresponding to at least one adjustment parameter type under the emotion type. Each emotion type corresponds to at least one adjustment parameter type; the adjustment parameter types corresponding to different emotion types can be the same or different. The adjustment parameter type refers to the type of the adjustment parameter. The voice feature parameter refers to a voice parameter capable of expressing emotion.

[0034] For example, the emotion type includes happy and surprised, the initial voice is input into the emotion unit corresponding to happy to obtain the adjustment parameter of at least one adjustment parameter type under the emotion type of happy; and the initial voice is input into the emotion unit corresponding to surprised to obtain the adjustment parameter of at least one adjustment parameter type under the emotion type of surprised.

[0035] In the embodiments of the present application, two fusion strategies are provided, which are a weighted fusion manner and an average fusion manner. The initial adjustment parameters of the same adjustment parameter type are fused based on the fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type, including: the initial adjustment parameters of the same adjustment parameter type are fused based on the weighted fusion manner or the average fusion manner to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type.

[0036] Further, the initial adjustment parameters of the same adjustment parameter type are fused based on the weighted fusion manner to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type, including: Step one, according to the matching degree of the text to be processed and each emotion type, determine at least one initial fusion weight corresponding to each emotion type.

[0037] In the embodiments of the present application, the higher the matching degree with the text to be processed, the greater the fusion weight corresponding to the emotion type. The present application provides three ways to determine the initial fusion weight corresponding to the emotion type. The greater the fusion weight corresponding to the emotion type, the greater the adjustment force of the initial adjustment parameter corresponding to the emotion type to the voice.

[0038] The first way to determine the initial fusion weight corresponding to the emotion type is to determine the initial fusion weight corresponding to each emotion type according to the preset weight setting rule based on the matching degree of the text to be processed and each emotion type. The preset weight setting rule explains how to determine the fusion weight corresponding to each emotion type based on the matching degree of the text to be processed and each emotion type.

[0039] For example, the preset weight setting rule can be that the fusion weight corresponding to the emotion type with the highest matching degree is 0.5, and the fusion weight corresponding to the emotion type with the second to fifth highest matching degree is 0.1. Assuming that the emotion type with the highest matching degree is happy, the initial fusion weight corresponding to the emotion type corresponding to happy is 0.5, and the initial fusion weight corresponding to the emotion type corresponding to the second to fifth matching degree is 0.1.

[0040] The second way to determine the initial fusion weight corresponding to the emotion type is to add the matching degrees of the text to be processed and all emotion types to obtain a sum value; and to obtain the initial fusion weight corresponding to each emotion type by dividing the matching degree of the text to be processed and each emotion type by the sum value.

[0041] For example, the emotion types include happy, excited and surprised, the matching degree of the text to be processed and the happy emotion type is 0.8, the matching degree of the text to be processed and the excited emotion type is 0.7, and the matching degree of the text to be processed and the surprised emotion type is 0.6. Add the matching degrees of the text to be processed and all emotion types to obtain a sum value, i.e. sum value = 0.8 + 0.7 + 0.6 = 2.1. Based on the sum value, the weight is normalized calculated, the fusion weight corresponding to the happy emotion type = 0.8 / 2.1 = 8 / 21. The fusion weight corresponding to the excited emotion type = 0.7 / 2.1 = 1 / 3. The fusion weight corresponding to the surprised emotion type = 0.6 / 2.1 = 2 / 7.

[0042] The third way to determine the fusion weight corresponding to the emotion type is to determine the initial fusion weight corresponding to each emotion type as the preset fusion weight corresponding to each emotion type set by the user.

[0043] Step two, obtaining a target fusion weight corresponding to each emotion type according to at least one fusion weight corresponding to each emotion type respectively.

[0044] In the embodiment of the present application, all fusion weights corresponding to each emotion type are weighted and summed or averaged to determine the target fusion weight corresponding to each emotion type.

[0045] Step three, for any adjustment parameter type, based on the target fusion weight corresponding to each emotion type, at least one initial adjustment parameter corresponding to the adjustment parameter type of each emotion type is weighted and fused to obtain the target adjustment parameter value corresponding to the adjustment parameter type.

[0046] Further, the initial adjustment parameters of the same adjustment parameter type are fused based on the average fusion manner to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type, including: for any adjustment parameter type, at least one initial adjustment parameter corresponding to the adjustment parameter type is averaged and fused to obtain the target adjustment parameter corresponding to the adjustment parameter type.

[0047] In the embodiment of the present application, if the user does not use the above three ways to determine the weight corresponding to the emotion type, the multiple emotion parameters are averaged and fused.

[0048] S103, input the to-be-processed text and the adjustment parameter, or the initial speech and the adjustment parameter into the text-to-speech unit, transform the speech feature parameter of the intermediate speech converted from the to-be-processed text or the speech feature parameter of the initial speech to obtain the target speech.

[0049] In the embodiment of the present application, the to-be-processed text and the adjustment parameter are input into the text-to-speech unit, the speech feature parameter of the intermediate speech converted from the to-be-processed text is transformed to obtain the target speech. Or the initial speech and the adjustment parameter are input into the text-to-speech unit, the speech feature parameter of the initial speech is transformed to obtain the target speech.

[0050] Here, the text-to-speech unit includes at least one intermediate layer, and different intermediate layers transform the speech feature parameters of the intermediate speech converted from the to-be-processed text or the speech feature parameters of the initial speech based on adjustment parameters of different adjustment parameter types; one intermediate layer transforms the speech feature parameters of the intermediate speech converted from the to-be-processed text or the speech feature parameters of the initial speech based on adjustment parameters of one adjustment parameter type. Specifically, the transformation of the speech feature parameters of the intermediate speech converted from the to-be-processed text or the speech feature parameters of the initial speech includes: the i th intermediate layer transforms the speech feature parameters of the intermediate speech converted from the to-be-processed text or the speech feature parameters of the initial speech based on corresponding adjustment parameters; i = i + 1; jump to "the i th intermediate layer transforms the speech feature parameters of the intermediate speech converted from the to-be-processed text or the speech feature parameters of the initial speech based on corresponding adjustment parameters" to continue execution; wherein the initial value of i is 1, and the maximum value of i is the number of intermediate layers.

[0051] Further, the speech generation model in the embodiment of the present application includes at least two different emotion units and a text-to-speech unit. The text-to-speech unit and all emotion units are connected respectively.

[0052] In the present application, the emotion unit delivers several adjustment parameters to the speech generation model, so that the emotional information is integrated with the final generated speech content, and the speech content that can reflect the emotion is generated. The training processes of each emotion unit are independent of each other, and the training process of the emotion unit is independent of the training process of the text-to-speech unit.

[0053] The text-to-speech unit can be a common TTS (text-to-speech) model, such as Tacotron series, DeepSpeech series, FastSpeech series, TransformerTTS, VALL-E series, TTS model based on generative flow or adversarial network, etc. which can convert text into speech, and no limitation is made on this.

[0054] Each emotion unit is composed of one or more multilayer perceptron modules (MLP) or convolutional neural network CNN modules, preferably, the emotion unit is composed of one multilayer perceptron module or convolutional neural network.

[0055] The selected one or more emotion units output different adaptive adjustment parameters to the text-to-speech unit to embed emotional information in the speech, rather than inputting the same parameters to each layer in the text-to-speech unit.

[0056] The connection between the text-to-speech model and each emotion unit is achieved through a structure based on AdaLN (Adaptive Layer Normalization) or LoRA (Low-Rank Adaptation).

[0057] The emotion unit transmits several adjustment parameters to the text-to-speech unit, integrating the emotion information with the resulting speech content. Specifically, each emotion unit outputs multiple adaptive adjustment parameters corresponding to the corresponding emotion type. These fine-tuning parameters are then injected into different intermediate layers within the text-to-speech unit. The results generated by each layer in the text-to-speech unit are then linearly transformed to adjust parameters such as intonation, speech rate, and tone of the generated speech. Note that this step does not alter the text content, only the speech parameters.

[0058] Furthermore, the text-to-speech unit is trained through the following steps: obtaining at least one text training sample and a converted voice label corresponding to each text training sample, wherein the several text training samples refer to texts that need to be converted into voice; the converted voice label refers to the voice obtained by converting the text training samples through the trained text-to-speech model; the text-to-speech unit is trained based on the text training samples and the converted voice label corresponding to each text training sample; and the verification parameters of the target emotion unit are evaluated based on at least the test text sample. If the evaluation result of the target emotion unit meets the preset verification conditions, the training of the target emotion unit is completed.

[0059] Specifically, the verification parameter of the target emotion unit is determined based on the test text sample, or the set of the test text sample and the verification text sample. If the evaluation result of the target emotion unit meets the preset verification condition, the training of the target emotion unit is completed.

[0060] Here, when training the text-to-speech unit, the text-to-speech unit is isolated from all emotion units, and only the text-to-speech unit is trained, and any parameters of the emotion unit are not changed due to the training of the text-to-speech unit.

[0061] Further, refer to Figure 2 As shown, it is a training flow chart of the emotion unit provided in an embodiment of the present application.

[0062] S201: Obtain at least one corresponding speech training sample.

[0063] The voice training sample at least includes a first voice sample and a second voice sample. The first voice sample is a neutral voice obtained by processing a text-to-speech model. The second voice sample is a voice corresponding to a target emotion type, and the text content of the second voice sample is the same as that of the first voice sample. The neutral voice refers to a voice without any emotion. The target emotion type refers to an emotion type corresponding to a target emotion unit to be trained.

[0064] In addition, the voice training sample can further include a third voice sample. The text content of the third voice sample is the same as that of the first voice sample and the second voice sample, and the similarity of the emotion type of the third voice sample to the emotion type of the second voice sample is greater than a similarity threshold.

[0065] In the embodiments of the present application, the similarity between two emotion types refers to the similarity between the voice features of two voices with the same text content.

[0066] Here, since human emotions have continuity and mixability, for example, the voice feature parameters of the voices of the emotion types of "happy" and "excited" are partially coincident. Therefore, some similar voice training samples are introduced as negative training samples to strengthen the understanding of the emotion unit corresponding to "happy" to "happy".

[0067] S202, training the target emotion unit through at least one voice training sample, evaluating a verification parameter of the target emotion unit based on at least a test voice sample, and if the evaluation result of the target emotion unit meets a preset verification condition, the training of the target emotion unit is completed.

[0068] In the embodiments of the present application, the first voice sample is input into the target emotion unit to obtain at least one adjustment parameter corresponding to the target emotion type. The at least one adjustment parameter corresponding to the target emotion type is input into the trained text-to-speech unit to transform the first voice sample to obtain a generated voice. The target emotion unit is updated according to the generated voice and the second voice sample, so that the generated voice can approach the second voice sample. Or, the target emotion unit is updated according to the generated voice, the second voice sample and the third voice sample, so that the generated voice can approach the second voice sample and move away from the third voice sample. The verification parameter of the target emotion unit is evaluated based on at least a test voice sample. If the evaluation result of the target emotion unit meets a preset verification condition, the training of the target emotion unit is completed.

[0069] Specifically, the verification parameter of the target emotion unit is determined based on a test voice sample or a set of test voice samples and verification voice samples. If the evaluation result of the target emotion unit meets a preset verification condition, the training of the target emotion unit is completed.

[0070] The verification parameters include at least one of consistency verification, confusion rate verification, and error rate verification.

[0071] For example, an emotion unit corresponding to a happy emotion type is trained, and the happy emotion unit is trained according to the first voice sample, the second voice sample of the happy emotion type, and the third voice sample, so that the happy emotion unit can output the adjustment parameter for making the emotion of the voice "speak happily".

[0072] Here, the trained emotion units are obtained, and each emotion unit can accurately convert the voice without emotion into the voice with the corresponding emotion. The mutual isolation training between the text-to-voice unit and the emotion unit and between each emotion unit can also make the training process simpler and reduce the complexity of the training.

[0073] In addition, if the emotion types covered by the trained emotion units are not enough, and the emotion unit needs to be added, the text-to-voice unit and other emotion units can be isolated, and the voice training sample corresponding to the emotion type added alone is trained according to steps S201-S202.

[0074] The embodiment of the application provides a voice generation method and device, the method comprising: obtaining a to-be-processed text, an initial voice corresponding to the to-be-processed text, and at least one emotion type; obtaining at least one adjustment parameter corresponding to the at least one emotion type; inputting the to-be-processed text and the adjustment parameter, or the initial voice and the adjustment parameter into a text-to-voice unit, and transforming the voice feature parameter of the intermediate voice converted from the to-be-processed text or the voice feature parameter of the initial voice to obtain a target voice. The application can generate a voice with mixed emotions, and does not change the voice generation model structure when generating a voice with a non-predefined emotion, and quickly adjusts the generation emotion of the voice generation model.

[0075] Based on the same inventive concept, the embodiment of the application also provides a voice generation device corresponding to the voice generation method. Since the principle of the device solving the problem in the embodiment of the application is similar to the voice generation method described above, the implementation of the device can be referred to the implementation of the method, and the repeated parts will not be described here.

[0076] Referring to Figure 3 As shown in FIG. 1, a voice generation device provided by the embodiment of the application comprises: The obtaining module 301 is configured to obtain a to-be-processed text, an initial voice corresponding to the to-be-processed text, and at least one emotion type. The input module 302 is configured to obtain at least one adjustment parameter corresponding to the at least one emotion type, and the adjustment parameter is used to adjust the voice feature parameter. The input module 302 is further configured to input the to-be-processed text and the adjustment parameter, or the initial voice and the adjustment parameter into a text-to-speech unit, transform the speech feature parameter of the intermediate voice converted from the to-be-processed text or the speech feature parameter of the initial voice, and obtain a target voice. In a possible implementation, the obtaining module 301 is specifically configured to accurately match and / or fuzzily match at least one emotion type of the to-be-processed text; the accurate matching is based on matching an emotion type based on an emotion word or an emotion phrase in the to-be-processed text; and the fuzzy matching is based on matching an emotion type based on semantic information and / or context information of the to-be-processed text.

[0077] The embodiment of the present application provides a speech generation device, which can generate a speech with mixed emotions, and does not change the structure of a speech generation model when a speech with a non-predefined emotion needs to be generated, and quickly adjusts the generated emotion of the speech generation model.

[0078] As shown in Figure 4 The embodiment of the present application provides an electronic device 400, which includes a processor 401, a memory 402 and a bus. The memory 402 stores machine readable instructions executable by the processor 401. When the electronic device is running, the processor 401 and the memory 402 communicate through the bus. The processor 401 executes the machine readable instructions to perform the steps of the above speech generation method.

[0079] Specifically, the memory 402 and the processor 401 can be general memory and processor, which are not specifically limited here. When the processor 401 runs the computer program stored in the memory 402, the above speech generation method can be executed.

[0080] Corresponding to the above speech generation method, the embodiment of the present application further provides a computer readable storage medium, which stores a computer program. When the computer program is run by a processor, the steps of the above speech generation method are executed.

[0081] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working process of the system and the device described above can refer to the corresponding process in the method embodiment, and will not be repeated in the present application. In the several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented by other means. The above-described device embodiments are only schematic, for example, the division of the modules is only a logical function division, and the actual implementation can have another division manner, for example, a plurality of modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed mutual elements can be indirect coupling or communication connection through some communication interface, device or module, which can be electrical, mechanical or other forms.

[0082] The modules described as separate components can or can not be physically separated, and the components shown as modules can or can not be physical units, i.e. can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0083] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.

[0084] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer readable storage medium executable by a processor. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the information processing method described in each embodiment of the present application. The foregoing storage medium includes: U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk and various program code storage media.

[0085] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A speech generation method, characterized in that: The method comprises: Obtaining a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; Obtaining at least one corresponding adjustment parameter according to at least one emotion type; the adjustment parameter is used to adjust the speech feature parameter; wherein each emotion type corresponds to an emotion unit, and the emotion unit corresponding to the emotion type is used to determine the adjustment parameter corresponding to at least one adjustment parameter type under the emotion type; Inputting the to-be-processed text and the adjustment parameters, or the initial speech and the adjustment parameters, into a text-to-speech unit, and transforming speech feature parameters of the intermediate speech converted from the to-be-processed text or the initial speech to obtain a target speech; The speech generation model includes at least two different emotion units and the text-to-speech unit, and the target emotion unit corresponding to the target emotion type in the speech generation model is trained by the following steps: obtaining at least one corresponding speech training sample, the speech training sample including a first speech sample and a third speech sample, the first speech sample is a neutral speech obtained by processing the text-to-speech model, the text content of the third speech sample is the same as the text content of the first speech sample and the text content of the second speech sample, and the similarity between the emotion type of the third speech sample and the emotion type of the second speech sample is greater than a similarity threshold; the target emotion unit is trained by the at least one speech training sample, and then the verification parameters of the target emotion unit are evaluated based on at least a test speech sample. If the evaluation result of the target emotion unit meets the preset verification condition, the training of the target emotion unit is completed.

2. The speech generation method according to claim 1, wherein: Obtain at least one emotion type corresponding to the text to be processed by the following steps: Based on a matching principle, matching at least one emotion type according to the content of the text to be processed; The matching principle refers to matching one or more emotion types based on emotion words, emotion phrases, semantic information, and context information in the text to be processed.

3. The speech generation method according to claim 2, wherein: The matching principle includes an exact matching principle and a fuzzy matching principle. The matching principle is based on the content of the text to be processed, and matches at least one emotion type, including: Analyze the text to be processed based on the exact matching principle to obtain a most matching mandatory emotion type; and / or, analyzing the text to be processed based on the fuzzy matching principle to obtain at least two candidate emotion types with the highest matching degree; The mandatory emotion type and / or at least one candidate emotion type are determined as final emotion types.

4. The speech generation method according to claim 2, wherein: The acquiring at least one corresponding adjustment parameter according to at least one emotion type includes: According to at least one emotion type, inputting the initial speech into at least one emotion unit corresponding to the emotion type to obtain at least one initial adjustment parameter corresponding to at least one adjustment parameter type; Initial adjustment parameters of the same adjustment parameter type are fused based on a fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type.

5. The speech generation method according to claim 4, characterized in that The fusing of initial adjustment parameters of the same adjustment parameter type based on the fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type includes: Determining at least one initial fusion weight corresponding to each emotion type according to the matching degree between the text to be processed and each emotion type; According to at least one fusion weight corresponding to each emotion type, a target fusion weight corresponding to each emotion type is obtained; For any adjustment parameter type, based on the target fusion weight corresponding to each emotion type, at least one initial adjustment parameter of the adjustment parameter type corresponding to each emotion type is weightedly fused to obtain the target adjustment parameter corresponding to the adjustment parameter type.

6. The speech generation method according to claim 4, characterized in that The step of fusing initial adjustment parameters of the same adjustment parameter type based on a fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type further includes: For any adjustment parameter type, at least one initial adjustment parameter corresponding to the adjustment parameter type is averaged and fused to obtain a target adjustment parameter corresponding to the adjustment parameter type.

7. The speech generation method according to any one of claims 4 to 6, characterized in that: The second voice sample is a voice corresponding to the target emotion type, and the text content of the second voice sample is the same as the text content of the first voice sample.

8. The speech generation method according to claim 7, characterized in that The training process of each emotion unit is independent of each other, and the training process of the emotion unit is independent of the training process of the text-to-speech unit.

9. A speech generating device, characterized in that: The device comprises: An acquisition module, configured to acquire a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; An input module, configured to obtain at least one adjustment parameter corresponding to at least one emotion type; the adjustment parameter is used to adjust a speech feature parameter; wherein each emotion type corresponds to an emotion unit, and the emotion unit corresponding to the emotion type is used to determine an adjustment parameter corresponding to at least one adjustment parameter type under the emotion type; The input module is further configured to input the text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters, into a text-to-speech unit, and transform the speech feature parameters of the intermediate speech converted from the text to be processed or the speech feature parameters of the initial speech to obtain a target speech; The speech generation model includes at least two different emotion units and the text-to-speech unit, and the input module trains the target emotion unit corresponding to the target emotion type in the speech generation model through the following steps: obtaining at least one corresponding speech training sample, the speech training sample including a first speech sample and a third speech sample, the first speech sample is a neutral speech obtained by processing the text-to-speech model, the text content of the third speech sample is the same as the text content of the first speech sample and the text content of the second speech sample, and the similarity between the emotion type of the third speech sample and the emotion type of the second speech sample is greater than a similarity threshold; training the target emotion unit through the at least one speech training sample, and then evaluating the verification parameters of the target emotion unit based on at least a test speech sample. If the evaluation result of the target emotion unit meets the preset verification condition, the training of the target emotion unit is completed.

Citation Information

Patent Citations

  • Speech synthesis method and device, electronic equipment and storage medium

    CN114678003A

  • Voice emotion adjustment method, device, equipment and product

    CN118197331A

  • Method and apparatus for processing speech

    US20200388283A1