A method and device for generating speech
By obtaining the emotion types of the pending text and initial speech, and adjusting the speech feature parameters using matching principles and fusion strategies, the problem that traditional speech generation technology cannot generate multiple emotional mixed speech is solved, and the effect of quickly adjusting the generation of emotions is achieved.
Patent Information
- Application Number
- CN202510503104.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Traditional speech generation technology cannot generate speech with mixed emotions, and cannot quickly adjust the generated emotions of the speech generation model when it is necessary to generate non-determined emotions, without changing the model structure.
By obtaining the emotion types of pending text and initial speech, using matching principles and fusion strategies to obtain adjustment parameters, adjust speech feature parameters, and generate speech mixed speech.
It realizes that without changing the structure of the speech generation model, quickly adjusting the generated emotions, generating voices with mixed emotions, and improving the user experience.
Smart Images

Figure CN120015012B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech generation, and in particular to a speech generation method and device. Background Art
[0002] Speech generation is a technology that converts text into speech. This technology is widely used in various scenarios, such as AI (artificial intelligence) companion chat. In each of these scenarios, we aim to generate richer, more immersive speech for users, enhancing the user experience.
[0003] However, traditional speech generation technology can only provide users with speech with a few predetermined emotions, and cannot provide speech with multiple mixed emotions. It is also unable to generate speech with non-predetermined emotions without changing the speech generation model structure, and cannot quickly adjust the generated emotions of the speech generation model. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a speech generation method and device that can generate speech with multiple mixed emotions, and when it is necessary to generate speech with non-predetermined emotions, the speech generation model structure does not need to be changed, and the generated emotions of the speech generation model can be quickly adjusted.
[0005] In a first aspect, an embodiment of the present application provides a speech generation method, the speech generation method comprising:
[0006] Obtaining a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type;
[0007] Acquire at least one corresponding adjustment parameter according to at least one emotion type; the adjustment parameter is used to adjust the voice feature parameter;
[0008] The text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters are input into a text-to-speech unit, and the speech feature parameters of the intermediate speech converted from the text to be processed or the speech feature parameters of the initial speech are transformed to obtain the target speech.
[0009] In a possible implementation, at least one emotion type corresponding to the text to be processed is obtained through the following steps:
[0010] Based on a matching principle, matching at least one emotion type according to the content of the text to be processed;
[0011] The matching principle refers to matching one or more emotion types based on emotion words, emotion phrases, semantic information, and context information in the text to be processed.
[0012] In a possible implementation, the matching principle includes an exact matching principle and a fuzzy matching principle. The matching principle-based method matches at least one emotion type according to the content of the text to be processed, including:
[0013] Analyze the text to be processed based on the exact matching principle to obtain a most matching mandatory emotion type;
[0014] and / or, analyzing the text to be processed based on the fuzzy matching principle to obtain at least two candidate emotion types with the highest matching degree;
[0015] The mandatory emotion type and / or at least one candidate emotion type are determined as final emotion types.
[0016] In a possible implementation, obtaining at least one corresponding adjustment parameter according to at least one emotion type includes:
[0017] According to at least one emotion type, inputting the initial speech into at least one emotion unit corresponding to the emotion type to obtain at least one initial adjustment parameter corresponding to at least one adjustment parameter type;
[0018] Initial adjustment parameters of the same adjustment parameter type are fused based on a fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type.
[0019] In a possible implementation, fusing initial adjustment parameters of the same adjustment parameter type based on a fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type includes:
[0020] Determining at least one initial fusion weight corresponding to each emotion type according to the matching degree between the text to be processed and each emotion type;
[0021] According to at least one fusion weight corresponding to each emotion type, a target fusion weight corresponding to each emotion type is obtained;
[0022] For any adjustment parameter type, based on the target fusion weight corresponding to each emotion type, at least one initial adjustment parameter of the adjustment parameter type corresponding to each emotion type is weightedly fused to obtain the target adjustment parameter corresponding to the adjustment parameter type.
[0023] In a possible implementation, the fusing of initial adjustment parameters of the same adjustment parameter type based on a fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type further includes:
[0024] For any adjustment parameter type, at least one initial adjustment parameter corresponding to the adjustment parameter type is averaged and fused to obtain a target adjustment parameter corresponding to the adjustment parameter type.
[0025] In one possible implementation, the speech generation model includes at least two different emotion units and the text-to-speech unit, and the target emotion unit corresponding to the target emotion type in the speech generation model is trained by the following steps:
[0026] Obtaining at least one corresponding speech training sample, the speech training sample including at least a first speech sample and a second speech sample, the first speech sample being a neutral speech obtained by processing a text-to-speech model, the second speech sample being a speech corresponding to the target emotion type, and the text content of the second speech sample being the same as the text content of the first speech sample;
[0027] The target emotion unit is trained using the at least one speech training sample, and a verification parameter of the target emotion unit is evaluated based on at least a test speech sample. If the evaluation result of the target emotion unit meets a preset verification condition, the training of the target emotion unit is completed.
[0028] In a possible implementation, the training processes of the emotion units are independent of each other, and the training process of the emotion unit is independent of the training process of the text-to-speech unit.
[0029] In a possible implementation, the speech training sample further includes a third speech sample, the text content of the third speech sample is the same as the text content of the first speech sample and the text content of the second speech sample, and the similarity between the emotion type of the third speech sample and the emotion type of the second speech sample is greater than a similarity threshold.
[0030] In a second aspect, an embodiment of the present application further provides a speech generating device, the speech generating device comprising:
[0031] An acquisition module, configured to acquire a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type;
[0032] An input module, configured to obtain at least one corresponding adjustment parameter according to at least one emotion type; the adjustment parameter is used to adjust the speech feature parameter;
[0033] The input module is also used to input the text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters into the text-to-speech unit, transform the speech feature parameters of the intermediate speech converted from the text to be processed or the speech feature parameters of the initial speech to obtain the target speech.
[0034] An embodiment of the present application provides a speech generation method and apparatus, comprising: obtaining a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; obtaining at least one adjustment parameter corresponding to the at least one emotion type; inputting the text to be processed and the adjustment parameter, or the initial speech and the adjustment parameter, into a text-to-speech unit, and transforming the speech feature parameters of the intermediate speech after the text to be processed or the speech feature parameters of the initial speech to obtain a target speech. The present application is capable of generating speech with a variety of mixed emotions. When speech with non-predetermined emotions needs to be generated, the speech generation model structure does not need to be changed, and the generated emotion of the speech generation model can be quickly adjusted. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0036] Figure 1 A flow chart of a speech generation method provided by an embodiment of the present application is shown;
[0037] Figure 2 A training flow chart of an emotion unit provided in an embodiment of the present application is shown;
[0038] Figure 3 A schematic structural diagram of a speech generating device provided in an embodiment of the present application is shown;
[0039] Figure 4 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of illustration and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowcharts can be implemented out of sequence, and steps without logical context can be reversed or implemented simultaneously. In addition, those skilled in the art, under the guidance of the contents of this application, can add one or more other operations to the flowchart, or remove one or more operations from the flowchart.
[0041] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.
[0042] To enable those skilled in the art to use the present disclosure, the following embodiments are provided in conjunction with a specific application scenario, the field of speech generation. Those skilled in the art will appreciate that the general principles defined herein may be applied to other embodiments and application scenarios without departing from the spirit and scope of this disclosure. While this disclosure is primarily described in the field of speech generation, it should be understood that this is merely an exemplary embodiment.
[0043] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.
[0044] The following is a detailed description of a speech generation method provided in an embodiment of the present application.
[0045] Reference Figure 1 FIG. 1 is a flow chart of a speech generation method provided in an embodiment of the present application. The exemplary steps of the embodiment of the present application are described below:
[0046] S101: Acquire a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type.
[0047] In the embodiments of the present application, the text to be processed refers to the text for which speech needs to be generated; the text to be processed can be text input by the user, or it can be text used to reply to the user in a conversation scenario. The initial speech is a neutral speech obtained by inputting the text to be processed into any trained text-to-speech model. Neutral speech refers to speech without any emotion, that is, emotionless speech. Therefore, the initial speech does not contain any emotion and is emotionless speech. Emotion types include but are not limited to happiness, anger, indifference, anxiety, surprise, etc., and can be any emotion that changes the speech feature parameters. Among them, the speech feature parameters may include tone, intonation, and speaking speed. Tone refers to the attitude and emotional tendency expressed by the speaker through speech. Intonation refers to the rise and fall of the voice when speaking. Speaking speed refers to the speed of the language when speaking.
[0048] The text to be processed may correspond to one or more emotion types, which may be obtained by using one or more methods selected from a matching principle, user specification, and an emotion analysis model. User specification refers to one or more emotion types selected by the user from among several emotion type options.
[0049] A matching principle is adopted to obtain at least one emotion type corresponding to the text to be processed, including: matching at least one emotion type based on the content of the text to be processed based on the matching principle; wherein the matching principle refers to matching one or more emotion types based on emotion vocabulary, emotion phrases, semantic information, and contextual information in the text to be processed.
[0050] In an embodiment of the present application, the matching principles include an exact matching principle and a fuzzy matching principle, wherein the text to be processed is analyzed based on the exact matching principle to obtain the most matching required emotion type; and / or the text to be processed is analyzed based on the fuzzy matching principle to obtain at least two alternative emotion types with the highest degree of matching; the required emotion type and / or at least one alternative emotion type are determined as the final emotion type.
[0051] Specifically, the emotional words and / or emotional phrases in the text to be processed are analyzed based on an exact matching principle to obtain a single mandatory emotion type that best matches the text. The semantic information and / or contextual information in the text to be processed are analyzed based on a fuzzy matching principle to obtain at least two candidate emotion types with the highest degree of matching. The candidate emotion types are sent to the user so that the user can select at least one candidate emotion type. The mandatory emotion type and the at least one candidate emotion type selected by the user are determined as the final emotion type.
[0052] This involves using AI models based on matching principles to match at least one emotion type to the content of the text being processed. For example, GPT4 (Generative Pre-trained Transformer 4) or BERT (Bidirectional Encoder Representations from Transformers) both support precise matching of a single emotion type and are part of the sentiment analysis module in natural language processing. BERT also supports fuzzy matching of at least one emotion based on the text content, meaning that a single text can be mapped to multiple emotion types, or several highly matched emotion types can be calculated simultaneously as alternatives.
[0053] A sentiment analysis model is used to obtain at least one sentiment type corresponding to the text to be processed, including: in a conversation scenario, determining the content and sentiment type of the voice sent by the user, or the content and sentiment type of the text sent by the user; the text to be processed refers to the text used to reply to the user; the content of the voice or text sent by the user and the text to be processed are combined to form the current conversation content; historical conversation content whose similarity to the current conversation content is greater than a preset similarity; and determining a preset number of historical sentiment types that reply to the user in the historical conversation content as the sentiment type corresponding to the text to be processed.
[0054] S102. Acquire at least one corresponding adjustment parameter according to at least one emotion type; the adjustment parameter is used to adjust the speech feature parameter.
[0055] In an embodiment of the present application, according to at least one emotion type, the initial speech is input into at least one emotion unit corresponding to the emotion type to obtain at least one initial adjustment parameter corresponding to at least one adjustment parameter type; based on the fusion strategy, the initial adjustment parameters with the same adjustment parameter type are fused to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type.
[0056] Each emotion type corresponds to an emotion unit; the emotion unit corresponding to the emotion type is used to determine the adjustment parameter corresponding to at least one adjustment parameter type under that emotion type. Each emotion type corresponds to at least one adjustment parameter type; the adjustment parameter types corresponding to different emotion types can be the same or different. The adjustment parameter type refers to the type of adjustment parameter. Speech feature parameters refer to speech parameters that can express emotions.
[0057] For example, if the emotion types include happiness and surprise, the initial speech is input into the emotion unit corresponding to happiness to obtain an adjustment parameter of at least one adjustment parameter type under the emotion type of happiness; the initial speech is input into the emotion unit corresponding to surprise to obtain an adjustment parameter of at least one adjustment parameter type under the emotion type of surprise.
[0058] In an embodiment of the present application, two fusion strategies are provided: a weighted fusion method and an average fusion method. Initial adjustment parameters of the same adjustment parameter type are fused based on the fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type, including: initial adjustment parameters of the same adjustment parameter type are fused based on the weighted fusion method or the average fusion method to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type.
[0059] Furthermore, initial adjustment parameters of the same adjustment parameter type are fused based on a weighted fusion method to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type, including:
[0060] Step 1: Determine at least one initial fusion weight corresponding to each emotion type based on the matching degree between the text to be processed and each emotion type.
[0061] In the embodiments of the present application, the higher the degree of match with the text to be processed, the greater the fusion weight corresponding to the emotion type. The embodiments of the present application provide three methods for determining the initial fusion weight corresponding to the emotion type. The greater the fusion weight corresponding to the emotion type, the greater the adjustment effect of the initial adjustment parameters corresponding to the emotion type on the speech.
[0062] The first method of determining the initial fusion weight corresponding to the emotion type is: according to the preset weight setting rules, based on the degree of matching between the text to be processed and each emotion type, the initial fusion weight corresponding to each emotion type is determined; wherein, the preset weight setting rules explain how to determine the fusion weight corresponding to each emotion type based on the degree of matching between the text to be processed and each emotion type.
[0063] For example, the preset weight setting rule can be that the fusion weight corresponding to the emotion type with the highest degree of matching is 0.5, and the fusion weights corresponding to the emotion types with the second highest degree of matching to the fifth highest degree of matching are all 0.1. Assuming that the emotion type with the highest degree of matching is happiness, the initial fusion weight corresponding to the emotion type corresponding to happiness is 0.5, and the initial fusion weights corresponding to the other emotion types with the second highest degree of matching to the fifth highest degree of matching are 0.1.
[0064] The second method to determine the initial fusion weight corresponding to the emotion type is to add the matching degree of the text to be processed with all emotion types to obtain the sum value; and to obtain the initial fusion weight corresponding to each emotion type by taking the ratio of the matching degree of the text to be processed with each emotion type to the sum value.
[0065] For example, the emotion types include happiness, excitement, and surprise. The matching degree between the processed text and the emotion type of happiness is 0.8, the matching degree between the processed text and the emotion type of excitement is 0.7, and the matching degree between the processed text and the emotion type of surprise is 0.6. The matching degree between the processed text and all emotion types is summed up to get the sum value: sum = 0.8 + 0.7 + 0.6 = 2.1. The weights are normalized based on the sum value: the fusion weight for the emotion type of happiness is 0.8 / 2.1 = 8 / 21. The fusion weight for the emotion type of excitement is 0.7 / 2.1 = 1 / 3. The fusion weight for the emotion type of surprise is 0.6 / 2.1 = 2 / 7.
[0066] The third method of determining the fusion weight corresponding to the emotion type is: the preset fusion weight corresponding to each emotion type set by the user is determined as the initial fusion weight corresponding to each emotion type.
[0067] Step 2: According to at least one fusion weight corresponding to each emotion type, a target fusion weight corresponding to each emotion type is obtained.
[0068] In the implementation manner of the present application, all fusion weights corresponding to each emotion type are weighted summed or averaged to determine the target fusion weight corresponding to each emotion type.
[0069] Step 3: For any adjustment parameter type, based on the target fusion weight corresponding to each emotion type, perform weighted fusion on at least one initial adjustment parameter of the adjustment parameter type corresponding to each emotion type to obtain the target adjustment parameter value corresponding to the adjustment parameter type.
[0070] Furthermore, initial adjustment parameters of the same adjustment parameter type are fused based on an average fusion method to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type, including: for any adjustment parameter type, at least one initial adjustment parameter corresponding to the adjustment parameter type is averaged and fused to obtain a target adjustment parameter corresponding to the adjustment parameter type.
[0071] In the embodiment of the present application, if the user does not adopt the above three methods of determining the weights corresponding to the emotion types, the multiple emotion parameters are averaged and fused.
[0072] S103: Input the text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters, into a text-to-speech unit, transform the speech feature parameters of the intermediate speech or the speech feature parameters of the initial speech after the text to be processed is converted, and obtain the target speech.
[0073] In the embodiments of the present application, the text to be processed and the adjustment parameters are input into the text-to-speech unit, and the speech feature parameters of the intermediate speech after the text to be processed are transformed to obtain the target speech. Alternatively, the initial speech and the adjustment parameters are input into the text-to-speech unit, and the speech feature parameters of the initial speech are transformed to obtain the target speech.
[0074] Here, the text-to-speech unit includes at least one intermediate layer, and different intermediate layers transform the speech feature parameters of the intermediate speech after the text is converted or the speech feature parameters of the initial speech based on adjustment parameters of different adjustment parameter types; an intermediate layer transforms the speech feature parameters of the intermediate speech after the text is converted or the speech feature parameters of the initial speech based on adjustment parameters of one adjustment parameter type. Specifically, the speech feature parameters of the intermediate speech after the text is converted or the speech feature parameters of the initial speech are transformed, including: the i-th intermediate layer transforms the speech feature parameters of the intermediate speech after the text is converted or the speech feature parameters of the initial speech based on the corresponding adjustment parameters; i=i+1; jump to "the i-th intermediate layer transforms the speech feature parameters of the intermediate speech after the text is converted or the speech feature parameters of the initial speech based on the corresponding adjustment parameters" to continue execution; wherein the initial value of i is 1, and the maximum value of i is the number of intermediate layers.
[0075] Furthermore, the speech generation model in the embodiment of the present application includes at least two different emotion units and a text-to-speech unit, and the text-to-speech unit and all emotion units are connected respectively.
[0076] In this embodiment, the emotion unit transmits several adjustment parameters to the speech generation model, so that the emotion information is integrated with the final generated speech content, generating speech content that can reflect the emotion. The training process of each emotion unit is independent of each other, and the training process of the emotion unit is independent of the training process of the text-to-speech unit.
[0077] The text-to-speech unit can be a common TTS (text-to-speech) model, such as the Tacotron series, DeepSpeech series, FastSpeech series, TransformerTTS, VALL-E series, TTS models based on generative streams or adversarial networks, and other models that can convert text to speech. There are no restrictions on this.
[0078] Each emotion unit is composed of one or more multilayer perceptron (MLP) modules or convolutional neural network (CNN) modules. Preferably, the emotion unit is composed of one multilayer perceptron module or convolutional neural network.
[0079] The selected one or more emotion units will output different adaptive adjustment parameters to the text-to-speech unit to embed the emotion information into the speech, rather than inputting the same parameters to each layer in the text-to-speech unit.
[0080] The connection between the text-to-speech model and each emotion unit is achieved through a structure based on AdaLN (Adaptive Layer Normalization) or LoRA (Low-Rank Adaptation).
[0081] The emotion unit transmits several adjustment parameters to the text-to-speech unit, integrating the emotion information with the resulting speech content. Specifically, each emotion unit outputs multiple adaptive adjustment parameters corresponding to the corresponding emotion type. These fine-tuning parameters are then injected into different intermediate layers within the text-to-speech unit. The results generated by each layer in the text-to-speech unit are then linearly transformed to adjust parameters such as intonation, speech rate, and tone of the generated speech. Note that this step does not alter the text content, only the speech parameters.
[0082] Furthermore, the text-to-speech unit is trained through the following steps: obtaining at least one text training sample and a converted voice label corresponding to each text training sample, wherein the several text training samples refer to texts that need to be converted into voice; the converted voice label refers to the voice obtained by converting the text training samples through the trained text-to-speech model; the text-to-speech unit is trained based on the text training samples and the converted voice label corresponding to each text training sample; and the verification parameters of the target emotion unit are evaluated based on at least the test text sample. If the evaluation result of the target emotion unit meets the preset verification conditions, the training of the target emotion unit is completed.
[0083] Specifically, the verification parameter of the target emotion unit is determined based on the test text sample, or the set of the test text sample and the verification text sample. If the evaluation result of the target emotion unit meets the preset verification condition, the training of the target emotion unit is completed.
[0084] Here, when training the text-to-speech unit, the text-to-speech unit is isolated from all emotion units, and only the text-to-speech unit is trained, and any parameters of the emotion unit are not changed due to the training of the text-to-speech unit.
[0085] Further, refer to Figure 2 As shown, it is a training flow chart of the emotion unit provided in an embodiment of the present application.
[0086] S201: Obtain at least one corresponding speech training sample.
[0087] The speech training samples include at least a first speech sample and a second speech sample. The first speech sample is neutral speech obtained through processing with a text-to-speech model, and the second speech sample is speech corresponding to the target emotion type. The text content of the second speech sample is the same as that of the first speech sample. Neutral speech refers to speech without any emotion. The target emotion type refers to the emotion type corresponding to the target emotion unit to be trained.
[0088] In addition, the speech training sample may also include a third speech sample, the text content of the third speech sample is the same as the text content of the first speech sample and the text content of the second speech sample, and the similarity between the emotion type of the third speech sample and the emotion type of the second speech sample is greater than the similarity threshold.
[0089] In the embodiment of the present application, the similarity between two emotion types refers to the similarity between the speech features of two voices with the same text content of the two emotion types.
[0090] Here, since human emotions are continuous and mixed, for example, the speech feature parameters of speech with the emotional types of "happy" and "excited" partially overlap, it is necessary to introduce some similar speech training samples here as negative training samples to strengthen the understanding of "happy" by the emotional unit corresponding to "happy".
[0091] S202: Train the target emotion unit using at least one speech training sample, and evaluate the verification parameters of the target emotion unit based on at least a test speech sample. If the evaluation result of the target emotion unit meets a preset verification condition, the training of the target emotion unit is completed.
[0092] In an embodiment of the present application, a first speech sample is input into a target emotion unit to obtain at least one adjustment parameter corresponding to the target emotion type; at least one adjustment parameter corresponding to the target emotion type is input into a trained text-to-speech unit to transform the first speech sample to obtain a generated speech; the target emotion unit is updated according to the generated speech and the second speech sample so that the generated speech can approach the second speech sample; or the target emotion unit is updated according to the generated speech, the second speech sample and the third speech sample so that the generated speech can approach the second speech sample and move away from the third speech sample; and the verification parameters of the target emotion unit are evaluated based on at least the test speech sample. If the evaluation result of the target emotion unit meets the preset verification conditions, the training of the target emotion unit is completed.
[0093] Specifically, the verification parameter of the target emotion unit is determined based on the test speech sample, or the set of the test speech sample and the verification speech sample. If the evaluation result of the target emotion unit meets the preset verification condition, the training of the target emotion unit is completed.
[0094] The verification parameters include at least one of consistency verification, confusion rate verification, and error rate verification.
[0095] For example, to train an emotion unit corresponding to a happy emotion type, the happy emotion unit is trained based on a first speech sample, a second speech sample of a happy emotion type, and a third speech sample, so that it can output adjustment parameters that make the emotion of the speech "speaking happily".
[0096] Here, we obtain several trained emotion units to ensure that each unit can accurately convert neutral speech into speech with the corresponding emotion. Training the text-to-speech units and emotion units separately, as well as each emotion unit, also simplifies the training process and reduces training complexity.
[0097] In addition, if the emotion types covered by the trained emotion units are insufficient and additional emotion units are needed, the text-to-speech units and other emotion units can be isolated, and the speech training samples corresponding to the added emotion types can be used to train the newly added emotion units according to steps S201~S202.
[0098] An embodiment of the present application provides a speech generation method and apparatus, comprising: obtaining a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; obtaining at least one adjustment parameter corresponding to the at least one emotion type; inputting the text to be processed and the adjustment parameter, or the initial speech and the adjustment parameter, into a text-to-speech unit, and transforming the speech feature parameters of the intermediate speech after the text to be processed or the speech feature parameters of the initial speech to obtain a target speech. The present application is capable of generating speech with a variety of mixed emotions. When speech with non-predetermined emotions needs to be generated, the speech generation model structure does not need to be changed, and the generated emotion of the speech generation model can be quickly adjusted.
[0099] Based on the same inventive concept, a speech generation device corresponding to the speech generation method is also provided in the embodiment of the present application. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the above-mentioned speech generation method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0100] Reference Figure 3 FIG. 1 is a schematic diagram of a speech generating device provided in an embodiment of the present application, wherein the speech generating device includes:
[0101] An acquisition module 301 is configured to acquire a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type;
[0102] An input module 302 is configured to obtain at least one corresponding adjustment parameter according to at least one emotion type; the adjustment parameter is used to adjust the speech feature parameter;
[0103] The input module 302 is further used to input the text to be processed and the adjustment parameters, or the initial speech and the adjustment parameters, into a text-to-speech unit, and transform the speech feature parameters of the intermediate speech after the text to be processed is converted or the speech feature parameters of the initial speech to obtain the target speech. In one possible embodiment, the acquisition module 301 is specifically used to accurately match and / or fuzzily match at least one emotion type of the text to be processed; wherein the accurate matching is based on matching emotion types based on emotion words or emotion phrases in the text to be processed; and the fuzzy matching is based on matching emotion types based on semantic information and / or contextual information of the text to be processed.
[0104] An embodiment of the present application provides a speech generation device that can generate speech with a variety of mixed emotions. When it is necessary to generate speech with non-predetermined emotions, the speech generation model structure does not need to be changed, and the generated emotions of the speech generation model can be quickly adjusted.
[0105] like Figure 4 As shown, an electronic device 400 provided in an embodiment of the present application includes: a processor 401, a memory 402 and a bus, wherein the memory 402 stores machine-readable instructions executable by the processor 401. When the electronic device is running, the processor 401 communicates with the memory 402 through the bus, and the processor 401 executes the machine-readable instructions to perform the steps of the above-mentioned speech generation method.
[0106] Specifically, the above-mentioned memory 402 and processor 401 can be general-purpose memories and processors, which are not specifically limited here. When the processor 401 runs the computer program stored in the memory 402, the above-mentioned speech generation method can be executed.
[0107] Corresponding to the above-mentioned speech generation method, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned speech generation method are executed.
[0108] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the method embodiment, and will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0109] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0110] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0111] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the information processing method described in each embodiment of this application. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks, or optical disks.
[0112] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A speech generation method, characterized in that: The method comprises: Obtaining a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; Acquire at least one corresponding target adjustment parameter according to at least one emotion type; the target adjustment parameter is used to adjust the speech feature parameter; Inputting the to-be-processed text and the target adjustment parameters, or the initial speech and the target adjustment parameters, into a text-to-speech unit, and transforming speech feature parameters of the intermediate speech converted from the to-be-processed text or the initial speech to obtain a target speech; Among them, obtaining at least one corresponding target adjustment parameter according to at least one emotion type includes: according to multiple emotion types, inputting the initial speech into at least one emotion unit corresponding to each emotion type, and obtaining an initial adjustment parameter corresponding to at least one adjustment parameter type under each emotion type; fusing the initial adjustment parameters with the same adjustment parameter type based on a fusion strategy, and obtaining at least one target adjustment parameter corresponding to at least one adjustment parameter type; each emotion type corresponds to at least one adjustment parameter type.
2. The speech generation method according to claim 1, wherein: Obtain at least one emotion type corresponding to the text to be processed by the following steps: Based on a matching principle, matching at least one emotion type according to the content of the text to be processed; The matching principle refers to matching one or more emotion types based on emotion words, emotion phrases, semantic information, and context information in the text to be processed.
3. The speech generation method according to claim 2, wherein: The matching principle includes an exact matching principle and a fuzzy matching principle. The matching principle is based on the content of the text to be processed, and at least one emotion type is matched, including: Analyze the text to be processed based on the exact matching principle to obtain a most matching mandatory emotion type; and / or, analyzing the text to be processed based on the fuzzy matching principle to obtain at least two candidate emotion types with the highest matching degree; The mandatory emotion type and / or at least one candidate emotion type are determined as final emotion types.
4. The speech generation method according to claim 1, wherein: The fusing of initial adjustment parameters of the same adjustment parameter type based on the fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type includes: Determining at least one initial fusion weight corresponding to each emotion type according to the matching degree between the text to be processed and each emotion type; According to at least one fusion weight corresponding to each emotion type, a target fusion weight corresponding to each emotion type is obtained; For any adjustment parameter type, based on the target fusion weight corresponding to each emotion type, at least one initial adjustment parameter of the adjustment parameter type corresponding to each emotion type is weightedly fused to obtain the target adjustment parameter corresponding to the adjustment parameter type.
5. The speech generation method according to claim 1, wherein: The step of fusing initial adjustment parameters of the same adjustment parameter type based on a fusion strategy to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type further includes: For any adjustment parameter type, at least one initial adjustment parameter corresponding to the adjustment parameter type is averaged and fused to obtain a target adjustment parameter corresponding to the adjustment parameter type.
6. The speech generation method according to claim 4 or 5, characterized in that: The speech generation model includes at least two different emotion units and the text-to-speech unit. The target emotion unit corresponding to the target emotion type in the speech generation model is trained by the following steps: Obtaining at least one corresponding speech training sample, the speech training sample including at least a first speech sample and a second speech sample, the first speech sample being a neutral speech obtained by processing a text-to-speech model, the second speech sample being a speech corresponding to the target emotion type, and the text content of the second speech sample being the same as the text content of the first speech sample; The target emotion unit is trained using the at least one speech training sample, and then a verification parameter of the target emotion unit is evaluated based on at least a test speech sample. If the evaluation result of the target emotion unit meets a preset verification condition, the training of the target emotion unit is completed.
7. The speech generation method according to claim 6, characterized in that The training process of each emotion unit is independent of each other, and the training process of the emotion unit is independent of the training process of the text-to-speech unit.
8. The speech generation method according to claim 6, wherein: The speech training sample also includes a third speech sample, the text content of the third speech sample is the same as the text content of the first speech sample and the text content of the second speech sample, and the similarity between the emotion type of the third speech sample and the emotion type of the second speech sample is greater than a similarity threshold.
9. A speech generating device, characterized in that: The device comprises: An acquisition module, configured to acquire a text to be processed, an initial speech corresponding to the text to be processed, and at least one emotion type; An input module, configured to obtain at least one corresponding target adjustment parameter according to at least one emotion type; the target adjustment parameter is used to adjust the speech feature parameter; The input module is further configured to input the text to be processed and the target adjustment parameters, or the initial speech and the target adjustment parameters, into a text-to-speech unit, and transform the speech feature parameters of the intermediate speech converted from the text to be processed or the speech feature parameters of the initial speech to obtain a target speech; The input module is used to obtain at least one corresponding target adjustment parameter according to the following steps: according to multiple emotion types, the initial speech is input into at least one emotion unit corresponding to each emotion type to obtain an initial adjustment parameter corresponding to at least one adjustment parameter type under each emotion type; based on a fusion strategy, the initial adjustment parameters with the same adjustment parameter type are fused to obtain at least one target adjustment parameter corresponding to at least one adjustment parameter type; each emotion type corresponds to at least one adjustment parameter type.
Citation Information
Patent Citations
Speech synthesis method and system and terminal equipment
CN108615524A