Interactive speech synthesis method and system based on emotion regulation and control, medium and product
By obtaining the speaker's audio in real time and controlling the dialogue text with a large language model, personalized voice is generated, and the problems of single and complexity in the existing technology are solved, real-time voice generation of multiple roles and multiple emotions is realized, and the naturalness and real-time nature of human-computer interaction is improved.
Patent Information
- Application Number
- CN202510493069.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing text-to-speech conversion technology lacks real-time emotion recognition and regulation mechanisms. The generated voice is relatively single in emotion, style and personalized expression. The traditional methods are complex and rely on large-scale high-quality labeled data, which is easy to introduce errors.
By obtaining the speaker's role audio in real time, performing voice text transcription and audio emotional intent recognition, combining large language models for emotion analysis and text generation, regulating dialogue text to generate target speech, integrating speaker's role features, and using the VoiceChat model to map directly from characters to the Mel spectrum, outputting personalized speech.
Real-time voice generation with multiple roles and multiple emotions is realized, improving the naturalness and real-time nature of human-computer interaction, without the need for fine-tuning of the specific speaker model, and the voice matches the current emotional state, enriching emotional expression and personalized characteristics.
Smart Images

Figure CN120260537A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech processing, and particularly to an interactive speech synthesis method, system, medium and product based on emotion regulation. Background Art
[0002] Currently, text-to-speech conversion technology has been widely applied. Its core process usually includes steps such as text preprocessing, phoneme conversion, prosody annotation, feature extraction and vocoder generation. However, most current methods lack real-time emotion recognition and regulation mechanisms, and the generated speech is relatively single in terms of emotion, style and personalized expression. In addition, traditional methods have complex preprocessing and rely on data. The phoneme conversion and prosody annotation processes not only increase the algorithm complexity, but also have a high dependence on large-scale high-quality annotated data, and are prone to introducing errors. Summary of the Invention
[0003] To solve the technical problems in the background art, the present invention proposes an interactive speech synthesis method, system, medium and product based on emotion regulation.
[0004] In a first aspect, an interactive speech synthesis method based on emotion regulation proposed by the present invention includes:
[0005] Obtaining the speaker role audio in real time;
[0006] Performing speech-to-text transcription on the speaker role audio to obtain a first text;
[0007] Performing audio emotion intention recognition on the speaker role audio to obtain a second text;
[0008] Concatenating the first text and the second text to obtain a dialogue text;
[0009] Performing emotion regulation on the dialogue text to obtain an emotion-regulated dialogue text;
[0010] Selecting a speaker role;
[0011] Outputting a target speech according to the emotion-regulated dialogue text and the speaker role.
[0012] Preferably, performing emotion regulation on the dialogue text to obtain an emotion-regulated dialogue text specifically includes:
[0013] Using a preset large language model, combined with historical dialogue records, performing emotion analysis and text generation processing on the dialogue text to obtain an emotion-regulated dialogue text;
[0014] Among them, the expression of the emotion-regulation dialogue text is t7 = LLM(History(t6)); in the formula, t6 represents the dialogue text, LLM is a large language model, History is the operation of obtaining the dialogue text t6 wrapped into the historical dialogue record, and t7 is the emotion-regulation dialogue text after emotion regulation.
[0015] Preferably, before selecting the speaker role, it further includes:
[0016] Adding the speaker role;
[0017] Obtaining the speaker video of the speaker role;
[0018] Performing feature extraction on the speaker video to obtain the speaker role features;
[0019] Combining the speaker role features as the overall speaker features.
[0020] Preferably, after combining the speaker role features as the overall speaker features, it further includes:
[0021] Saving the speaker role and the corresponding overall speaker features.
[0022] Preferably, the speaker role features include: speaker ID, speaker face feature vector, speaker age information, audio data content of the speaker video material, description of the person's style, and text content of the speaker video material.
[0023] Preferably, performing feature extraction on the speaker video to obtain the speaker role features specifically includes:
[0024] Automatically sorting the speaker videos according to the input order of the speaker videos, and using the order of the speaker videos as the speaker ID;
[0025] Using a preset face recognition model to perform face feature extraction on the speaker video to obtain the speaker face feature vector;
[0026] Using a preset age prediction model to perform age prediction on the speaker video to obtain the speaker age information;
[0027] Using an audio-video separation tool to perform audio separation on the speaker video to obtain the audio data content of the video material;
[0028] Using a preset graphic-text dialogue large model to perform person style extraction on the speaker video to obtain the description of the person's style;
[0029] Using a preset text transcription model to perform speech-to-text transcription on the speaker video to obtain the text content of the speaker video material.
[0030] Preferably, the expression of the dialogue text is
[0031] t6 = Concat(ASR(input audio ), ALM(input audio ));
[0032] Wherein, ASR is a text transcription model, ALM is a large audio-visual dialogue model, input audio is the speaker role audio, t6 is the dialogue text, and Concat means concatenation.
[0033] Preferably, according to the emotional regulation of the dialogue text and the speaker role, the target voice is output, specifically including:
[0034] According to the speaker role, the overall characteristics of the speaker are obtained;
[0035] The overall characteristics of the speaker, the speaker ID t1 and the emotionally regulated dialogue text are formed into input features;
[0036] The input features are input into a preset VoiceChat model to obtain the target voice;
[0037] The target voice is output in a streaming manner.
[0038] In a second aspect, the present invention also provides an interactive voice synthesis system based on emotional regulation, including:
[0039] A speaker role module for selecting a speaker role;
[0040] A dialogue text acquisition module for real-time acquisition of the speaker role audio; performing speech-to-text transcription on the speaker role audio to obtain a first text; performing audio emotional intention recognition on the speaker role audio to obtain a second text; concatenating the first text and the second text to obtain a dialogue text;
[0041] An interactive emotion regulation module for emotionally regulating the dialogue text to obtain an emotionally regulated dialogue text;
[0042] A voice synthesis module for outputting a target voice according to the emotionally regulated dialogue text and the speaker role.
[0043] In a third aspect, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the interactive voice synthesis method based on emotional regulation described in any item of the first aspect are implemented.
[0044] In a fourth aspect, the present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the interactive voice synthesis method based on emotional regulation described in any item of the first aspect are implemented.
[0045] In the present invention, the proposed interactive speech synthesis method, system, medium and product based on emotion regulation can capture the emotional information input by the user in real time by first selecting a speaker role and then obtaining the audio of the speaker role in real time, so as to perform real-time dynamic emotion adjustment on the dialogue text according to the emotional information input by the user, and adjust parameters such as volume, speech rate, and intonation of the target speech in the speech generation process in real time, so that the output target speech always matches the current emotional state, greatly improving the naturalness and real-time performance of human-computer interaction, effectively making up for the deficiencies of the existing technologies in aspects such as emotional expression and personalization, and realizing the interactive speech generation ability of multiple roles and multiple emotions without the need for specific speaker model fine-tuning. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a schematic flowchart of an interactive speech synthesis method based on emotion regulation in an embodiment proposed by the present invention.
[0047] Figure 2 It is a schematic diagram for extracting the overall characteristics of a speaker in an embodiment proposed by the present invention.
[0048] Figure 3 It is a schematic diagram for emotion regulation of a dialogue text in an embodiment proposed by the present invention.
[0049] Figure 4 It is a schematic diagram of the large language model structure in an embodiment proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.
[0051] Referring to Figure 1 , an interactive speech synthesis method based on emotion regulation proposed by the present invention includes:
[0052] Select a speaker role;
[0053] Obtain a dialogue text;
[0054] Perform emotion regulation on the dialogue text to obtain an emotion-regulated dialogue text;
[0055] Output a target speech according to the emotion-regulated dialogue text and the speaker role.
[0056] The present invention first selects a speaker role, then obtains a dialogue text, and performs emotional regulation on the dialogue text to adjust characteristics such as the volume, speech rate, and intonation of the target speech output, effectively making up for the deficiencies of the prior art in aspects such as emotional expression and personalization, and can achieve the ability of multi-role and multi-emotion interactive speech generation without performing fine-tuning of the model for a specific speaker.
[0057] Among them, the dialogue text t6 obtained in this embodiment includes two parts: the first part is the first text generated by calling a microphone to record and then using a text transcription model for the speaker role audio; the second part is a descriptive text obtained by performing audio emotional intention recognition on the speaker role audio through a speech and text dialogue large model, that is, the second text; finally, the first text and the second text are concatenated as the input dialogue text t6.
[0058] Among them, the expression of the dialogue text t6 is t6 = Concat(ASR(input audio ), ALM(input audio )); in the formula, ASR is the text transcription model, ALM is the speech and text dialogue large model, input audio is the speaker role audio; t6 is the finally determined dialogue text, and Concat represents concatenation.
[0059] Therefore, in this embodiment, obtaining the dialogue text specifically includes:
[0060] Obtaining the speaker role audio in real time;
[0061] Performing speech-to-text transcription on the speaker role audio to obtain the first text;
[0062] Performing audio emotional intention recognition on the speaker role audio to obtain the second text;
[0063] Concatenating the first text and the second text to obtain the dialogue text.
[0064] In this embodiment, by obtaining the speaker role audio in real time, the emotional information input by the user can be captured in real time, so as to perform real-time dynamic emotional adjustment on the dialogue text according to the emotional information input by the user, so as to real-time adjust parameters such as the volume, speech rate, and intonation of the target speech in the speech generation process, so that the output target speech always matches the current emotional state, greatly improving the naturalness and real-time nature of human-computer interaction.
[0065] Specifically, in this embodiment, the speaker role audio is obtained in real time by real-time recording through a microphone.
[0066] In this embodiment, performing emotional regulation on the dialogue text to obtain the emotionally regulated dialogue text t7 specifically includes:
[0067] Using a pre-set large language model (LLM), combined with historical conversation records, perform sentiment analysis and text generation processing on t6 to obtain an emotion-regulated conversation text t7, such that the emotion-regulated conversation text t7 embeds such as <break> 、 <long-break> 、 <breath> 、 <laugh>Special characters such as etc. are used to control the voice parameters (volume, speech rate, intonation) of the emotion-regulated dialogue text t7.
[0068] Among them, the expression of the emotion-regulated dialogue text t7 is t7 = LLM(History(t6)); in the formula, LLM is a large language model, History is the operation of obtaining the dialogue text t6 wrapped into the historical dialogue record, and t7 is the emotion-regulated dialogue text after emotion regulation.
[0069] The emotion-regulated dialogue text generated in this embodiment integrates the current dialogue text and the historically processed dialogue text, and can provide a more accurate and more in line with the current emotional state dialogue text.
[0070] Such as Figure 4 As shown, the large language model in this embodiment is based on the Transformer architecture, adopts a decoder-only structure, and realizes sequence modeling by stacking self-attention layers and feed-forward networks. The large language model in this embodiment also introduces LE grouped query attention (GQA) to optimize the efficiency of traditional multi-head attention through grouped calculation. Among them, the position encoding method of the large language model in this embodiment adopts a variant YaRN of traditional rotational position encoding.
[0071] The model in this embodiment supports a context length of 128K tokens, can generate 8K tokens of coherent text, and is suitable for long text tasks.
[0072] In one specific embodiment, before selecting the speaker role, it further includes:
[0073] Adding a speaker role;
[0074] Obtaining the speaker video of the speaker role;
[0075] Performing feature extraction on the speaker video to obtain speaker role features;
[0076] Combining the speaker role features as the overall speaker features.
[0077] Specifically, first, record the speaker video material through a camera to add a speaker role; secondly, extract the speaker role feature information, including the speaker ID, the speaker's face features, the speaker's age information, the speaker's style description, the text content of the speaker video material, and the audio data content of the speaker video material, and encode the above feature information as the overall features of the speaker, so as to be applicable to the speaker role in the subsequent emotion-regulated dialogue text to realize the personalized speech synthesis function.
[0078] In this embodiment, obtaining the speaker video specifically includes:
[0079] By calling the camera to shoot real-time speaker video materials, the speaker video is obtained.
[0080] In this embodiment, the speaker role characteristics include: speaker ID t1, speaker face feature vector f0, speaker age information t2, audio data content t3 of the speaker video material, character style description t4, and text content t5 of the speaker video material.
[0081] This embodiment is set up in this way, integrating the speaker image, text description, and voice data, realizing the extraction and integration of multi-dimensional features, so that the generated voice is richer and more natural in terms of emotional expression and personalized features, meeting diverse application requirements.
[0082] It should be understood that the speaker ID t1 in this embodiment is automatically sorted according to the input order, and the numerical type is an integer. Among them, the expression of t1 is t1 = n (n = 1, 2, 3,...); in the formula, n is a positive integer.
[0083] The age information t2 in this embodiment is obtained by a preset age prediction model, and the result is an integer value. Among them, the calculation formula of t2 is t2 = AgePredModel(image); in the formula, AgePredModel is the age prediction model, image is a certain frame of the picture containing the speaker, and t2 is the predicted age information.
[0084] The audio data content t3 of the speaker video material in this embodiment is separated from the speaker video, and the audio data is saved in tensor format. Among them, the calculation formula of t3 is t3 = ToTensor(Tools(video)); in the formula, Tools is an audio-video separation tool for obtaining pure audio information; ToTensor is an audio-tensor conversion function.
[0085] The character style description t4 in this embodiment is obtained by a preset text-image dialogue large model. By setting a prompt word template, the model describes the style of the speaker in the image, and the result is a string containing keywords of the style description. Among them, the calculation formula of t4 is t4 = VLM(image); in the formula, VLM is the text-image dialogue large model, image is a certain frame of the picture containing the speaker, and t4 is the keyword of the style description.
[0086] The text content t5 of the speaker video material in this embodiment is obtained by a preset text transcription model, and the result is a string containing the speech content of the speaker in the audio. Among them, the calculation formula of the text content t5 of the speaker video material is t5 = ASR(audio); in the formula, ASR is the text transcription model; audio is the audio containing the clear human voice of the speaker, separated from the video material; t5 is the speech content of the speaker in the audio.
[0087] Therefore, in this embodiment, feature extraction is performed on the speaker video to obtain the speaker role features, specifically including:
[0088] Automatically sort the speaker videos according to the input order of the speaker videos, and use the order of the speaker videos as the speaker ID t1;
[0089] Use a preset face recognition model to extract face features from the speaker video to obtain the speaker face feature vector f0;
[0090] Use a preset age prediction model to predict the age of the speaker video to obtain the speaker age information t2;
[0091] Use an audio-video separation tool to separate the audio from the speaker video to obtain the audio data content t3 of the video material;
[0092] Use a preset graphic-text dialogue large model to extract the character style of the speaker video to obtain the character style description t4;
[0093] Use a preset text transcription model to transcribe the speech of the speaker video into text to obtain the text content t5 of the speaker video material.
[0094] It should be noted that all kinds of models in this embodiment have been trained.
[0095] In this embodiment, the overall speaker feature F is the combination of the above features, separated by special characters, and the expression is: F = Concat(<speaker t1>, <face>,f0,<\face>, <age>,t2,<\age>, <audio>,t3,<\audio>,<style>,t4,<\style>,<text>,t5,<\text>,<\speaker t1>); where Concat is a concatenation operation.
[0096] like Figure 2 As shown, in one of the specific embodiments, a speaker material video is obtained; feature extraction is performed on the speaker material video to obtain speaker role features; wherein the speaker role features include speaker ID (001), speaker facial features, speaker age information (23 years old), speaker style description, speaker video material text content, and speaker video material audio data content, and the above feature information is encoded as the overall feature of the speaker.
[0097] Among them, the speaker's style description creates an image of a Q-version girl with an exaggerated head-to-body ratio. The details of her curly hair, sneakers, and shopping bags combine the Japanese Harajuku style with street trends, showing youthful vitality while implying symbols of consumer culture, and conveying a sense of healing through warm light and shadow and dynamic postures.
[0098] In a further embodiment, after combining the speaker role features as the speaker overall features, the method further includes:
[0099] The speaker roles and the corresponding overall speaker characteristics are saved.
[0100] Specifically, this embodiment saves the overall speaker feature F in a pickle file format on a local disk.
[0101] This embodiment is configured in this way so that when the added speaker role is used subsequently, the relevant feature information of the speaker role can be directly read and loaded from the local storage file.
[0102] Therefore, in another specific embodiment, selecting a speaker role specifically includes: selecting a speaker role from saved speaker roles according to the dialogue text.
[0103] In this embodiment, the target speech is output according to the emotion regulation dialogue text and the speaker role, specifically including:
[0104] According to the speaker role, the overall characteristics of the speaker are obtained;
[0105] The overall characteristics of the speaker, the speaker ID t1 and the emotion regulation dialogue text are used to form input features;
[0106] Input the input features into the preset VoiceChat model to obtain the target speech;
[0107] Output the target audio in streaming format.
[0108] In this embodiment, the overall characteristics of the selected speaker role and the emotion-regulated dialogue text are combined as the input features of the VoiceChat model. The target voice is obtained through the VoiceChat model and the target voice is streamed out, realizing the personalized speech synthesis function, and the interactive speech generation ability with multiple roles and multiple emotions can be realized without fine-tuning the model of a specific speaker.
[0109] It should be noted that the VoiceChat model in this embodiment is pre-trained on a large-scale multilingual dataset and follows the above input data format. Therefore, it has good generalization performance and can adapt to new samples without fine-tuning. Moreover, the VoiceChat model is based on a general large model architecture, does not rely on phonemes for text-to-speech conversion, directly realizes the mapping from characters to mel spectrograms, can process any language represented by text, and reduces the complexity of the method operation.
[0110] In one specific embodiment, when the role settings include Speaker A and Speaker B, the user selects the speaker role, such as Speaker A (the other party's role) and Speaker B (one's own role);
[0111] Next, as Figure 3 shown, the user records in real time through the microphone to obtain the real-time audio of Speaker B (one's own role); among them, Speaker A (the other party's role) uses the preset feature data of the system and does not need to obtain real-time audio;
[0112] The text transcription model is used to obtain the dialogue text according to the real-time audio of Speaker B;
[0113] The emotion of the dialogue text is regulated through the audio-visual dialogue large model to obtain the emotion-regulated dialogue text t7;
[0114] The overall characteristics F of Speaker A are loaded respectively A and the overall characteristics F of Speaker B B ;
[0115] The features of Speaker A, the features of Speaker B, the ID t1 of Speaker A and the emotion-regulated dialogue text t7 are concatenated to form the model input I according to the following formula: where, I = Concat(F A , F B , t1, t7); in the formula, Concat is the concatenation operation; among them, after encoding, each type of data is adjusted to 1024 dimensions;
[0116] After receiving the input feature I, the VoiceChat model directly maps from characters to mel spectrograms, generates the target voice without traditional phoneme conversion and prosody annotation;
[0117] Finally, the generated target voice is played in real time through streaming output, realizing a highly natural and personalized interaction between the user and the method.
[0118] Among them, the text of the user's real-time audio transcription in this embodiment is "I didn't do well in today's exam. I'm in a very bad mood. Is there any way to relieve it?".
[0119] The analysis result of the text by the audio-visual large model in this embodiment is "This audio expresses the loss and helplessness emotions caused by the failure in the exam, and at the same time reveals the desire to seek solutions. The tone and emotion may be low, helpless with a trace of eagerness.".
[0120] The input t6 of the model in this embodiment is "I didn't do well in today's exam. I'm in a very bad mood. Is there any way to relieve it? This audio expresses the loss and helplessness emotions caused by the failure in the exam, and at the same time reveals the desire to seek solutions. The tone and emotion may be low, helpless with a trace of eagerness.".
[0121] The text content t7 of the target voice in this embodiment is " <break>I fully understand how you feel at this moment. The sense of frustration brought about by failing the exam is like a heavy stone pressing on your chest, right? First, gently hug the self who has been working hard, and allow yourself to daze for half an hour, listen to a depressing song, or sigh towards the sky. <breath>..These emotions are all allowed <laugh>。Make a funny face to encourage yourself in front of the mirror, or run three laps on the playground to let the sweat wash away your irritability - when your body moves, your brain will secrete dopamine that makes you happy. <break> 。 <long-break>Remember? Exams are just a supply station in the long race of life, not the key line. Last year, a senior student scored 47 points in physics. He turned his wrong question book into a notebook and used comics to record each knowledge point. Later, he made a comeback and ranked among the top ten in the grade. You see, the confusion at this moment may be the beginning of transformation? Believe that there are always stars behind the dark clouds. "
[0122] In a second aspect, the present invention further proposes an interactive speech synthesis system based on emotion regulation, comprising:
[0123] Speaker role module, used to select speaker role;
[0124] The dialogue text acquisition module is used to acquire the speaker role audio in real time; perform voice-to-text transcription on the speaker role audio to obtain a first text; perform audio emotion intention recognition on the speaker role audio to obtain a second text; and concatenate the first text and the second text to obtain a dialogue text;
[0125] An interactive emotion regulation module is used to perform emotion regulation on the dialogue text to obtain an emotion regulation dialogue text;
[0126] The speech synthesis module is used to regulate the dialogue text and speaker role according to the emotion and output the target speech.
[0127] The dialogue text acquisition module in this embodiment is preset with a text transcription model and a large audio-text dialogue model. The text transcription model is used to transcribe the speaker character audio into speech and text to obtain a first text; the large audio-text dialogue model is used to perform audio emotion intention recognition on the speaker character audio to obtain a second text.
[0128] The expression of the dialogue text is t6 = Concat(ASR(input audio ),ALM(input audio ));
[0129] In the formula, ASR is the text transcription model, ALM is the audio-text dialogue model, and input audio is the speaker role audio, t6 is the dialogue text, and Concat means concatenation.
[0130] Among them, the expression of the emotion regulation dialogue text is t7=LLM(History(t6)); wherein t6 represents the dialogue text, LLM is the large language model, History is the operation of obtaining the dialogue text t6 and wrapping it into the historical dialogue record, and t7 is the emotion regulation dialogue text after emotion regulation.
[0131] In this embodiment, the speaker role module is further configured to add a speaker role; obtain the speaker video of the speaker role; extract features from the speaker video to obtain speaker role features; and combine the speaker role features as the overall speaker features.
[0132] In a further embodiment, a storage module is further included, and the storage module is configured to save the speaker role and the corresponding overall speaker features.
[0133] The speaker role features in this embodiment include: speaker ID, speaker face feature vector, speaker age information, audio data content of the speaker video material, description of the person's style, and text content of the speaker video material.
[0134] Among them, the process of obtaining the speaker role features is the same as the process of obtaining the speaker role features in the first aspect.
[0135] In a third aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the interactive voice synthesis method based on emotion regulation described in any one of the first aspects are implemented.
[0136] In a fourth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the interactive voice synthesis method based on emotion regulation described in any one of the first aspects are implemented.
[0137] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention. < / break> < / laugh> < / breath> < / break> < / audio> < / age> < / face> < / laugh> < / breath> < / long-break> < / break>
Claims
1. An interactive speech synthesis method based on emotion regulation, characterized in that, It includes: Select a speaker role; Obtain the audio of the speaker role in real time; Transcribe the audio of the speaker role into speech text to obtain the first text; Identify the audio emotional intention of the speaker role to obtain the second text; Concatenate the first text and the second text to obtain the dialogue text; Perform emotional regulation on the dialogue text to obtain the emotionally regulated dialogue text; Output the target speech according to the emotionally regulated dialogue text and the speaker role.
2. The interactive voice synthesis method based on emotion regulation according to claim 1, wherein Performing emotional regulation on the dialogue text to obtain the emotionally regulated dialogue text specifically includes: Using a preset large language model and combining historical dialogue records, perform emotional analysis and text generation processing on the dialogue text to obtain the emotionally regulated dialogue text; Among them, the expression of the emotionally regulated dialogue text is t7 = LLM(History(t6)); in the formula, t6 represents the dialogue text, LLM is the large language model, History is the operation of obtaining the dialogue text t6 wrapped into the historical dialogue record, and t7 is the emotionally regulated dialogue text after emotional regulation.
3. The interactive voice synthesis method based on emotion regulation according to claim 1, characterized in that, Before selecting the speaker role, it also includes: Add a speaker role; Obtain the speaker's video of the speaker role; Extract features from the speaker's video to obtain the speaker role features; Combine the speaker role features as the overall speaker features. Preferably, after combining the speaker role features as the overall speaker features, it also includes: Save the speaker role and the corresponding overall speaker features.
4. The interactive voice synthesis method based on emotion regulation according to claim 3, wherein The speaker role features include: speaker ID, speaker face feature vector, speaker age information, audio data content of the speaker's video material, description of the person's style, and text content of the speaker's video material.
5. The interactive voice synthesis method based on emotion regulation according to claim 4, wherein Extracting features from the speaker's video to obtain the speaker role features specifically includes: Automatically sort the speaker's videos according to the input order of the speaker's videos, and use the order of the speaker's videos as the speaker ID; Use a preset face recognition model to extract face features from the speaker's video to obtain the speaker's face feature vector; Use a preset age prediction model to predict the age of the speaker from the speaker's video to obtain the speaker's age information; Use an audio-video separation tool to separate the audio from the speaker's video to obtain the audio data content of the video material; Use a preset graphic-text dialogue large model to extract the person's style from the speaker's video to obtain the description of the person's style; Use a preset text transcription model to transcribe the speech of the speaker's video into text to obtain the text content of the speaker's video material.
6. The interactive voice synthesis method based on emotion regulation according to claim 1, characterized in that The expression of the dialogue text is t6 = Concat(ASR(input audio ), ALM(input audio )); Wherein, ASR is a text transcription model, ALM is a speech-to-text dialogue large model, and input audio is the speaker role audio, t6 is the dialogue text, and Concat represents concatenation.
7. The interactive voice synthesis method based on emotion regulation according to claim 1, wherein Outputting the target speech according to the emotionally regulated dialogue text and the speaker role specifically includes: Obtain the overall speaker features according to the speaker role; Form input features by combining the overall speaker features and the emotionally regulated dialogue text; Input the input features into a preset VoiceChat model to obtain the target speech; Output the target speech in a streaming manner.
8. The interactive voice synthesis system based on emotion regulation according to claim 1, characterized in that It includes: A speaker role module for selecting a speaker role; A dialogue text acquisition module for obtaining the audio of the speaker role in real time; Performing speech-to-text transcription on the audio of the speaker role to obtain a first text; performing audio emotion intention recognition on the audio of the speaker role to obtain a second text; splicing the first text and the second text to obtain a dialogue text; An interactive emotion regulation module for regulating the emotion of the dialogue text to obtain an emotion-regulated dialogue text; A speech synthesis module for outputting a target speech according to the emotion-regulated dialogue text and the speaker role.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the emotion-regulated interactive speech synthesis method described in any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the emotion-regulated interactive speech synthesis method described in any one of claims 1-7.
Citation Information
Patent Citations
Intelligent robot-oriented voice synthesis method and device
CN109461435A
Man-machine interaction method and man-machine interaction device
CN110110169A
Voice generation method and device based on big data, equipment and medium
CN111445906A
Speech synthesis model, model training method and speech synthesis method
CN113920977A
Face recognition method and device, electronic equipment and storage medium
CN116469146A