A speech generation method, apparatus, device, and storage medium
By directly embedding emotional information when generating prompts in smart voice devices, emotional recognition errors are avoided. An encoder-decoder model is used to generate response speech that matches the dialogue content, solving the problem of inaccurate emotional expression in speech synthesis and achieving more accurate emotional expression and a better human-computer interaction experience.
Patent Information
- Application Number
- CN202310418107.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-13
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2043-04-13
AI Technical Summary
Existing intelligent voice devices rely on the recognition of user voice emotions during the speech synthesis process, which leads to the gradual aggravation of errors and results in inaccurate emotional expression in the synthesized speech data.
By generating prompts based on the first dialogue content and the reply text, emotional information is directly embedded into the prompts to generate reply speech that matches the emotion of the first dialogue content, thus avoiding errors in the emotion recognition process. An encoder-decoder structure such as the Transformer model is used for speech generation.
It improves the accuracy of speech synthesis, and the generated speech data expresses emotions more accurately, enhancing the human-computer interaction experience and empathy.
Smart Images

Figure CN116798402B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech synthesis technology, and in particular to a speech generation method, apparatus, device and storage medium. Background Technology
[0002] With the rapid development of computer and internet technology, speech synthesis has been widely applied in interactive fields such as smart homes and intelligent robots. When using smart voice devices, users are no longer satisfied with a single, neutral emotional response from synthesized speech; they desire responses that correspond to the user's emotions. This would make human-computer interaction systems more vivid and relatable. For example, when a user is angry, the smart voice device could respond with comforting or apologetic emotions.
[0003] Currently, intelligent voice devices typically synthesize speech data by combining emotion tags with response text. Emotion tags are determined by recognizing the emotion in the user's voice. This emotion is then combined with the response text to generate the emotion in the response speech. Finally, the emotion in the response speech and the response text are used to synthesize speech with the desired emotion. It is evident that this speech synthesis relies heavily on recognizing the user's voice emotion. Therefore, any deviation in the recognition of the user's voice emotion leads to a cascading error during the speech synthesis process, resulting in inaccurate emotional expression in the synthesized speech data. Summary of the Invention
[0004] To address the aforementioned problems, this application proposes a speech generation method, apparatus, device, and storage medium that can significantly improve the accuracy of synthesized speech.
[0005] According to a first aspect of the embodiments of this application, a speech generation method is provided, comprising:
[0006] Determine the response text corresponding to the first dialogue content; wherein, the first dialogue content includes dialogue voice and / or voice text corresponding to the dialogue voice;
[0007] Based on the first dialogue content and the reply text, a prompt message is generated, the prompt message including the emotional information of the first dialogue content;
[0008] Based on the prompt information, a response voice corresponding to the first dialogue content is generated, and the emotion of the response voice matches the emotion of the first dialogue content according to a preset emotion matching relationship.
[0009] According to a second aspect of the embodiments of this application, a speech generation apparatus is provided, comprising:
[0010] The determining module is used to determine the response text corresponding to the first dialogue content; wherein, the first dialogue content includes dialogue voice and / or voice text corresponding to the dialogue voice;
[0011] A generation module is used to generate prompt information based on the first dialogue content and the reply text, wherein the prompt information includes the emotional information of the first dialogue content;
[0012] The synthesis module is used to generate a response voice corresponding to the first dialogue content based on the prompt information, wherein the emotion of the response voice conforms to a preset emotion matching relationship with the emotion of the first dialogue content.
[0013] A third aspect of this application provides an electronic device, comprising:
[0014] Memory and processor;
[0015] The memory is connected to the processor and is used to store programs;
[0016] The processor implements the above-described speech generation method by running the program in the memory.
[0017] A fourth aspect of this application provides a storage medium storing a computer program, which, when executed by a processor, implements the above-described speech generation method.
[0018] One embodiment of the above application has the following advantages or beneficial effects:
[0019] Based on the first dialogue content and the corresponding reply text, a prompt message is generated. This prompt message contains the emotional information of the first dialogue content, which is directly carried into the prompt message through the original content of the first dialogue. Then, the reply speech corresponding to the first dialogue content is directly generated using the prompt message. In the above process, the emotional information carried by the first dialogue content itself can be directly injected into the prompt message. Then, a reply speech matching the emotion of the first dialogue content is generated based on the prompt message. This process does not require the recognition of the emotional label of the speech data, avoiding the gradual aggravation of errors during speech synthesis, and making the emotional expression of the synthesized speech data more accurate. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0021] Figure 1 A schematic flowchart illustrating a speech generation method provided in an embodiment of this application;
[0022] Figure 2 A schematic diagram of speech synthesis in the prior art provided in the embodiments of this application;
[0023] Figure 3 A schematic diagram of step S130 of a speech generation method provided in an embodiment of this application;
[0024] Figure 4 This is a schematic diagram illustrating a specific process of a speech generation method provided in an embodiment of this application;
[0025] Figure 5 This is a schematic diagram of the structure of a speech generation device provided in an embodiment of this application;
[0026] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0027] The technical solutions of this application are applicable to various speech generation scenarios, such as human-computer interaction scenarios and conference scenarios. Using the technical solutions of this application can improve the accuracy of speech synthesis.
[0028] The technical solutions of this application can be applied, by way of example, to hardware devices such as processors, electronic devices, and servers (including cloud servers), or packaged as software programs and run. When the hardware device executes the processing procedure of the technical solutions of this application, or when the aforementioned software program is run, the purpose of generating a reply voice corresponding to the first dialogue content based on the prompt information can be achieved. This application only provides an exemplary description of the specific processing procedure of the technical solutions of this application, and does not limit the specific implementation form of the technical solutions of this application. Any technical implementation form that can execute the processing procedure of the technical solutions of this application can be adopted by this application.
[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0030] Exemplary methods
[0031] Figure 1 This is a flowchart of a speech generation method according to an embodiment of this application. In an exemplary embodiment, a speech generation method is provided, including:
[0032] S110. Determine the reply text corresponding to the content of the first dialogue;
[0033] S120. Generate a prompt message based on the first dialogue content and the reply text;
[0034] S130. Based on the prompt information, generate a response voice corresponding to the first dialogue content, wherein the emotion of the response voice conforms to a preset emotion matching relationship with the emotion of the first dialogue content.
[0035] In step S110, exemplarily, the first dialogue content refers to the user-inputted voice or text, which can be in Chinese, English, Spanish, etc. The first dialogue content includes: dialogue voice or corresponding voice-text, or dialogue voice and its corresponding voice-text. Optionally, the dialogue voice can be obtained through a device with sound recording capabilities (e.g., a microphone). The voice-text can be text data input by the user, or it can be obtained by recognizing dialogue voice using a speech recognition model in the prior art. The response text refers to the answer to the first dialogue content.
[0036] Specifically, the correspondence between dialogue content and response text can be pre-stored. This way, after obtaining the first dialogue content, the corresponding response text can be determined based on the correspondence. Alternatively, a neural network model can be trained using sample dialogue content training data and the corresponding sample response texts to obtain a trained model. The training data can be speech data obtained from an open-source speech database. Inputting the obtained first dialogue content into the trained model will then output the response text.
[0037] In step S120, for example, the prompt information includes emotional information of the first dialogue content. For instance, when the first dialogue content includes spoken dialogue, the tone of voice in the spoken dialogue can reflect emotional information; therefore, the prompt information generated based on the first dialogue content and the reply text includes the emotional information of the first dialogue content. As another example, when the first dialogue content includes spoken text, the emotional information corresponding to the spoken text can also be analyzed. Therefore, the prompt information generated based on the first dialogue content and the reply text includes the emotional information of the first dialogue content. Optionally, the prompt information can be generated by concatenating the first dialogue content and the reply text; alternatively, the prompt information can be generated by merging the first dialogue content and the reply text.
[0038] It is understood that, in this embodiment of the application, generating prompt information through the first dialogue content and the corresponding reply text allows for the embedding of emotional components contained in the first dialogue content into the generated prompt information using the original information of the first dialogue content. In this process, the emotional information of the first dialogue content does not undergo secondary recognition and processing; therefore, the generated prompt information can retain the emotional information from the first dialogue content to the greatest extent possible, avoiding the problem of inaccurate emotional information extraction caused by extracting emotional information from the first dialogue content.
[0039] In step S130, for example, the preset emotion matching relationship can be achieved by pre-designing some emotion response mechanisms, that is, setting the corresponding emotion of the reply according to the different emotions expressed by the user. For example, if the emotion of the first dialogue content output by the user is anger, then the reply voice should be in a comforting tone; if the emotion of the first dialogue content output by the user is anxiety or fear, then the reply voice should use an encouraging or motivating tone, etc.
[0040] Specifically, sample prompts can be generated in advance based on sample dialogue content and sample response text, thus including emotional information in the prompts. The neural network model is then trained using the sample prompts and corresponding sample response speech to obtain a trained speech generation model. Finally, inputting the prompts into the trained speech generation model will produce the corresponding response speech.
[0041] In existing technologies, such as Figure 2 As shown, the speech synthesis process is as follows: First, the user's speech is converted into text content, and the corresponding emotion in the speech is identified as an emotion tag (i.e., the user's speech emotion). The text content and emotion tag are input into the dialogue management response generation module to obtain the response text and emotion tag. Then, the response text and emotion tag are input into the speech synthesis module, where the emotion tag is encoded by the emotion encoding module to obtain emotion encoding, and the response text is encoded by the synthesis encoding module to form text encoding. Both are then simultaneously sent to the synthesis decoding module to form spectral features with emotion. Finally, the vocoder decodes the spectral features with emotion to obtain speech. The above speech synthesis relies on the emotion tag identified in the user's speech emotion. That is, if the initial speech emotion recognition is incorrect, subsequent response text and speech synthesis will use incorrect input, inevitably leading to incorrect output. Thus, the error gradually increases during the speech synthesis process, resulting in inaccurate emotional expression in the synthesized speech data.
[0042] In the technical solution of this application, prompt information is generated based on the first dialogue content and the corresponding reply text. The prompt information contains the emotional information of the first dialogue content. Furthermore, the emotional information of the first dialogue content is directly carried into the prompt information through the original content of the first dialogue content. Then, the reply voice corresponding to the first dialogue content is directly generated using the prompt information. In the above processing, the emotional information carried by the first dialogue content itself can be directly injected into the prompt information. Then, a reply voice matching the emotion of the first dialogue content is generated based on the prompt information. This processing does not require the recognition of the emotional label of the voice data, avoiding the gradual aggravation of errors during the voice synthesis process, and making the emotional expression of the synthesized voice data more accurate.
[0043] In one implementation, a prompt message is generated based on the first dialogue content and the reply text. Step S120 includes:
[0044] The first dialogue content and the reply text are encoded respectively to obtain a first code corresponding to the first dialogue content and a second code corresponding to the reply text. The first code contains the sentiment encoding information of the first dialogue content.
[0045] The first code and the second code are fused together to obtain the prompt information.
[0046] In existing technologies, the emotion tags generated through emotion recognition and dialogue management are ultimately fed into the speech synthesis system in the form of discrete tags for speech synthesis. This results in a relatively fixed emotion in the synthesized speech, while human emotions are quite complex. The fixed form of discrete tags reduces the diversity of emotional expression.
[0047] Preferably, the spoken dialogue and its corresponding text in the first dialogue content are encoded separately to obtain spoken dialogue encoding and text-to-speech encoding. The response text is then encoded to obtain a second encoding. The spoken dialogue encoding, text-to-speech encoding, and second encoding are concatenated to obtain the prompt information. This approach uses both spoken and text-based prompts, rather than a tagging-based approach, resulting in more nuanced and diverse emotional responses from the synthesized voice, making it more relatable and empathetic to the user.
[0048] Optionally, the dialogue speech is encoded to obtain a first code, and the reply text is encoded to obtain a second code. Since the dialogue speech contains the user's tone (i.e., emotional information), the first code contains emotional encoding information. The first code and the second code are concatenated to obtain the prompt information. The above encoding method can employ, but is not limited to, various deep neural networks.
[0049] Optionally, the speech text corresponding to the dialogue is encoded to obtain a first code, and the response text is encoded to obtain a second code. Since the user's emotional information can be analyzed from the speech text, the first code includes emotional encoding information. The first code and the second code are concatenated to obtain the prompt information. In this way, without extracting emotional tags, the dialogue speech or speech text is directly encoded, ensuring the accuracy of the emotion in the speech. As a result, the prompt information can contain accurate emotion, which helps to generate an accurate emotional response speech based on the prompt information.
[0050] In one implementation, generating prompt information based on the first dialogue content and the reply text, and generating a reply voice corresponding to the first dialogue content based on the prompt information, includes:
[0051] The first dialogue content and the reply text are input into a pre-trained speech generation model to obtain the reply speech corresponding to the first dialogue content.
[0052] The speech generation model generates prompt information based on the first dialogue content and the reply text, and generates a reply speech corresponding to the first dialogue content based on the prompt information.
[0053] For example, the speech generation model can be formed by a single model or by combining two models. For instance, an encoder-decoder structure can be used to encode the first dialogue content and the response text to generate prompt information, and then decode the prompt information to obtain the response speech. Alternatively, a first model can be used to fuse the first dialogue content and the response text to obtain prompt information, and then a second model can be used to decode the prompt information to obtain the response speech.
[0054] Optionally, the dialogue speech and response text can be input into a pre-trained speech generation model to obtain the corresponding response speech; in this case, the training data for the pre-trained speech generation model uses sample dialogue speech and sample response text. The sample response text is determined based on the sample dialogue speech.
[0055] Alternatively, the dialogue speech, its corresponding speech text, and the response text can be input into a pre-trained speech generation model to obtain the corresponding response speech; in this case, the training data for the pre-trained speech generation model uses sample dialogue speech, sample speech text, and sample response text. The sample speech text is obtained by translating the sample dialogue speech.
[0056] Alternatively, the speech text and response text can be input into a pre-trained speech generation model to obtain the corresponding response speech; in this case, the training data for the pre-trained speech generation model uses sample speech text and sample response text. Thus, the sample data contains complete emotional information, enabling the model to perceive the correct emotions during training, thereby facilitating the generation of response speech corresponding to the aforementioned emotional information.
[0057] Furthermore, the speech generation model is trained through a first training process;
[0058] The first training process aims to at least make the first emotion of the response speech output by the speech generation model match the second emotion of the sample dialogue content input to the speech generation model according to a preset emotion matching relationship.
[0059] In this embodiment, the speech generation model can adopt an encoder-decoder structure, such as the Transformer model.
[0060] Taking the training data of a pre-trained speech generation model as an example, which includes sample dialogue speech, sample speech text, sample response text, and sample response speech, the emotion in the sample response speech matches the emotion in the sample dialogue speech. Sample dialogue speech, sample speech text, and sample response text generate sample prompt information. The encoder of the Transformer model encodes the sample prompt information to obtain sample speech encoding, and the decoder of the Transformer model decodes the sample speech encoding to generate the response speech. The response speech is compared with the sample response speech, and the Transformer model is optimized based on the comparison results, thus generating a trained speech generation model.
[0061] In one implementation, the step S130 of generating a response voice corresponding to the first dialogue content based on the prompt information includes:
[0062] The prompt information is encoded to obtain speech code;
[0063] The voice encoding is decoded to obtain the response voice corresponding to the first dialogue content.
[0064] Specifically, a prompting-based learning scheme is used to train the encoder-decoder structure (such as the Transformer model), that is, by giving the model prompts, it learns directly based on the task. In this embodiment, such as... Figure 3As shown, the dialogue speech, its corresponding speech text, and the response text are encoded separately to obtain speech encoding, speech text encoding, and response text encoding (i.e., the second encoding). These three encodings are then concatenated to obtain the prompt information. The encoder encodes the prompt information to obtain the speech encoding, and the decoder decodes it to obtain the corresponding response speech. Therefore, by using the dialogue speech, speech text, and response text as prompt information, the trained encoder-decoder will encode the prompt information to obtain the speech encoding, which contains the complete emotional information of the dialogue speech. Decoding this speech encoding allows the determination of the response speech that matches the emotional information of the dialogue speech.
[0065] In one implementation, step S110 of determining the reply text corresponding to the first dialogue content includes:
[0066] Determine the audio text corresponding to the content of the first dialogue;
[0067] Determine the semantic information of the spoken text, wherein the semantic information includes at least one of intent, entity, and grammar;
[0068] Based on the semantic information, determine the response text corresponding to the content of the first dialogue.
[0069] Specifically, when the first dialogue content is speech, the speech can be recognized using existing speech recognition models (such as Whisper model, HMM speech recognition model, acoustic model, etc.) to obtain speech text. The speech text is then preprocessed to obtain corresponding semantic information, that is, converted into a form that a computer can understand for further processing and response. Preprocessing includes at least one of the following: word segmentation, part-of-speech tagging, named entity recognition, syntactic analysis, and intent recognition. Optionally, word segmentation divides the speech text into individual words or phrases. Part-of-speech tagging determines the part of speech of the segmented words, such as nouns, verbs, or adjectives. Named entity recognition identifies entities in the speech text, such as names of people, places, and organizations. Syntactic analysis analyzes the text structure of the speech text according to grammatical rules to understand the relationships between words. Intent recognition analyzes the content and context of the speech text to determine the user's intent. Then, based on the intent expressed in the speech text, the corresponding response text is determined from a pre-set corpus. The pre-set corpus can be a collection of multiple sets of correspondences between intents and texts. In this way, since the semantic information is determined based on the speech text, the response text determined based on the semantic information has a higher degree of matching with the speech text.
[0070] In one implementation, determining the response text corresponding to the first dialogue content based on the semantic information includes:
[0071] Obtain the historical dialogue content preceding the first dialogue content;
[0072] Based on the semantic information, the first dialogue content, and the historical dialogue content, the target intent of the first dialogue content is determined;
[0073] Generate a response text corresponding to the content of the first dialogue based on the stated target intent.
[0074] For example, the historical dialogue content represents the content received at the previous moment. Optionally, continuous text can be crawled from web pages or public articles beforehand as sample context texts, and the intent of the text can be determined. The sample context texts are used as input to the model, and the intent of the text is used as the output to train the neural network model, resulting in a trained text intent recognition model. Optionally, the neural network model (such as a language model) can be trained in advance based on the intent of the text and its corresponding sample response text, resulting in a trained text response model.
[0075] Specifically, the dialogue speech is subjected to speech recognition to obtain speech text. The speech text is preprocessed to obtain the corresponding intent. The historical dialogue content and the first dialogue content are input into a trained text intent recognition model to obtain the text intent. The text intent and the target intent are combined. It is evident that because the first dialogue content and the historical dialogue content have a contextual relationship, understanding this contextual relationship helps in understanding the intent of the first dialogue content.
[0076] Alternatively, the semantic information of the preceding and following text of the sample can be used as input to the model, while the intent of the text can be used as the output to train the neural network model, resulting in a trained text intent recognition model. In this way, after obtaining the semantic information of the spoken text, the semantic information, the content of the first dialogue, and the content of the historical dialogue can be input into the text intent recognition model to output the target intent.
[0077] Then, as Figure 4As shown, the target intent is input into a trained text response model, which outputs the corresponding response text. The dialogue speech, its corresponding speech-text, and the response text are then encoded separately to obtain speech encoding, speech-text encoding, and response text encoding (i.e., the second encoding). These three encodings are concatenated to obtain the prompt information. This prompt information is then input into a trained speech generation model (i.e., the Transformer model), where it serves as the prompt content for the attention mechanism. The Transformer model's encoder autoregressively generates the corresponding speech encoding, which is then decoded by the Transformer model's decoder to generate the corresponding response speech. Since the emotion in the response speech is no longer obtained through a concatenation of emotion recognition, response text, and emotional speech synthesis, there are no cascading errors, resulting in more accurate emotional expression. Furthermore, because the emotional flow is no longer achieved through discrete emotion tags, the expressed emotions are more nuanced and diverse. Based on these two points, the speech generated by this method enhances the human-computer interaction experience, enabling a more empathetic interactive experience.
[0078] Exemplary device
[0079] Correspondingly, Figure 5 This is a schematic diagram of a speech generation apparatus according to an embodiment of the present application. In an exemplary embodiment, a speech generation apparatus is provided, including:
[0080] The determining module 510 is used to determine the response text corresponding to the first dialogue content; wherein, the first dialogue content includes dialogue voice and / or voice text corresponding to the dialogue voice;
[0081] The generation module 520 is used to generate prompt information based on the first dialogue content and the reply text, wherein the prompt information includes the emotional information of the first dialogue content;
[0082] The synthesis module 530 is used to generate a response voice corresponding to the first dialogue content based on the prompt information, wherein the emotion of the response voice conforms to a preset emotion matching relationship with the emotion of the first dialogue content.
[0083] In one implementation, generating prompt information based on the first dialogue content and the reply text, and generating a reply voice corresponding to the first dialogue content based on the prompt information, includes:
[0084] The first dialogue content and the reply text are input into a pre-trained speech generation model to obtain the reply speech corresponding to the first dialogue content.
[0085] The speech generation model generates prompt information based on the first dialogue content and the reply text, and generates a reply speech corresponding to the first dialogue content based on the prompt information.
[0086] In one implementation, the speech generation model is trained through a first training process;
[0087] The first training process aims to at least make the first emotion of the response speech output by the speech generation model match the second emotion of the sample dialogue content input to the speech generation model according to a preset emotion matching relationship.
[0088] In one implementation, the generation module 520 includes:
[0089] The encoding module is used to encode the first dialogue content and the reply text respectively to obtain a first code corresponding to the first dialogue content and a second code corresponding to the reply text. The first code contains sentiment encoding information of the first dialogue content.
[0090] The first encoding module is used to fuse the first encoding and the second encoding to obtain the prompt information.
[0091] In one implementation, the first dialogue content includes dialogue voice and voice text corresponding to the dialogue voice;
[0092] The first encoding module is further configured to encode the dialogue speech in the first dialogue content and the corresponding speech text, respectively, to obtain the dialogue speech encoding and the speech text encoding.
[0093] In one embodiment, the synthesis module 530 includes:
[0094] The second encoding module is used to encode the prompt information to obtain voice encoding;
[0095] The decoding module is also used to decode the voice encoding to obtain a response voice corresponding to the first dialogue content.
[0096] In one implementation, the determining module includes:
[0097] The speech recognition module is used to determine the speech text corresponding to the content of the first dialogue;
[0098] A semantic conversion module is used to determine the semantic information of the speech text, wherein the semantic information includes at least one of intent, entity, and grammar;
[0099] The processing module is used to determine the response text corresponding to the first dialogue content based on the semantic information.
[0100] In one implementation, the processing module is further configured to:
[0101] Obtain the historical dialogue content preceding the first dialogue content;
[0102] Based on the semantic information, the first dialogue content, and the historical dialogue content, the target intent of the first dialogue content is determined;
[0103] Generate a response text corresponding to the content of the first dialogue based on the stated target intent.
[0104] The speech generation apparatus provided in this embodiment belongs to the same concept as the speech generation method provided in the above embodiments of this application. It can execute the speech generation method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects for executing the speech generation method. Technical details not described in detail in this embodiment can be found in the specific processing content of the speech generation method provided in the above embodiments of this application, and will not be repeated here.
[0105] Exemplary electronic devices
[0106] Another embodiment of this application also provides an electronic device, see [link to relevant documentation] Figure 6 As shown, the device includes:
[0107] Memory 600 and processor 610;
[0108] The memory 600 is connected to the processor 610 and is used to store programs;
[0109] The processor 610 is configured to implement the speech generation method disclosed in any of the above embodiments by running the program stored in the memory 600.
[0110] Specifically, the aforementioned electronic device may also include: a bus, a communication interface 620, an input device 630, and an output device 640.
[0111] The processor 610, memory 600, communication interface 620, input device 630, and output device 640 are interconnected via a bus. Among them:
[0112] A bus can include a pathway for transmitting information between various components of a computer system.
[0113] The processor 610 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0114] The processor 610 may include a main processor, as well as a baseband chip, modem, etc.
[0115] The memory 600 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 600 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.
[0116] Input device 630 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.
[0117] Output device 640 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.
[0118] The communication interface 620 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.
[0119] The processor 610 executes the program stored in the memory 600 and calls other devices, and can be used to implement the various steps of any of the speech generation methods provided in the above embodiments of this application.
[0120] Exemplary computer program products and storage media
[0121] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the speech generation methods according to various embodiments of this application as described in the "Exemplary Methods" section of this specification.
[0122] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this application. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0123] Furthermore, embodiments of this application may also be storage media storing computer programs, the computer programs being executed by a processor in the steps of the speech generation methods according to various embodiments of this application described in the "Exemplary Methods" section above. The specific working content of the above-mentioned electronic device, as well as the specific working content of the above-mentioned computer program product and the computer program on the storage medium being run by a processor, can all be found in the content of the above-mentioned method embodiments, and will not be repeated here.
[0124] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0125] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0126] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.
[0127] The modules and sub-modules in the various embodiments of the present application's devices and terminals can be merged, divided, and deleted according to actual needs.
[0128] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.
[0129] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.
[0130] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.
[0131] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0132] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0133] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0134] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A speech generation method, characterized in that, include: Determine the response text corresponding to the first dialogue content; wherein, the first dialogue content includes dialogue voice and / or voice text corresponding to the dialogue voice; Based on the first dialogue content and the reply text, a prompt message is generated. The prompt message is obtained by concatenating the encoding results corresponding to the first dialogue content and the reply text respectively. The prompt message carries the emotional information of the first dialogue content, and the emotional information is represented by the original content of the first dialogue content. Based on the prompt information, a response voice corresponding to the first dialogue content is generated, and the emotion of the response voice matches the emotion of the first dialogue content according to a preset emotion matching relationship.
2. The method according to claim 1, characterized in that, The step of generating prompt information based on the first dialogue content and the reply text, and generating a reply voice corresponding to the first dialogue content based on the prompt information, includes: The first dialogue content and the reply text are input into a pre-trained speech generation model to obtain the reply speech corresponding to the first dialogue content. The speech generation model generates prompt information based on the first dialogue content and the reply text, and generates a reply speech corresponding to the first dialogue content based on the prompt information.
3. The method according to claim 2, characterized in that, The speech generation model is trained through a first training process; The first training process aims to at least make the first emotion of the response speech output by the speech generation model match the second emotion of the sample dialogue content input to the speech generation model according to a preset emotion matching relationship.
4. The method according to any one of claims 1 to 3, characterized in that, Based on the first dialogue content and the reply text, a prompt message is generated, including: The first dialogue content and the reply text are encoded respectively to obtain a first code corresponding to the first dialogue content and a second code corresponding to the reply text. The first code contains the sentiment encoding information of the first dialogue content. The first code and the second code are fused together to obtain the prompt information.
5. The method according to claim 4, characterized in that, The first dialogue content includes the dialogue voice and the corresponding voice text; Encoding the first dialogue content and the reply text respectively yields a first encoding corresponding to the first dialogue content, including: The dialogue voice in the first dialogue content and the corresponding voice text are encoded respectively to obtain the dialogue voice code and the voice text code.
6. The method according to any one of claims 1 to 3, characterized in that, The step of generating a response voice corresponding to the first dialogue content based on the prompt information includes: The prompt information is encoded to obtain speech code; The voice encoding is decoded to obtain the response voice corresponding to the first dialogue content.
7. The method according to any one of claims 1 to 3, characterized in that, The determination of the reply text corresponding to the content of the first dialogue includes: Determine the audio text corresponding to the content of the first dialogue; Determine the semantic information of the spoken text, wherein the semantic information includes at least one of intent, entity, and grammar; Based on the semantic information, determine the response text corresponding to the content of the first dialogue.
8. The method according to claim 7, characterized in that, The step of determining the response text corresponding to the first dialogue content based on the semantic information includes: Obtain the historical dialogue content preceding the first dialogue content; Based on the semantic information, the first dialogue content, and the historical dialogue content, the target intent of the first dialogue content is determined; Generate a response text corresponding to the content of the first dialogue based on the stated target intent.
9. A speech generation device, characterized in that, include: The determining module is used to determine the response text corresponding to the first dialogue content; wherein, the first dialogue content includes dialogue voice and / or voice text corresponding to the dialogue voice; The generation module is used to generate prompt information based on the first dialogue content and the reply text. The prompt information is obtained by concatenating the encoding results corresponding to the first dialogue content and the reply text respectively. The prompt information carries the emotional information of the first dialogue content, and the emotional information is represented by the original content of the first dialogue content. The synthesis module is used to generate a response voice corresponding to the first dialogue content based on the prompt information, wherein the emotion of the response voice conforms to a preset emotion matching relationship with the emotion of the first dialogue content.
10. An electronic device, characterized in that, include: Memory and processor; The memory is connected to the processor and is used to store programs; The processor implements any one of the speech generation methods as described in claims 1 to 8 by running the program in the memory.
11. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements any one of the speech generation methods as described in claims 1 to 8.
Citation Information
Patent Citations
Speech synthesis method, device and equipment and readable storage medium
CN113593521A