Voice interaction method and device, intelligent terminal and readable storage medium
By collecting and analyzing the user's multimodal information, identifying the user's emotional state, and generating corresponding digital human interaction information, the problem of digital human action triggering is solved, and the interaction ability is improved.
Patent Information
- Application Number
- CN202510427452.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Digital people’s movement triggers are not flexible enough under complex emotional changes, resulting in poor interaction ability.
By collecting multimodal information of users, including voice data and face data, extracting voice features and face features, identifying the user's emotional state, and generating digital human interaction information based on the emotional state, controlling the digital human interaction with the user.
It improves the flexibility of digital people to trigger movements under complex emotional changes, enhances interaction ability, and allows digital people to respond more accurately to users' emotional state.
Smart Images

Figure CN120148560A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and specifically relates to a voice interaction method, device, intelligent terminal, and readable storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology, emotion recognition has become a research direction in the field of human-computer interaction. Multimodal emotion recognition systems use multiple data sources such as images, sounds, voices, and postures to comprehensively analyze the important emotional states of users, and on this basis, adjust the human-computer interaction mode to fully realize multiple fields such as intelligent customer service, virtual assistants, entertainment, and education.
[0003] However, current interaction solutions usually regard emotion recognition and action triggering as efficient and independent channels, lacking a mechanism to achieve a close connection between emotional states and action performances. This results in inflexible action triggering of digital humans under complex emotional changes and poor interaction capabilities. Summary of the Invention
[0004] To address the above technical problems, this application provides a voice interaction method, device, intelligent terminal, and readable storage medium, which can solve the problem of inflexible action triggering of digital humans under complex emotional changes, thereby improving the interaction ability.
[0005] To solve the above technical problems, this application provides a voice interaction method, including:
[0006] Respond to the user's voice trigger operation and collect the user's multimodal information;
[0007] Obtain the voice features corresponding to the voice data in the multimodal information and the face features corresponding to the face data in the multimodal information;
[0008] Determine the emotional state corresponding to the user according to the voice features and face features;
[0009] Generate the interaction information corresponding to the digital human based on the protocol parsing message corresponding to the emotional state, and control the digital human to perform dynamic interaction with the user according to the interaction information.
[0010] Optionally, in some embodiments of this application, the determining the emotional state corresponding to the user according to the voice features and face features includes:
[0011] Process the voice features based on a preset voice model to identify the voice emotion corresponding to the user;
[0012] Identify the face emotion corresponding to the user based on the face features and preset reference face features.
[0013] Optionally, in some embodiments of the present application, identifying the facial emotion corresponding to the user based on the facial features and preset reference facial features includes:
[0014] Matching the facial features with the preset reference facial features;
[0015] Determining the matched reference facial feature as the target facial feature;
[0016] Determining the emotion corresponding to the target facial feature as the facial emotion corresponding to the user.
[0017] Optionally, in some embodiments of the present application, it further includes:
[0018] Obtaining the sample data corresponding to multiple sample users;
[0019] Annotating the sample data and converting the annotated sample data into data in a preset format;
[0020] Dividing the sample data after format conversion into a training set and a validation set;
[0021] Training a preset basic model using the training set and validating the trained basic model using the validation set to obtain an expression detection model, where the expression detection model is used to detect facial expressions.
[0022] Optionally, in some embodiments of the present application, generating interaction information corresponding to the digital human based on the protocol parsing message corresponding to the emotional state, and controlling the digital human to perform dynamic interaction with the user according to the interaction information includes:
[0023] Converting the emotional state into standardized protocol information;
[0024] Parsing the protocol information;
[0025] Generating interaction information corresponding to the digital human based on a preset large language model and the parsing result, and controlling the digital human to perform dynamic interaction with the user according to the interaction information.
[0026] Optionally, in some embodiments of the present application, generating interaction information corresponding to the digital human based on a preset large language model and the parsing result, and controlling the digital human to perform dynamic interaction with the user according to the interaction information includes:
[0027] Generating interaction actions and interaction texts corresponding to the digital human based on a preset large language model and the parsing result;
[0028] Generating interaction voice corresponding to the interaction text;
[0029] Control the digital human to perform dynamic interaction with the user according to the interactive voice and interactive actions.
[0030] Optionally, in some embodiments of the present application, generating the interactive voice corresponding to the interactive text includes:
[0031] Convert the interactive text into text audio according to speech synthesis technology;
[0032] Adjust the text audio based on the emotional state of the user to obtain the interactive voice corresponding to the interactive text.
[0033] Optionally, in some embodiments of the present application, generating the interactive actions corresponding to the digital human based on the preset large language model and parsing result includes:
[0034] Based on the preset large language model and parsing result, obtain the interactive actions corresponding to the digital human from the preset action sequence.
[0035] Correspondingly, the present application also provides a voice interaction device, including:
[0036] A collection module, configured to collect the multi-modal information of the user in response to a voice trigger operation of the user;
[0037] An acquisition module, configured to acquire the voice feature corresponding to the voice data in the multi-modal information and the face feature corresponding to the face data in the multi-modal information;
[0038] An identification module, configured to determine the emotional state corresponding to the user according to the voice feature and the face feature;
[0039] An interaction module, configured to generate the interactive information corresponding to the digital human based on the protocol parsing message corresponding to the emotional state, and control the digital human to perform dynamic interaction with the user according to the interactive information.
[0040] The present application also provides an intelligent terminal, including a memory and a processor, where the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0041] The present application also provides a computer storage medium, where the computer storage medium stores a computer program, and the computer program implements the steps of the above method when executed by a processor.
[0042] As described above, the present application provides a voice interaction method, apparatus, intelligent terminal, and readable storage medium. After responding to a user's voice trigger operation and collecting the user's multi-modal information, the voice features corresponding to the voice data in the multi-modal information and the face features corresponding to the face data in the multi-modal information are obtained. Then, based on the voice features and face features, the emotional state corresponding to the user is determined. Finally, based on the protocol parsing message corresponding to the emotional state, the interaction information corresponding to the digital human is generated, and the digital human is controlled to perform dynamic interaction with the user according to the interaction information. In the voice interaction solution provided by the present application, the emotional state corresponding to the user can be determined based on the voice features and face features, and the interaction information corresponding to the digital human can be generated based on the protocol parsing message corresponding to the emotional state, which can ensure that the interaction information of the digital human is associated with the emotional state of the user, so as to ensure that in subsequent interactions, the digital human performs dynamic interaction with the user based on the interaction information. Therefore, the problem that the action trigger of the digital human is not flexible enough under complex emotional changes can be solved, thereby improving the interaction ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0044] Figure 1 is a schematic structural diagram of the voice interaction system provided by the embodiment of the present application;
[0045] Figure 2 is a schematic flowchart of the voice interaction method provided by the embodiment of the present application;
[0046] Figure 3 is a schematic flowchart of emotion recognition in the voice interaction method provided by the embodiment of the present application;
[0047] Figure 4 is a schematic flowchart of voice emotion detection in the voice interaction method provided by the embodiment of the present application;
[0048] Figure 5 is a schematic flowchart of face emotion detection in the voice interaction method provided by the embodiment of the present application;
[0049] Figure 6 is a schematic flowchart of intonation synthesis in the voice interaction method provided by the embodiment of the present application;
[0050] Figure 7It is a schematic structural diagram of the voice interaction device provided by an embodiment of the present application
[0051] Figure 8 It is a schematic structural diagram of the intelligent terminal provided by an embodiment of the present application.
[0052] The realization of the purpose of the present application, functional features and advantages will be further described in conjunction with the embodiments with reference to the accompanying drawings. Through the above-mentioned accompanying drawings, the specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Specific Embodiments
[0053] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0054] It should be noted that in this document, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, components, features, and elements with the same name in different embodiments of the present application may have the same meaning or different meanings, and their specific meanings need to be determined by their explanations in the specific embodiments or further in combination with the context in the specific embodiments.
[0055] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0056] In subsequent descriptions, the suffixes such as "module", "component" or "unit" used to represent elements are only for the convenience of the description of the present application, and they have no specific meaning in themselves. Therefore, "module", "component" or "unit" can be used interchangeably.
[0057] The following specifically describes the embodiments related to the present application. It should be noted that the order of description of the embodiments in the present application does not limit the priority order of the embodiments.
[0058] The embodiments of the present application provide a voice interaction method, device, storage medium, and intelligent terminal. Specifically, the voice interaction method of the embodiments of the present application can be executed by an intelligent terminal or a server. Among them, the intelligent terminal can be a terminal. The terminal can be an intelligent terminal such as a smart phone, a tablet computer, a laptop computer, a touch screen, a game console, a personal computer (PC), a personal digital assistant (PDA), etc. The terminal can also include a client, and the client can be a media player client or an instant messaging client, etc.
[0059] For example, when the voice interaction method runs on an intelligent terminal, the intelligent terminal can respond to the user's voice trigger operation, collect the user's multi-modal information, and then obtain the voice features corresponding to the voice data in the multi-modal information and the face features corresponding to the face data in the multi-modal information. Then, according to the voice features and the face features, determine the emotional state corresponding to the user. Finally, based on the protocol parsing message corresponding to the emotional state, generate the interaction information corresponding to the digital human, and control the digital human to perform dynamic interaction with the user according to the interaction information. The manner in which the intelligent terminal provides the graphical user interface to the user can include various ways. For example, it can be rendered and displayed on the display screen of the intelligent terminal, or the graphical user interface can be presented through holographic projection. For example, the intelligent terminal can include a touch display screen and a processor, and the touch display screen is used to present the graphical user interface and receive the operation instructions generated by the user acting on the graphical user interface.
[0060] Please refer to Figure 1 , Figure 1Schematic diagram of the system of the voice interaction device provided by the embodiment of the present application. The system may include at least one intelligent terminal 1000 and at least one server or personal computer 2000. The intelligent terminal 1000 held by the user can be connected to different servers or personal computers through a network. The intelligent terminal 1000 can be an intelligent terminal with computing hardware, and the computing hardware can support and execute software products corresponding to multimedia. In addition, the intelligent terminal 1000 can also have one or more multi-touch sensitive screens for sensing and obtaining the input of touch or slide operations performed by the user at multiple points on one or more touch display screens. In addition, the intelligent terminal 1000 can be interconnected with the server or personal computer 2000 through a network. The network can be a wireless network or a wired network. For example, the wireless network is a wireless local area network (WLAN), local area network (LAN), cellular network, 2G network, 3G network, 4G network, 5G network, etc. In addition, different intelligent terminals 1000 can also use their own Bluetooth network or hotspot network to connect to other embedded platforms or to connect to servers and personal computers, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0061] The embodiment of the present application provides a voice interaction method, and this method can be executed by an intelligent terminal or a server. The embodiment of the present application takes the voice interaction method being executed by an intelligent terminal as an example for illustration. Among them, the intelligent terminal includes a touch display screen and a processor, and the touch display screen is used to present a graphical user interface and receive operation instructions generated by the user acting on the graphical user interface. When the user operates on the graphical user interface through the touch display screen, the graphical user interface can control the content on the local side of the intelligent terminal in response to the received operation instructions, or can also control the content on the server side in response to the received operation instructions. For example, the operation instructions generated by the user acting on the graphical user interface include instructions for processing initial audio data, and the processor is configured to start the corresponding application program after receiving the instructions provided by the user. In addition, the processor is configured to render and draw the graphical user interface associated with the application program on the touch display screen. The touch display screen is a multi-touch sensitive screen capable of sensing touch or slide operations performed simultaneously at multiple points on the screen. When the user performs a touch operation on the graphical user interface with a finger, when the graphical user interface detects the touch operation, it controls the corresponding operation to be displayed in the graphical user interface of the application.
[0062] The voice interaction solution provided by this application can determine the emotional state corresponding to the user based on voice features and facial features, and parse messages based on the protocol corresponding to the emotional state to generate interaction information corresponding to the digital human, which can ensure that the interaction information of the digital human is associated with the emotional state of the user, so as to ensure that in subsequent interactions, the digital human conducts dynamic interactions with the user based on this interaction information. Therefore, the problem that the action triggering of the digital human is not flexible enough under complex emotional changes can be solved, thereby improving the interaction ability.
[0063] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the priority order of the embodiments.
[0064] A voice interaction method includes: responding to a voice trigger operation of a user, and collecting multimodal information of the user; obtaining a voice feature corresponding to voice data in the multimodal information and a facial feature corresponding to face data in the multimodal information; recognizing an emotional state corresponding to the user based on the voice emotion corresponding to the voice feature and the facial expression corresponding to the facial feature; parsing a message based on the protocol corresponding to the emotional state to generate interaction information corresponding to the digital human, and controlling the digital human to conduct dynamic interactions with the user according to the interaction information.
[0065] Please refer to Figure 2 , Figure 2 , which is a schematic flowchart of the voice interaction method provided by the embodiment of this application. The specific process of this voice interaction method can be as follows:
[0066] 101. Respond to a voice trigger operation of a user, and collect multimodal information of the user.
[0067] Among them, the voice trigger operation refers to the behavior of the user having a conversation with the digital human through voice input. When the user emits voice, this behavior will start a series of subsequent processing processes of the system, including multimodal information collection, emotion recognition, digital human action triggering, intonation synthesis, etc., which is the starting point of the entire interaction process. For example, when the user asks the digital human "What's the weather like tomorrow", the behavior of saying this sentence is the voice trigger operation.
[0068] Multimodal information is information from multiple sensors and data sources, mainly including content in aspects such as vision and voice. When the user interacts with the digital human, the visual information refers to the facial features of the user collected by the camera, and these features can assist in judging the user's facial expression and then inferring emotions; the voice information not only includes the content of the user's input words, but also covers the emotional information in the words and the surrounding background sounds, such as music, applause, etc. These multimodal information will be integrated and transmitted to the large language model (LLM) as prompt words for multimodal judgment of the emotion of the user's input voice, providing a basis for subsequent interactions of the digital human.
[0069] When the user converses with the digital human through voice input, the system will activate the multi-modal information collection mechanism. On the one hand, the camera will collect the user's facial features, which can reflect the user's facial expressions and thus assist in judging the user's emotional state at this time. On the other hand, the system will read the audio input by the user, not only judging the emotion in the user's words, but also collecting other background sounds, such as music, applause and other information. These multi-modal information from vision (facial features) and audition (voice and background sounds) are transmitted to the large language model (LLM) as prompt words to achieve the purpose of multi-modal judgment of the emotion of the user's input voice.
[0070] When the user excitedly shouts: "The special effects of that movie are simply amazing!" At this time, the voice trigger operation is started. The camera captures the user's facial features such as wide-open eyes, upturned corners of the mouth, and slightly trembling facial muscles due to excitement; at the same time, the system reads the user's audio, identifies the excited emotion in the voice, and also collects the exciting melody of the background music generated by the movie playback segment in the surrounding environment and the slight exclamations of other audiences. These multi-modal information such as facial features, voice emotion, and background sounds are all collected and transmitted to the LLM as prompt words to help more accurately judge the user's emotional state at this time.
[0071] 102. Obtain the voice features corresponding to the voice data in the multi-modal information and the face features corresponding to the face data in the multi-modal information.
[0072] Voice features are a series of information extracted from the user's input audio, including the emotion, content and background sound information of the words themselves. In the speech emotion recognition (SER) task, the system will analyze the acoustic features such as the pitch, volume, and speech rate of the user's speech to judge the emotion in the words. For example, when angry, the volume may be higher and the speech rate may be faster. At the same time, it will also collect the background sounds in the surrounding environment, such as music, applause, etc. These information together constitute the voice features and are used for comprehensive analysis of the user's emotional state. For example, in a concert scene, the user shouts excitedly at the digital human. In addition to the excited emotion expressed in the user's words, the background sounds such as the music rhythm and the cheers of the audience at the scene are all part of the voice features.
[0073] Facial features refer to the features related to the user's face captured by the camera, mainly used to judge facial expressions and then assist in judging the user's emotions. The system detects the face in the camera, extracts and matches feature points, and then compares them with known expression patterns to obtain the emotions presented by the face. For example, raised eyebrows, widened eyes, and upturned corners of the mouth may represent joy; frowning eyebrows and a serious look in the eyes may indicate dissatisfaction or worry. During the interaction with the digital human, these facial features can intuitively reflect the user's current emotions, complementing the voice features to more accurately identify the user's emotions. For example, when the user shares interesting travel stories with the digital human, facial features such as a smiling face and bright eyes can help the system better understand the user's happy mood.
[0074] Obtaining the voice features and facial features in the multimodal information is to comprehensively analyze the user's emotions from different dimensions. The voice features are obtained by reading the audio input by the user. In addition to identifying the voice content, it also includes judging the emotions in the speech and collecting background sound information such as music and applause. These information are integrated as the voice basis for analyzing the user's emotions. The facial features are obtained by the camera capturing the user's facial features, and the captured facial features are judged to obtain the user's facial expression, so as to assist in judging the user's emotions.
[0075] For example, the user is watching a live sports game and communicating with the digital human assistant about their feelings. The user shouts excitedly: "That goal just now was so beautiful!" In this scenario, when the system obtains the voice features corresponding to the voice data, it will identify the high-frequency and high-volume characteristics in the user's voice, and combine the content of the speech to judge the excited emotion. At the same time, the cheers in the background and the excited voices of the commentators will also be collected as part of the voice features. In terms of obtaining the facial features corresponding to the facial data, the camera captures features such as the user's widened eyes, upturned corners of the mouth, and slightly trembling cheek muscles due to excitement. These features are all used to assist in judging that the user is in an excited emotional state. These voice features and facial features combined can more accurately judge the user's emotions, providing a more accurate basis for the subsequent interaction of the digital human.
[0076] 103. Identify the emotional state corresponding to the user based on the voice emotion corresponding to the voice features and the facial expression corresponding to the facial features.
[0077] The emotional state is a judgment result of the user's emotional state during the interaction with the digital human, and is obtained through comprehensive analysis based on multimodal information.
[0078] Specifically, a speech model can be used to analyze the collected speech information. The speech model can perform tasks such as automatic speech recognition (ASR), language identification (LID), speech emotion recognition (SER), and audio event detection (AED). In addition, the corresponding facial emotion of the user can be recognized based on the preset reference facial features and the collected facial features. That is, optionally, in some embodiments of the present application, the step of "identifying the emotional state corresponding to the user based on the speech emotion corresponding to the speech feature and the facial expression corresponding to the facial feature" may specifically include:
[0079] Processing the speech feature based on a preset speech model to identify the speech emotion corresponding to the user;
[0080] Identifying the facial expression corresponding to the user based on the facial feature and the preset reference facial feature;
[0081] Using the speech emotion and the facial expression as emotion cue words, and inputting the emotion cue words into a preset large language model to obtain the emotional state corresponding to the user.
[0082] For example, specifically, a pre-constructed speech model can be obtained. The speech model is designed to analyze speech features, which may include information such as the pitch, volume, speech rate, intonation changes of the speech, as well as vocabulary, semantics, etc. The model is trained with a large amount of speech data to learn the emotion patterns represented by different combinations of speech features. When new speech is input, the model extracts the speech features therein and matches them with the learned patterns to determine the speech emotion of the user. For example, a large number of angry speech in the training data has the characteristics of loud volume, fast speech rate, and more high-frequency words. After the model remembers these patterns, when encountering new speech that conforms to such characteristics, it may recognize that the user is in an angry state.
[0083] It should be noted that the speech model includes automatic speech recognition (ASR), language identification speech recognition (LID), speech emotion recognition (SER), and audio event detection (AED) <lid>Indicates the LID task. If <lid>is prefixed, then train a model to predict the language token at the corresponding position in the output. During the training phase, randomly replace <lid>Based on the ground-truth language tokens with a probability of 0.8, the model can arbitrarily predict language tokens or be configured with specified language tokens during the inference phase. <ser>Indicates the SER task. If <ser>If pre-added, train a model to predict speech emotion labels and output them at the corresponding positions. <AEC> represents the AEC task. If <AEC> is preposed, train a model to predict audio event labels at the corresponding positions in the output. <ITN> or <NoITN> specify the transcription style. If <itn>If <ITN> is provided, the model is trained to use inverse text normalization (ITN) and punctuation in the text. If <NoITN> is provided, the model is trained to transcribe without ITN and punctuation. During the training phase, cross-entropy loss is used to optimize the LID, SER, and AEC tasks. CTC loss is used for optimization to improve the ability to recognize speech emotions. For example, by analyzing acoustic features such as the pitch, volume, and speaking speed of the speech, as well as background sounds, the emotional elements in the speech are recognized. If the speech has a large volume, a fast speaking speed, and the background sound is cheering, the model may recognize the user's excited emotion.
[0084] In addition, the reference face features may come from a large number of face samples with labeled emotion categories, covering the face feature data corresponding to different expressions (such as happy, sad, angry, etc.), including features such as facial muscle movements and changes in the morphology of facial features. When the user's face features are obtained, they are compared with the preset reference face features. By calculating similarities and other methods, the facial expression corresponding to the most matching reference feature is found to determine the user's facial emotion. For example, in the preset reference face features of a happy expression, there are characteristics such as the corners of the mouth turning up, the eyes squinting, and the cheekbones rising. If the user's face features are highly similar to them, it can be judged that the user is in a happy emotional state at this time.
[0085] For example, in an intelligent customer service scenario, the user consults a digital human customer service. When the user speaks, the voice is very loud and the speaking speed is very fast, constantly emphasizing "How could your service be so bad". The preset speech model analyzes these speech features and finds that they conform to the speech feature pattern of the angry emotion in the model, thus recognizing the user's angry speech emotion. At the same time, the camera captures the user's face features, with the eyebrows frowning, the eyes sharp, and the corners of the mouth downturned, which match the features of the angry expression in the preset reference face features, and then determines that the user's corresponding facial emotion is also angry. Combining the speech and facial emotion recognition results, the digital human customer service can more accurately judge that the user is in an angry state, and then adjust the reply strategy, such as using a more gentle and soothing tone to reply to the user to improve the user experience.
[0086] Next, combine the preprocessed voice emotion and facial expression information into a prompt suitable for input to a Large Language Model (LLM). This can be text concatenation, such as "The user's voice emotion is happy, and the user's facial expression is a smile", or more complex format design can be carried out according to the characteristics and requirements of the LLM. Some LLMs may have specific requirements for the input format, such as adding specific instructions or tags before the prompt to guide the model to better understand the task. When constructing the prompt, some context information can also be added, such as the topic of the conversation, scene description, etc., to help the LLM more accurately analyze the emotional state. For example, "In a conversation about travel plans, the user's voice emotion is happy, and the user's facial expression is a smile". The LLM uses its pre-trained knowledge and language understanding ability to analyze and process the prompt, and combines the language patterns and emotional association knowledge it has learned to infer the user's emotional state.
[0087] Optionally, in some embodiments of the present application, the step of "identifying the user's corresponding facial expression based on the facial features and the preset reference facial features" may specifically include:
[0088] Match the facial features with the preset reference facial features;
[0089] Determine the matched reference facial feature as the target facial feature, and obtain the facial expression corresponding to the target facial feature;
[0090] Determine the obtained facial expression as the user's corresponding facial expression.
[0091] Specifically, a series of reference facial features representing different facial expressions are set in advance. When the facial features of the user are obtained, they are compared with these preset reference features to find the most similar reference feature, and the facial expression corresponding to this reference feature is determined as the user's current facial expression.
[0092] Optionally, in some embodiments of the present application, the interaction method of the present application may specifically further include:
[0093] Obtain sample data corresponding to multiple sample users;
[0094] Annotate the sample data and convert the annotated sample data into data in a preset format;
[0095] Divide the sample data after format conversion into a training set and a validation set;
[0096] Use the training set to train a preset basic model, and use the validation set to validate the trained basic model to obtain an expression detection model for detecting facial expressions.
[0097] Among them, the facial expression detection model can be the YOLOv8 model. Specifically, the sample data should be data related to the user's emotions, covering facial image data of different users in various emotional states. Then, the collected sample data is labeled, and the labeling content may include key information such as the emotional category corresponding to the sample. After that, the labeled sample data is converted into a preset format to make the data meet the input requirements of the YOLOv8 model. In some embodiments of the present application, the data labeling can be converted into the format required by YOLO (such as a.txt file, containing the class number, normalized center coordinates, and width and height), and the data after format conversion is divided into a training set and a validation set. The training set is used to train the YOLOv8 model, enabling the model to learn the features and patterns in the data; the validation set is used to validate the trained model, evaluate the performance of the model, check whether the model accurately learns the patterns in the data, and whether there are problems such as overfitting, ensuring the generalization ability of the model so that it can also perform well when facing new data. The preset YOLOv8 base model is trained using the training set. During the training process, the model is continuously optimized by setting hyperparameters, loading pre-trained weights for transfer learning, etc. After training is completed, the trained YOLOv8 model is validated using the validation set, and the model parameters are further adjusted according to the validation results. Through such a training and validation process, the final model obtained is the facial expression detection model for detecting human facial expressions, that is, the YOLOv8 model. It can process the facial data in the real-time frames obtained by the camera to assist in determining the corresponding facial expressions of the human face, providing support for the subsequent digital human to make corresponding actions and intonation responses according to the user's emotions.
[0098] 104. Parse the protocol message corresponding to the emotional state to generate the interactive information corresponding to the digital human, and control the digital human to perform dynamic interaction with the user according to the interactive information.
[0099] For example, specifically, the emotional state of the user can be parsed through a large language model (LLM). Based on the parsed message, the LLM generates the interactive information corresponding to the digital human. This process comprehensively considers factors such as the user's emotions and the conversation context. For example, if the parsed message indicates that the user is happy because a problem has been solved, the interactive information may be that the digital human expresses congratulations in a cheerful tone and mentions the relevant solution content to enhance the pertinence and naturalness of the interaction. After the digital human display terminal receives the interactive information generated by the LLM, it parses it according to the specified protocol. After parsing, it triggers the actions and intonation of the digital human to achieve dynamic interaction with the user. The digital human may cooperate with the cheerful tone and make actions such as smiling and applauding, allowing the user to feel a more real and natural interaction experience and improving the automation level and user satisfaction of the interaction.
[0100] Optionally, in some embodiments of the present application, the step of "parsing the message based on the protocol corresponding to the emotional state, generating interactive information corresponding to the digital human, and controlling the digital human to dynamically interact with the user according to the interactive information" may specifically include:
[0101] Convert emotional states into standardized protocol information;
[0102] Parse the protocol information;
[0103] Based on the preset large language model and the analysis results, the corresponding interactive information of the digital human is generated, and the digital human is controlled to dynamically interact with the user according to the interactive information.
[0104] Specifically, the emotional state is converted into standardized protocol information in a specific format according to pre-set rules. This process is similar to "translating" the actual emotional state into an information form that the computer system can understand and process, such as converting emotions such as "happy" and "angry" into codes or data structures containing parameters such as emotion category and intensity, so as to facilitate subsequent transmission and processing in the system. The parsing process breaks down complex protocol information into specific instructions or parameters based on established protocol rules. For example, it parses from the protocol information whether the user's current emotion is positive or negative, how strong the emotion is, and other information. LLM will process the parsing results based on internal algorithms and training data. For example, if the parsing results show that the user is in an angry state, LLM may generate interactive information such as "respond to the user with a gentle, apologetic tone, and express that the problem will be solved as soon as possible", which contains appropriate language content, tone style and other information to adapt to different emotional scenarios. According to the parsing results, the actions and intonation of the digital human are triggered. For example, if the interactive information requires a response in a gentle tone, the digital human will adjust the speech synthesis module and output a soft voice; if a soothing action is required, the digital human will show corresponding body movements, thereby conducting natural and smooth dynamic interaction with the user, improving the naturalness and automation level of the user experience.
[0105] Optionally, in some embodiments of the present application, the step of "generating interactive information corresponding to the digital human based on the preset large language model and the analysis result, and controlling the digital human to dynamically interact with the user according to the interactive information" may specifically include:
[0106] Generate interactive actions and interactive texts corresponding to the digital human based on the preset large language model and analysis results;
[0107] Generate interactive speech corresponding to the interactive text;
[0108] Based on interactive voice and interactive actions, the digital human is controlled to dynamically interact with the user.
[0109] Specifically, LLM analyzes user emotions according to established protocols. For example, if the user is identified as being in an angry state, it can be analyzed that the user is dissatisfied with the product or service. Based on this analysis result, LLM generates corresponding interactive actions and interactive texts. The interactive text may be soothing words, such as "I am very sorry for giving you a bad experience, we will solve the problem for you immediately"; the interactive action may be the digital human bowing its head and putting its hands together to express apology, so that the digital human's response is in line with the user's current emotional state, more targeted and natural.
[0110] After the interactive text is generated, the appropriate voice parameters can be selected according to the content, tone and emotional state of the user analyzed. For the above text to appease the user, a gentle and sincere tone, moderate speed and tone may be selected to make the voice sound more friendly, so as to better convey the soothing emotion and enhance the emotional communication effect with the user. After the digital human display terminal receives the interactive voice and interactive action instructions, it parses and performs the corresponding operations according to the specified protocol. The digital human will play the synthesized interactive voice and display the corresponding interactive actions at the same time to interact with the user in real time. For example, the digital human plays the soothing voice while making an apology gesture, giving the user feedback from both visual and auditory aspects, improving the naturalness and automation level of the interactive experience, allowing users to feel the more real and considerate service of the digital human, and realizing efficient human-computer interaction.
[0111] Optionally, in some embodiments of the present application, the step of “generating interactive speech corresponding to the interactive text” may specifically include:
[0112] Convert interactive text into text audio based on speech synthesis technology;
[0113] The text audio is adjusted based on the user's emotional state to obtain the interactive voice corresponding to the interactive text.
[0114] After generating the interactive text corresponding to the digital human through a large language model (LLM), it is necessary to convert this text into an audio form so that the digital human can "speak" it out to interact with the user. Text-to-speech technology can convert text information into sound signals. For example, when the LLM generates an interactive text like "Hello, nice to serve you!", the text-to-speech technology will synthesize this text into an initial text audio according to preset voice parameters, such as selecting a suitable voice tone (possibly a gentle female voice or a kind male voice), setting a certain speaking speed and intonation. This process is the basis for converting text information into audible sounds and is a fundamental step in human-computer voice interaction. After obtaining the user's emotional state, the previously generated text audio will be adjusted based on this state. If it is recognized that the user is in an anxious state, then for the text audio of "Hello, nice to serve you!", the speaking speed may be increased and the intonation may be made more urgent, so that the digital human's response better matches the user's current emotion, allowing the user to feel that the digital human can understand and respond to their emotions, enhancing the naturalness and affinity of the interaction. The text audio adjusted in this way becomes the interactive voice corresponding to the interactive text and is finally played by the digital human to interact with the user.
[0115] Optionally, in some embodiments of the present application, the step of "generating the interactive actions corresponding to the digital human based on the preset large language model and the parsing result" may specifically include:
[0116] Based on the preset large language model and the parsing result, obtain the interactive actions corresponding to the digital human from the preset action sequences.
[0117] In some embodiments of the present application, a series of action sequences are preset, and these action sequences are associated with factors such as different emotional states and dialogue scenarios. The action sequences include the digital human's body movements, facial expressions, etc. For example, actions such as the digital human nodding, smiling, waving, etc., as well as expressions such as surprise, happiness, sadness, etc., are organized into different combinations corresponding to various possible situations. Then, based on the parsing result of the user's emotion by the LLM, the matching interactive actions are selected from the preset action sequences. If the LLM parses that the user is in a happy state, the system will find the actions corresponding to the happy emotion from the preset action sequences, which may be actions such as the digital human showing a bright smile and applauding. These interactive actions can intuitively show the digital human's understanding and response to the user's emotion, enhancing the naturalness and automation level of the interaction, and allowing the user to feel a more real and considerate service experience during the interaction with the digital human.
[0118] To further understand the voice interaction solution of the present application, please refer to Figure 3 , when the user converses with the digital human through voice input, two mechanisms will be triggered. The first mechanism is that the camera will collect the user's facial features at this time, judge the facial features, and thus obtain the user's facial expression at this time to assist in judging the user's emotion at this time. The second mechanism is to read the audio input by the user, judge the emotion in the user's words at this time, and collect some other background sounds such as music and applause. Then, the information collected will be used as a prompt word and transmitted to the LLM to achieve the purpose of multimodal judgment of the emotion of the user's input voice.
[0119] Further, please refer to Figure 4 and Figure 5 , as Figure 4 shown, the speech emotion discrimination module is designed to perform various speech understanding tasks, including automatic speech recognition (ASR), language identification speech recognition (LID), speech emotion recognition (SER), and audio event detection (AED). <lid>Indicates the LID task. If <lid>is prefixed, then train a model to predict the language tokens at the corresponding positions in the output. During the training phase, we randomly replace <lid>Based on the ground truth language tokens with a probability of 0.8, the model can arbitrarily predict language tokens or be configured with specified language tokens during the inference phase. <SER> represents the SER task. If <SER> is pre-added, the model is trained to predict speech emotion labels and output them at the corresponding positions. <aec>Indicates an AEC task. If <aec>If it is preposed, then train a model to predict the audio event label at the corresponding position of the output. <itn>or <noitn>Specify the transcription style. If <itn>is provided, the model is trained to use inverse text normalization (ITN) and punctuation for the text. If <noitn>If provided, the model is trained to transcribe without ITN and punctuation. During the training phase, cross-entropy loss is used to optimize the LID, SER, and AEC tasks. CTC loss is used for optimization.
[0120] As Figure 5 shown, in face detection and emotion recognition, first detect whether there is a human face in the camera. After successful matching, extract and match the feature points, and then conduct a comparison to obtain the emotion of the human face. Finally, refer to Figure 6 , and according to the answer of the emotion for parsing. After successful parsing, modify the reference audio to obtain a speech stream with different intonations, and finally send it back to the user side.
[0121] The above completes the voice interaction process of this application.
[0122] As can be seen from the above, this application provides a voice interaction method. After responding to the user's voice trigger operation and collecting the user's multi-modal information, the voice features corresponding to the voice data in the multi-modal information and the face features corresponding to the face data in the multi-modal information are obtained. Then, based on the voice emotion corresponding to the voice features and the face expression corresponding to the face features, the emotional state corresponding to the user is recognized. Finally, based on the protocol parsing message corresponding to the emotional state, the interactive information corresponding to the digital human is generated, and the digital human is controlled to perform dynamic interaction with the user according to the interactive information. In the voice interaction solution provided by this application, the emotional state corresponding to the user can be determined according to the voice emotion and the face expression, and the interactive information corresponding to the digital human can be generated based on the protocol parsing message corresponding to the emotional state, which can ensure that the interactive information of the digital human is associated with the emotional state of the user, so as to ensure that in subsequent interactions, the digital human performs dynamic interaction with the user based on this interactive information. Therefore, the problem that the action trigger of the digital human is not flexible enough under complex emotional changes can be solved, thereby improving the interaction ability.
[0123] To facilitate better implementation of the voice interaction method of this application, this application also provides a voice interaction device based on the above. The meanings of the nouns are the same as those in the above voice interaction method, and the specific implementation details can refer to the description in the method embodiments.
[0124] Please refer to Figure 7 , Figure 7 which is the structural schematic diagram of the voice interaction device provided by this application. The voice interaction device may include a collection module 201, an acquisition module 202, an identification module 203, and an interaction module 204, specifically as follows:
[0125] The collection module 201 is used to respond to the user's voice trigger operation and collect the user's multi-modal information;
[0126] An acquisition module 202, configured to acquire a voice feature corresponding to voice data in the multimodal information and a face feature corresponding to face data in the multimodal information;
[0127] An identification module 203, configured to determine an emotional state corresponding to a user according to the voice feature and the face feature;
[0128] An interaction module 204, configured to parse a message based on a protocol corresponding to the emotional state, generate interaction information corresponding to the digital human, and control the digital human to perform dynamic interaction with the user according to the interaction information.
[0129] Optionally, in some embodiments of the present application, the identification module 203 may specifically include:
[0130] A first identification unit, configured to process the voice feature based on a preset voice model to identify the voice emotion corresponding to the user;
[0131] A second identification unit, configured to identify the face emotion corresponding to the user based on the face feature and a preset reference face feature.
[0132] Optionally, in some embodiments of the present application, the first identification unit is specifically configured to:
[0133] Match the face feature with a preset reference face feature;
[0134] Determine the matched reference face feature as the target face feature, and obtain the face expression corresponding to the target face feature;
[0135] Determine the obtained face expression as the face expression corresponding to the user.
[0136] Optionally, in some embodiments of the present application, a training unit is further included, and the training unit is specifically configured to:
[0137] Acquire sample data corresponding to multiple sample users;
[0138] Label the sample data, and convert the labeled sample data into data in a preset format;
[0139] Divide the sample data after format conversion into a training set and a validation set;
[0140] Train a preset basic model using the training set, and verify the trained basic model using the validation set to obtain an expression detection model for detecting face expressions.
[0141] The above completes the voice interaction process of the present application.
[0142] As described above, the present application provides a voice interaction device. After the acquisition module 201 responds to the voice trigger operation of the user and acquires the multimodal information of the user, the acquisition module 202 obtains the voice features corresponding to the voice data in the multimodal information and the face features corresponding to the face data in the multimodal information. Then, the recognition module 203 recognizes the emotional state corresponding to the user based on the voice emotion corresponding to the voice features and the face expression corresponding to the face features. Finally, the interaction module 204 generates the interaction information corresponding to the digital human based on the protocol parsing message corresponding to the emotional state, and controls the digital human to perform dynamic interaction with the user according to the interaction information. In the voice interaction solution provided by the present application, the emotional state corresponding to the user can be determined according to the voice emotion and the face expression, and the interaction information corresponding to the digital human can be generated based on the protocol parsing message corresponding to the emotional state, which can ensure that the interaction information of the digital human is associated with the emotional state of the user, so as to ensure that in the subsequent interaction, the digital human performs dynamic interaction with the user based on the interaction information. Therefore, the problem that the action trigger of the digital human is not flexible enough under complex emotional changes can be solved, thereby improving the interaction ability.
[0143] Those of ordinary skill in the art can understand that all or part of the steps in the above methods of the embodiments can be completed by instructions, or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0144] The embodiment of the present invention also provides an intelligent terminal 500, as Figure 8 shown. The intelligent terminal 500 can integrate the above voice interaction device, and can further include a radio frequency (RF) circuit 501, a memory 502 including one or more computer-readable storage media, an input unit 503, a display unit 504, a sensor 505, an audio circuit 506, a wireless fidelity (WiFi) module 507, a processor 508 including one or more processing cores, and a power supply 509 and other components. Those skilled in the art can understand that Figure 8 the structure of the intelligent terminal 500 shown in
[0145] The RF circuit 501 can be used for receiving and transmitting information or signals during a call. Specifically, after receiving the downlink information from the base station, it is handed over to one or more processors 508 for processing. Additionally, data related to the uplink is sent to the base station. Generally, the RF circuit 501 includes, but is not limited to, an antenna, at least one amplifier, a tuner, one or more oscillators, a Subscriber Identity Module (SIM) card, a transceiver, a coupler, a Low Noise Amplifier (LNA), a duplexer, etc. In addition, the RF circuit 501 can also communicate with the network and other devices via wireless communication. The wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0146] The memory 502 can be used to store software programs and modules. The processor 508 executes various functional applications and information processing by running the software programs and modules stored in the memory 502. The memory 502 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, applications required for at least one function (such as a voice playback function, a target data playback function, etc.), etc.; the data storage area can store data created according to the use of the smart terminal 500 (such as audio data, a phone book, etc.). In addition, the memory 502 can include a high-speed random access memory and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory 502 can also include a memory controller to provide access to the memory 502 for the processor 508 and the input unit 503.
[0147] The input unit 503 can be used to receive input numerical or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control. Specifically, in a specific embodiment, the input unit 503 may include a touch-sensitive surface and other input devices. The touch-sensitive surface, also known as a touch display screen or a touchpad, can collect touch operations of a user on or near it (such as operations of the user using any suitable object or accessory such as a finger or a stylus on or near the touch-sensitive surface), and drive corresponding connection devices according to a preset program. Optionally, the touch-sensitive surface may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 508, and can receive and execute commands sent by the processor 508. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch-sensitive surface. In addition to the touch-sensitive surface, the input unit 503 may further include other input devices. Specifically, the other input devices may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power switch keys, etc.), trackballs, mice, joysticks, etc.
[0148] The display unit 504 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the smart terminal 500. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 504 may include a display panel. Optionally, the display panel may be configured in forms such as a liquid crystal display (LCD) or an organic light-emitting diode (OLED). Further, the touch-sensitive surface may cover the display panel. When the touch-sensitive surface detects a touch operation on or near it, it is transmitted to the processor 508 to determine the type of touch event. Subsequently, the processor 508 provides a corresponding visual output on the display panel according to the type of touch event. Although in Figure 4 the touch-sensitive surface and the display panel are implemented as two independent components to achieve input and input functions, in some embodiments, the touch-sensitive surface and the display panel can be integrated to achieve input and output functions.
[0149] The intelligent terminal 500 may further include at least one sensor 505, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel according to the brightness of the ambient light, and the proximity sensor can turn off the display panel and / or the backlight when the intelligent terminal 500 is moved to the ear. As a kind of motion sensor, the gravity acceleration sensor can detect the magnitude of the acceleration in each direction (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used in applications for identifying the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors that the intelligent terminal 500 can also be configured with, they will not be elaborated here.
[0150] The audio circuit 506, the speaker, and the microphone can provide an audio interface between the user and the intelligent terminal 500. The audio circuit 506 can transmit the electrical signal converted from the received audio data to the speaker, and the speaker converts it into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 506 and then converted into audio data. After the audio data is output to the processor 508 for processing, it is sent through the RF circuit 501 to, for example, another intelligent terminal 500, or the audio data is output to the memory 502 for further processing. The audio circuit 506 may also include an earphone jack to provide communication between the peripheral earphone and the intelligent terminal 500.
[0151] WiFi belongs to short-range wireless transmission technology. The intelligent terminal 500 can help users send and receive emails, browse the web, and access streaming media through the WiFi module 507, which provides users with wireless broadband Internet access. Although Figure 4 the WiFi module 507 is shown, it can be understood that it does not belong to the essential components of the intelligent terminal 500 and can be omitted entirely within the scope of not changing the essence of the invention according to needs.
[0152] The processor 508 is the control center of the intelligent terminal 500, connecting various parts of the entire mobile phone through various interfaces and lines. By running or executing the software programs and / or modules stored in the memory 502, and by calling the data stored in the memory 502, it executes various functions of the intelligent terminal 500 and processes data, thereby monitoring the mobile phone as a whole. Optionally, the processor 508 may include one or more processing cores; preferably, the processor 508 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, the user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 508 either.
[0153] The intelligent terminal 500 further includes a power supply 509 (such as a battery) for supplying power to each component. Preferably, the power supply can be logically connected to the processor 508 through a power management system, so as to manage functions such as charging, discharging, and power consumption management through the power management system. The power supply 509 may further include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power data indicator.
[0154] Although not shown, the intelligent terminal 500 may further include a camera, a Bluetooth module, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 508 in the intelligent terminal 500 will load the executable files corresponding to the processes of one or more application programs into the memory 502 according to the following instructions, and the processor 508 will run the application programs stored in the memory 502 to implement various functions:
[0155] In response to the user's voice trigger operation, collect the user's multi-modal information; obtain the voice features corresponding to the voice data in the multi-modal information and the face features corresponding to the face data in the multi-modal information; based on the voice emotion corresponding to the voice features and the facial expression corresponding to the face features, identify the emotional state corresponding to the user; based on the protocol parsing message corresponding to the emotional state, generate the interaction information corresponding to the digital human, and control the digital human to perform dynamic interaction with the user according to the interaction information.
[0156] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not elaborated in a certain embodiment, reference can be made to the detailed description of the above voice interaction method, which will not be elaborated here.
[0157] As can be seen from the above, the intelligent terminal 500 in the embodiment of the present invention, after responding to the user's voice trigger operation and collecting the user's multi-modal information, obtains the voice features corresponding to the voice data in the multi-modal information and the face features corresponding to the face data in the multi-modal information. Then, according to the voice emotion corresponding to the voice features and the facial expression of the face features, the emotional state corresponding to the user is identified. Finally, based on the protocol parsing message corresponding to the emotional state, the interaction information corresponding to the digital human is generated, and the digital human is controlled to perform dynamic interaction with the user according to the interaction information. In the voice interaction solution provided by the present application, the emotional state corresponding to the user can be determined according to the voice emotion and the facial expression, and the interaction information corresponding to the digital human can be generated based on the protocol parsing message corresponding to the emotional state, which can ensure that the interaction information of the digital human is associated with the emotional state of the user, so as to ensure that in the subsequent interaction, the digital human performs dynamic interaction with the user based on the interaction information. Therefore, the problem that the action trigger of the digital human is not flexible enough under complex emotional changes can be solved, thereby improving the interaction ability.
[0158] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling related hardware through instructions. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0159] Therefore, an embodiment of the present application further provides a storage medium, on which multiple instructions are stored. The instructions are suitable for being loaded by a processor to execute the steps in the above voice interaction method.
[0160] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated here.
[0161] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0162] Since the instructions stored in the storage medium can execute the steps in any of the voice interaction methods provided by the embodiments of the present invention, the beneficial effects that can be achieved by any of the voice interaction methods provided by the embodiments of the present invention can be realized. For details, refer to the previous embodiments and will not be elaborated here.
[0163] The above has introduced in detail the voice interaction method, device, system and storage medium provided by the embodiments of the present invention. Specific examples are used herein to elaborate the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.< / noitn> < / itn> < / noitn> < / itn> < / aec> < / aec> < / lid> < / lid> < / lid> < / itn> < / ser> < / ser> < / lid> < / lid> < / lid>
Claims
1. A voice interaction method, characterized in that: include: In response to a voice trigger operation of a user, collecting multimodal information of the user; Acquire voice features corresponding to the voice data in the multimodal information and facial features corresponding to the facial data in the multimodal information; Based on the voice emotion corresponding to the voice feature and the facial expression corresponding to the facial feature, identifying the emotional state corresponding to the user; Based on the protocol parsing message corresponding to the emotional state, interactive information corresponding to the digital human is generated, and according to the interactive information, the digital human is controlled to dynamically interact with the user.
2. The voice interaction method according to claim 1, characterized in that: The identifying the emotional state of the user based on the voice emotion corresponding to the voice feature and the facial expression corresponding to the facial feature includes: Performing sentiment analysis on the speech feature based on a preset speech model to obtain the speech sentiment corresponding to the user; Based on the facial features and preset reference facial features, identifying the facial expression corresponding to the user; The voice emotion and facial expression are used as emotion prompt words, and the emotion prompt words are input into a preset large language model to obtain the corresponding emotion state of the user.
3. The voice interaction method according to claim 2, characterized in that: The identifying the facial expression corresponding to the user based on the facial features and the preset reference facial features includes: Matching the facial features with preset reference facial features; Determine the matched reference facial feature as the target facial feature, and obtain the facial expression corresponding to the target facial feature; Determine that the acquired facial expression is the facial expression corresponding to the user.
4. The voice interaction method according to claim 2, characterized in that: Also includes: Obtain sample data corresponding to multiple sample users; Annotating the sample data, and converting the annotated sample data into data in a preset format; Divide the sample data after format conversion into a training set and a validation set; The training set is used to train a preset basic model, and the verification set is used to verify the trained basic model to obtain an expression detection model for detecting facial expressions.
5. The voice interaction method according to claim 1, characterized in that: The protocol parsing message corresponding to the emotional state, generating interactive information corresponding to the digital human, and controlling the digital human to dynamically interact with the user according to the interactive information, includes: converting the affective state into standardized protocol information; Parsing the protocol information; Generate interactive information corresponding to the digital human based on the preset large language model and the analysis result, and control the digital human to dynamically interact with the user according to the interactive information.
6. The voice interaction method according to claim 5, characterized in that: The generating of interactive information corresponding to the digital human based on the preset large language model and the analysis result, and controlling the dynamic interaction between the digital human and the user according to the interactive information, includes: Generate interactive actions and interactive texts corresponding to the digital human based on the preset large language model and analysis results; Generating interactive speech corresponding to the interactive text; According to the interactive voice and interactive action, the digital human is controlled to dynamically interact with the user.
7. The voice interaction method according to claim 6, characterized in that: The generating of the interactive speech corresponding to the interactive text includes: Converting the interactive text into text audio according to speech synthesis technology; The text audio is adjusted based on the emotional state of the user to obtain interactive speech corresponding to the interactive text.
8. The voice interaction method according to claim 6, characterized in that: The generation of interactive actions corresponding to the digital human based on the preset large language model and the analysis results includes: Based on the preset large language model and parsing results, the corresponding interactive actions of the digital human are obtained from the preset action sequence.
9. A voice interaction device, characterized in that: include: A collection module, used to respond to a user's voice trigger operation and collect the user's multimodal information; An acquisition module, used to acquire voice features corresponding to the voice data in the multimodal information and facial features corresponding to the facial data in the multimodal information; A recognition module, used to determine the corresponding emotional state of the user according to the voice features and facial features; The interaction module is used to parse the message based on the protocol corresponding to the emotional state, generate interaction information corresponding to the digital human, and control the dynamic interaction between the digital human and the user according to the interaction information.
10. An intelligent terminal, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the voice interaction method described in any one of claims 1 to 8 are implemented.
11. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the voice interaction method according to any one of claims 1 to 8 are implemented.