Voice interaction method, device, apparatus and storage medium
By collecting voice and image data to identify user emotions and performing emotional reassurance actions, the problem of voice assistants being unable to provide feedback has been solved, thus improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN TCL DIGITAL TECH CO LTD
- Filing Date
- 2024-12-09
- Publication Date
- 2026-04-14
AI Technical Summary
In existing voice interaction processes, voice assistants cannot perform feedback operations based on the user's emotions, resulting in a poor user experience.
By collecting interactive voice data and object image data, and combining voice emotion recognition and facial expression recognition, the system determines the user's emotion recognition data and executes soothing operations based on emotion soothing strategies to improve the user experience.
It reduces the probability of users experiencing anger during voice interaction, and improves users' voice interaction experience and satisfaction.
Smart Images

Figure CN119673160B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, specifically to a voice interaction method, apparatus, device, and storage medium. Background Technology
[0002] Currently, with the rapid development of smart TVs and other smart terminal technologies, as well as voice recognition technology, more and more smart TVs and other smart terminals are integrating voice assistants. This allows users to interact with these terminals via voice, thus optimizing the interaction scenarios. However, in existing voice interaction processes, the voice assistant may repeatedly misrecognize content or fail to find the resources the user needs, leading to emotional fluctuations in the user. Furthermore, when the user experiences emotional fluctuations, the voice assistant may be unable to provide corresponding feedback, causing even more intense emotional fluctuations and resulting in a poor user experience. Summary of the Invention
[0003] This application provides a voice interaction method, apparatus, device, and storage medium, aiming to solve the technical problem in the prior art that the voice interaction experience is poor because the user cannot perform corresponding feedback operations based on the user's emotions when interacting with a smart terminal.
[0004] On one hand, embodiments of this application provide a voice interaction method, which includes the following steps:
[0005] In response to a voice soothing request triggered by the target object, the system collects the interactive voice data and object image data of the target object.
[0006] The emotion data determines the emotion recognition data of the target object based on the interactive voice features corresponding to the interactive voice data and the image emotion data corresponding to the object image data.
[0007] The target soothing operation is performed based on the emotion recognition data and the corresponding emotion soothing strategy to obtain the emotion soothing result.
[0008] In one possible implementation of this application, the emotion data determines the emotion recognition data of the target object based on the interactive voice features corresponding to the interactive voice data and the image emotion data corresponding to the object image data, including:
[0009] Perform emotion classification on the interactive voice features corresponding to the interactive voice data to obtain the voice emotion data corresponding to the interactive voice features;
[0010] Perform facial expression recognition on the object image features corresponding to the object image data to obtain image emotion data corresponding to the object image data;
[0011] The target object's emotion recognition data is obtained by matching the voice emotion data and the image emotion data.
[0012] In one possible implementation of this application, the emotion data determines the emotion recognition data of the target object based on the interactive voice features corresponding to the interactive voice data and the image emotion data corresponding to the object image data, including:
[0013] Perform emotion classification on the interactive voice features corresponding to the interactive voice data to obtain the voice emotion data corresponding to the interactive voice features;
[0014] Perform facial expression recognition on the object image features corresponding to the object image data to obtain image emotion data corresponding to the object image data;
[0015] The target object's emotion recognition data is obtained by matching the voice emotion data and the image emotion data.
[0016] In one possible implementation of this application, the step of performing emotion classification on the interactive voice features corresponding to the interactive voice data to obtain voice emotion data corresponding to the interactive voice features includes:
[0017] The interactive voice data is subjected to voice feature extraction to obtain the interactive voice features of the interactive voice data.
[0018] The interactive speech features are initially classified into emotions to obtain the initial emotion features of the interactive speech features;
[0019] If the initial emotional feature is a negative emotional feature, then the initial emotional feature is subjected to secondary emotional classification according to the preset speech recognition model to obtain the speech emotion data corresponding to the interactive speech feature.
[0020] In one possible implementation of this application, facial expression recognition is performed on the object image features corresponding to the object image data to obtain image emotion data corresponding to the object image data, including:
[0021] Facial features are extracted from the object image data to obtain the object image features corresponding to the object image data;
[0022] The target convolution module, stochastic gradient descent module, and target feature extractor are used to perform facial expression recognition on the object image features to obtain the image emotion data corresponding to the object image data.
[0023] In one possible implementation of this application, the step of matching the voice emotion data and the image emotion data to obtain the emotion recognition data of the target object includes:
[0024] Obtain the voice emotion level corresponding to the voice emotion data, and obtain the image emotion level corresponding to the image emotion data;
[0025] If the difference between the voice emotion level and the image emotion level is less than or equal to a preset first difference threshold, then the voice emotion level or the image emotion level with the larger level value is used as the emotion recognition data of the target object.
[0026] If the difference between the voice emotion level and the image emotion level is greater than or equal to a preset second difference threshold, then the voice recognition model is used to perform secondary recognition on the interactive voice data to obtain the emotion recognition data of the target object.
[0027] In one possible implementation of this application, the step of using a speech recognition model to perform secondary recognition on the interactive speech data and the interactive speech data to obtain the emotion recognition data of the target object includes:
[0028] The interactive voice data and the interactive voice data are re-identified using a speech recognition model to obtain an updated emotion level, which includes an updated voice emotion level and / or an updated image emotion level.
[0029] The average value of the voice emotion level, image emotion level, updated voice emotion level, and / or updated image emotion level is calculated to obtain the emotion recognition data of the target object.
[0030] In one possible implementation of this application, the step of performing a target soothing operation based on the emotion recognition data and the emotion soothing strategy corresponding to the emotion recognition data to obtain an emotion soothing result includes:
[0031] Obtain the target emotion level from the emotion recognition data, and generate an emotion soothing test strategy corresponding to the target emotion level;
[0032] Based on the stated emotion-soothing strategy, corresponding emotion-soothing content is output, and the emotion-soothing result is obtained.
[0033] In one possible implementation of this application, the step of outputting corresponding emotional soothing content based on the emotional soothing strategy to obtain an emotional soothing result includes:
[0034] Access the cloud database to obtain the emotion soothing strategy and the emotion soothing content corresponding to the interactive voice data;
[0035] The emotional comforting content is output to the target object, and the target object's emotional update data is collected;
[0036] By comparing the emotion update data and the emotion recognition data, the emotion soothing result of the target object is obtained.
[0037] On the other hand, this application provides a voice interaction device, the voice interaction device comprising:
[0038] The data acquisition module is configured to respond to a voice reassurance request triggered by the target object and acquire the interactive voice data and object image data of the target object;
[0039] The emotion recognition module is configured to determine the emotion recognition data of the target object based on the interactive voice features corresponding to the interactive voice data and the image emotion data corresponding to the object image data.
[0040] The voice interaction module is configured to perform a target soothing operation based on the emotion recognition data and the corresponding emotion soothing strategy, and obtain an emotion soothing result.
[0041] On the other hand, this application also provides a voice interaction device, the voice interaction device comprising:
[0042] One or more processors;
[0043] Memory; and
[0044] One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the steps of the voice interaction method.
[0045] On the other hand, this application also provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to execute the steps in the voice interaction method.
[0046] This application collects interactive voice data and object image data of the target object in response to a voice soothing request triggered by the target object. Emotional data is used to determine the target object's emotion recognition data based on the interactive voice features corresponding to the interactive voice data and the image emotion data corresponding to the object image data. A target soothing operation is then performed based on the emotion recognition data and the corresponding emotion soothing strategy to obtain the emotion soothing result. This approach achieves the goal of acquiring the target object's interactive voice data and object image data during voice interaction, performing multi-dimensional emotion recognition based on the interactive voice data and object image data, thereby determining the target object's emotion recognition data during voice interaction, and performing corresponding soothing operations based on the emotion soothing strategy corresponding to the emotion mouse data. This reduces the probability of users experiencing anger during voice interaction, improving the user's voice interaction experience and satisfaction. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1 This is a schematic diagram illustrating a scenario of the voice interaction method according to an embodiment of this application;
[0049] Figure 2 This is a flowchart illustrating one embodiment of the voice interaction method in this application.
[0050] Figure 3 A flowchart illustrating an embodiment of image emotion data recognition in the voice interaction method provided in this application;
[0051] Figure 4 A flowchart illustrating an embodiment of the voice interaction method provided in this application for determining the emotion recognition data of a target object through emotion matching;
[0052] Figure 5 A schematic diagram of the structure of one embodiment of the voice interaction device provided in this application;
[0053] Figure 6 This is a schematic diagram of the structure of one embodiment of the voice interaction device provided in this application. Detailed Implementation
[0054] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0055] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0056] In this application, the term "exemplary" is used to mean "serving as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0057] Currently, with the rapid development of smart TVs and other smart terminal technologies, as well as voice recognition technology, more and more smart TVs and other smart terminals are integrating voice assistants. This allows users to interact with these terminals via voice, thus optimizing the interaction scenarios. However, in existing voice interaction processes, the voice assistant may repeatedly misrecognize content or fail to find the resources the user needs, leading to emotional fluctuations in the user. Furthermore, when the user experiences emotional fluctuations, the voice assistant may be unable to provide corresponding feedback, causing even more intense emotional fluctuations and resulting in a poor user experience.
[0058] Based on this, this application proposes a voice interaction method, apparatus, device, and computer-readable storage medium to solve the technical problem in the prior art that the voice interaction experience is poor because the user cannot perform corresponding feedback operations based on the user's emotions when interacting with a smart terminal.
[0059] The voice interaction method in this embodiment of the invention is applied to a voice interaction device, which is set in a voice interaction equipment. The voice interaction equipment is provided with one or more processors, a memory, and one or more applications, wherein one or more applications are stored in the memory and configured to be executed by the processor to implement the voice interaction method. The voice interaction equipment can be a smart terminal, such as a mobile phone, tablet computer, network device, and smart computer. Optionally, the voice interaction equipment can also be a server or a service cluster composed of multiple servers.
[0060] like Figure 1 As shown, Figure 1 This is a schematic diagram of a scenario for a voice interaction method according to an embodiment of this application. In this embodiment of the invention, the voice interaction scenario includes a voice interaction device 100 (the voice interaction device 100 integrates a voice interaction apparatus), and the voice interaction device 100 runs a computer-readable storage medium corresponding to the voice interaction method to execute the steps of the voice interaction method.
[0061] Understandable, Figure 1 The voice interaction device in the scenario of the voice interaction method shown, or the device included in the voice interaction device, does not constitute a limitation on the embodiments of the present invention. That is, the number or type of the voice interaction device included in the scenario of the voice interaction method, or the number or type of the device included in each device, does not affect the overall implementation of the technical solution in the embodiments of the present invention, and can all be considered as equivalent substitutions or derivatives of the technical solutions claimed in the embodiments of the present invention.
[0062] In this embodiment of the invention, the voice interaction device 100 is mainly used to: respond to a voice soothing request triggered by a target object, and collect the interactive voice data and object image data of the target object;
[0063] The emotion data determines the emotion recognition data of the target object based on the interactive voice features corresponding to the interactive voice data and the image emotion data corresponding to the object image data.
[0064] The target soothing operation is performed based on the emotion recognition data and the corresponding emotion soothing strategy to obtain the emotion soothing result.
[0065] The voice interaction device 100 in this embodiment of the invention can be an independent voice interaction device, such as a mobile phone, tablet computer, network device, server and smart computer, or a voice interaction network or voice interaction cluster composed of multiple voice interaction devices.
[0066] This application provides a voice interaction method, apparatus, device, and computer-readable storage medium, which will be described in detail below.
[0067] It will be understood by those skilled in the art that Figure 1 The application environment shown is only one application scenario related to the solution of this application and does not constitute a limitation on the application scenario of this application. Other application environments may include more than one application scenario. Figure 1 The number of more or fewer voice interaction devices shown, or the voice interaction network connections, for example Figure 1 Only one voice interaction device is shown in the figure. It can be understood that the scenario of this voice interaction method may also include one or more voice interaction devices, which are not limited here. The voice interaction device 100 may also include a memory for storing interaction data and other data.
[0068] It should be noted that, Figure 1 The schematic diagram of the voice interaction method shown is merely an example. The scenarios of the voice interaction method described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention and do not constitute a limitation on the technical solutions provided in the embodiments of the present invention.
[0069] Based on the scenarios described above for voice interaction methods, various embodiments of the voice interaction method disclosed in this invention are proposed.
[0070] like Figure 2 As shown, Figure 2 This is a flowchart illustrating one embodiment of the voice interaction method in this application. The voice interaction method includes the following steps 201 to 203:
[0071] 201. In response to a voice soothing request triggered by the target object, collect the interactive voice data and object image data of the target object;
[0072] The voice interaction method in this embodiment is applied to a voice interaction device. The type and number of voice interaction devices are not specifically limited; that is, the voice interaction device can be one or more smart terminals that include a smart voice assistant and are capable of voice interaction with the user. In one specific embodiment, the voice interaction device is a smart TV. Optionally, in other embodiments, the voice interaction device can be a mobile phone, smart computer, tablet computer, or other smart device.
[0073] The intelligent voice assistant can be a voice assistant application built on technologies such as artificial intelligence (e.g., generative large models), capable of generating corresponding interactive content based on the interactive voice data of the target object, and outputting the interactive content using a voice output module or a display output module. That is, the voice interaction device can converse with the user through the intelligent voice assistant and execute corresponding interactive operations through the user's voice dialogue. For example, the user can wake up the intelligent voice assistant and ask it a question, thereby obtaining answers generated by the intelligent voice assistant (including text, recommended information, and image information, etc.). Optionally, the user can also wake up the intelligent voice assistant and send interactive commands to the intelligent voice assistant, driving the intelligent voice assistant to execute the interactive commands. Optionally, the user can also perform other interactions with the intelligent voice assistant that are not mentioned, which are not specifically limited in this embodiment.
[0074] Specifically, during operation, the voice interaction device can respond to voice soothing requests triggered by a target object. This voice soothing request is an event that drives the voice interaction device to detect and identify the target object's emotion data in the voice interaction scenario triggered by the target object, determine the target object's emotional state based on this emotion data, and perform emotional soothing operations when the target object is in a depressed or angry emotional state. The triggering method of this voice soothing request is not specifically limited here; that is, the voice soothing request can be actively triggered by the user. For example, the user can actively trigger the voice soothing request by clicking the voice soothing button on the smart voice assistant settings page of the voice interaction device. Optionally, in other embodiments, the voice soothing request can also be automatically triggered by the voice interaction device. For example, the voice interaction device is set up with an emotion soothing process that automatically triggers the voice soothing request when it detects that the user has started the smart voice assistant. The target object is the user who is performing voice interaction operations with the smart voice assistant in the voice interaction device.
[0075] Specifically, after receiving the voice-based soothing request, the voice interaction device collects the target object's interactive voice data and object image data when the target object uses the intelligent voice assistant to perform voice interaction. The interactive voice data consists of the voice data of the target object during its voice interaction with the intelligent voice assistant. The object image data consists of the target object's facial image data during its voice interaction with the intelligent voice assistant.
[0076] Specifically, the voice interaction device includes a voice acquisition module, an image acquisition module, and a computer vision module. When the target user interacts with the intelligent voice assistant using voice, the voice acquisition module collects the target user's real-time interactive voice data. Furthermore, when the target user interacts with the intelligent voice assistant using voice, the image acquisition module and the computer vision module also collect the target user's image data.
[0077] 202. The emotion data determines the emotion recognition data of the target object based on the interactive voice features corresponding to the interactive voice data and the image emotion data corresponding to the object image data;
[0078] Specifically, after acquiring the interactive voice data and object image data corresponding to the target object, the voice interaction device also uses the interactive voice data and interactive voice data to analyze the emotional state of the target object when using the smart voice assistant for voice interaction, and obtains emotion recognition data. That is, the voice interaction device determines the emotion recognition data of the target object based on the voice emotion data corresponding to the interactive voice data and the image emotion data corresponding to the object image data.
[0079] Specifically, after acquiring the interactive voice data and the object image data, the voice interaction device performs emotion classification on the interactive voice features corresponding to the interactive voice data, thereby determining the voice emotion data corresponding to the interactive voice features. That is, the voice interaction device is pre-set with a voice emotion recognition model capable of performing emotion classification on the interactive voice data, and after acquiring the interactive voice data, the interactive voice data is input into the voice emotion recognition model for emotion classification to obtain the voice emotion data corresponding to the interactive voice data.
[0080] Specifically, after the voice interaction device inputs the interactive voice data into the voice emotion recognition model, it uses the model to extract voice features from the data. Specifically, the voice encoding module within the model encodes the voice features of the interactive voice data to obtain its interactive voice features. These interactive voice features can include speech characteristic information such as the spectrogram, Mel-frequency cepstral coefficients (MFCC), and fundamental frequency (pitch) corresponding to the interactive voice data.
[0081] After acquiring the interactive voice features, the voice interaction device also uses the initial emotion classification module in the voice emotion recognition model to perform initial emotion classification on the interactive voice features, thereby determining the emotion category of the interactive voice features and obtaining the initial emotion features of the interactive voice features. These initial emotion features can be divided into positive emotion features and negative emotion features.
[0082] Specifically, after the voice interaction device uses the initial emotion classification module to determine the initial emotion feature corresponding to the interactive voice feature, it also determines whether the initial emotion feature is a negative emotion feature, and then determines whether further emotion recognition and corresponding soothing operations are needed for the initial emotion feature.
[0083] Optionally, if the voice interaction device determines that the initial emotional feature is a negative emotional feature, the voice interaction device performs a secondary emotional classification on the initial emotional feature according to the speech recognition model, thereby obtaining the voice emotion data corresponding to the interactive voice feature. Here, the voice emotion data is a quantitative parameter representing the degree of negative emotional state of the target object when interacting with the intelligent voice assistant. Optionally, in a specific embodiment, the voice emotion data includes a voice emotion level. Here, the voice emotion level is a quantitative parameter representing the negative emotional state and degree of negative emotion of the target object obtained through the voice data.
[0084] Specifically, after determining that the initial emotional feature is a negative emotional feature, the voice interaction device further utilizes the speech convolutional module and recurrent neural network module in the speech recognition model to perform secondary emotional classification on the initial emotional feature, thereby obtaining the speech emotion data corresponding to the interactive speech feature. The speech convolutional module can be a trained CNN model, and the recurrent neural network module can be an LSTM model.
[0085] Specifically, the voice interaction device can also perform facial expression recognition on the object image features corresponding to the object image data, thereby determining the user's facial expression state when the target object interacts with the intelligent voice assistant, and obtaining image emotion data corresponding to the object image data, which represents the user's emotional state of the target object. This image emotion data includes an image emotion level, which is a quantitative parameter representing the negative emotional state and degree of negative emotion of the target object obtained through facial image data.
[0086] Specifically, after acquiring the voice emotion data and image emotion data, the voice interaction device performs multi-dimensional emotion matching based on the voice emotion data and image emotion data to determine the emotion recognition data of the target object, thereby improving the accuracy of emotion recognition of the target object.
[0087] The emotion recognition data consists of quantitative parameters characterizing the negative emotional state and degree of the target object. Optionally, in one specific embodiment, the emotion recognition data can be a specific emotion type and a target emotion level corresponding to that specific emotion type. For example, the emotion recognition data can be anger level 1, anger level 2, ..., anger level n, etc. Furthermore, the emotion recognition data can also be other emotion types and their corresponding target emotion levels.
[0088] 203. Perform the target soothing operation based on the emotion recognition data and the corresponding emotion soothing strategy to obtain the emotion soothing result.
[0089] Specifically, after acquiring the emotion recognition data of the target object when using the intelligent voice assistant for voice interaction, the voice interaction device also performs the target soothing operation based on the emotion recognition data and the corresponding emotion soothing strategy to obtain the emotion soothing result.
[0090] Specifically, when a user interacts with a smart voice assistant but the assistant fails to provide satisfactory content, leading to anger or other emotional states, the voice interaction device, upon recognizing this emotion data, analyzes it to obtain the target emotion level. It then generates a corresponding emotion-soothing strategy and outputs soothing content based on this strategy, resulting in an emotional reassurance outcome. The emotion-soothing level is the interaction strategy that drives the smart voice assistant to change the corresponding interaction content or perform a specified soothing interaction based on the target emotion level.
[0091] Specifically, after acquiring the emotion-soothing strategy, the voice interaction device accesses the cloud database to retrieve the emotion-soothing content corresponding to the strategy and the interactive voice data. That is, the voice interaction device can retrieve the emotion-soothing content stored in the cloud database for the same voice interaction scenario as the emotion-soothing strategy and interactive voice data. This emotion-soothing content can be either replacing the interactive voice data with the correct interactive content or using a question tone to ask whether to search for specific content.
[0092] Specifically, after acquiring emotional comforting content, the voice interaction device can output the emotional comforting content to the target object through the voice output module or the display output module, and identify the target object's emotional update data. The emotional update data is real-time emotional recognition data that represents the target object's updated emotion after perceiving the emotional comforting content.
[0093] Optionally, after acquiring the emotion update data, the voice interaction device will also compare the emotion update data with the emotion recognition data to determine whether the emotional state of the target object has improved. If the emotional state of the target object has improved, the device will output the emotion soothing result of successful emotion soothing and record the emotion soothing content.
[0094] In this embodiment, the voice interaction device collects the interactive voice data and image data of the target object in response to a voice soothing request triggered by the target object. Emotional data is used to determine the target object's emotion recognition data based on the interactive voice features corresponding to the interactive voice data and the image emotion data corresponding to the object image data. A target soothing operation is then performed based on the emotion recognition data and the corresponding emotion soothing strategy to obtain the emotion soothing result. This approach achieves the goal of acquiring the target object's interactive voice data and object image data during voice interaction, performing multi-dimensional emotion recognition based on these data, determining the target object's emotion recognition data during voice interaction, and executing corresponding soothing operations based on the emotion soothing strategy corresponding to the emotion mouse data. This reduces the probability of users experiencing anger during voice interaction, improving the user's voice interaction experience and satisfaction.
[0095] like Figure 3 As shown, Figure 3 A flowchart illustrating an embodiment of image emotion data recognition in the voice interaction method provided in this application. Specifically, in this embodiment, the voice interaction method further includes steps 301 to 302:
[0096] 301. Extract facial features from the object image data to obtain the object image features corresponding to the object image data;
[0097] 302. Based on the target convolution module, stochastic gradient descent module and target feature extractor, facial expression recognition is performed on the object image features to obtain the image emotion data corresponding to the object image data.
[0098] Based on the above embodiments, in this embodiment, the voice interaction device can also perform facial expression recognition on the object image features corresponding to the object image data, thereby determining the user's facial expression state when the target object interacts with the intelligent voice assistant, so as to obtain image emotion data corresponding to the object image data that represents the user's emotional state of the target object.
[0099] Specifically, the voice interaction device pre-builds an image recognition model for facial expression recognition of object image data. It generates a training dataset by collecting facial image data containing different negative emotions under various factors such as different age groups and genders. The facial image data and voice training data of the same scene in the training dataset are divided into the same level. The facial image data is also divided into different age groups and corresponding training subsets are constructed. The image recognition model is then trained based on the training subsets and the training dataset to obtain the trained image recognition model.
[0100] Specifically, after acquiring the trained image recognition model, the voice interaction device also uses the image recognition model to extract facial features from the object image data, thereby obtaining the object image features corresponding to the object image data. That is, the voice interaction device uses the image recognition model to perform face detection on the object image data, and after detecting the face region image of the object image data, it extracts features from the face region image to obtain the object image features corresponding to the object image data.
[0101] Specifically, after acquiring the features of an object image, the voice interaction device also uses the image recognition model to perform facial expression recognition on the object image features, thereby obtaining the image emotion data corresponding to the object image data.
[0102] Specifically, the image recognition model also includes a target convolutional module, a stochastic gradient descent module, and a target feature extractor for facial expression recognition based on image features. After acquiring the object image features, the image recognition model inputs these features into the target convolutional module for convolution processing, then inputs the convolutionally processed object image features into the stochastic gradient descent module for optimization, and finally adds them to the target feature extractor for facial expression recognition, thereby obtaining the image emotion data corresponding to the object image data. In one specific embodiment, the target convolutional module is a CNN deep convolutional model, the stochastic gradient descent module can be an SGD optimized model, and the target feature extractor can be a VGG model.
[0103] Specifically, after acquiring the voice emotion data and image emotion data, the voice interaction device performs multi-dimensional emotion matching based on the voice emotion data and image emotion data to determine the emotion recognition data of the target object, thereby improving the accuracy of emotion recognition of the target object.
[0104] In this embodiment, the voice interaction device extracts facial features from the object image data to obtain object image features corresponding to the object image data; then, based on the object image features obtained by the target convolution module, stochastic gradient descent module, and target feature extractor, it performs facial expression recognition to obtain image emotion data corresponding to the object image data. This enables the identification of the target object's emotional state when using a smart voice assistant for voice interaction based on the target object's facial expression features, providing data support for subsequent emotional reassurance.
[0105] like Figure 4 As shown, Figure 4 A flowchart illustrating an embodiment of the voice interaction method provided in this application for determining the emotion recognition data of a target object through emotion matching. Specifically, in this embodiment, the voice interaction method includes steps 401 to 403:
[0106] 401. Obtain the voice emotion level corresponding to the voice emotion data, and obtain the image emotion level corresponding to the image emotion data;
[0107] 402. If the difference between the voice emotion level and the image emotion level is less than or equal to a preset first difference threshold, then the voice emotion level or the image emotion level with the larger level value is taken as the emotion recognition data of the target object.
[0108] 403. If the difference between the voice emotion level and the image emotion level is greater than or equal to a preset second difference threshold, then the voice recognition model is used to perform secondary recognition on the interactive voice data and the interactive voice data to obtain the emotion recognition data of the target object.
[0109] Based on the above embodiments, in this embodiment, after acquiring the voice emotion data and image emotion data, the voice interaction device also performs multi-dimensional emotion matching based on the voice emotion data and image emotion data to determine the emotion recognition data of the target object, thereby improving the accuracy of emotion recognition of the target object.
[0110] Specifically, after acquiring the voice emotion data and image emotion data, the voice interaction device obtains the voice emotion level corresponding to the voice emotion data and the image emotion level corresponding to the image emotion data. By comparing the voice emotion level and the image emotion level, it determines the emotion recognition data of the target object. The voice emotion level is a quantitative parameter representing the negative emotional state and degree of the target object obtained through the voice data. The image emotion level is a quantitative parameter representing the negative emotional state and degree of the target object obtained through the facial image data.
[0111] Specifically, the voice interaction device calculates the level difference between the voice emotion level and the image emotion level, and compares the level difference with a preset first difference threshold and a second difference threshold to determine whether the error between the voice emotion recognition and the image emotion recognition is too large, and then determines the emotion recognition data of the target object based on the comparison result.
[0112] Optionally, if the difference between the voice emotion level and the image emotion level is less than or equal to a preset first difference threshold, the voice interaction device determines that there is no error in the voice emotion recognition and image emotion recognition, and then uses the voice emotion level or image emotion level with the larger emotion level value as the emotion recognition data of the target object. In one specific embodiment, the first difference threshold is 1.
[0113] That is, when the difference between the voice emotion level and the image emotion level is less than or equal to a preset first difference threshold, the voice interaction device compares the voice emotion level and the image emotion level. If the voice emotion level is greater than the image emotion level, the voice emotion level is set as the emotion recognition data of the target object. If the voice emotion level is less than the image emotion level, the voice emotion level is set as the emotion recognition data of the target object.
[0114] Optionally, if the difference between the voice emotion level and the image emotion level is greater than or equal to a preset second difference threshold, the voice interaction device determines that there is a large error in the voice emotion recognition and image emotion recognition. Then, the voice interaction device uses the voice recognition model to perform secondary recognition on the interactive voice data and interactive voice data, thereby determining the emotion recognition data of the target object based on the recognition result, the voice emotion level, and the image emotion level.
[0115] Specifically, the voice interaction device uses a speech recognition model to perform secondary recognition on the interactive voice data and the interactive voice data to obtain an updated emotion level, which includes an updated voice emotion level and / or an updated image emotion level.
[0116] Specifically, after obtaining the updated emotion level, the voice interaction device calculates the voice emotion level, the image emotion level, and the average of the updated voice emotion level and / or the updated image emotion level in the updated emotion level to obtain the emotion recognition data of the target object.
[0117] In this embodiment, the voice interaction device acquires the voice emotion level corresponding to the voice emotion data and the image emotion level corresponding to the image emotion data. If the difference between the voice emotion level and the image emotion level is less than or equal to a preset first difference threshold, the voice emotion level or image emotion level with the larger value is used as the emotion recognition data of the target object. If the difference between the voice emotion level and the image emotion level is greater than or equal to a preset second difference threshold, a voice recognition model is used to perform secondary recognition on the interactive voice data to obtain the emotion recognition data of the target object. This achieves the determination of the target emotion recognition result of the target object based on the voice emotion recognition result and the image emotion recognition result, thereby reducing the error of the emotion recognition result.
[0118] To better implement the voice interaction method in the embodiments of this application, based on the voice interaction method, the embodiments of this application also provide a voice interaction device, such as... Figure 5 As shown, Figure 5 This is a schematic diagram of the structure of the voice interaction device provided in the embodiments of this application. Specifically, the voice interaction device 500 includes:
[0119] The data acquisition module 501 is configured to collect interactive voice data and object image data of the target object in response to a voice soothing request triggered by the target object.
[0120] The emotion recognition module 502 is configured to determine the emotion recognition data of the target object based on the interactive voice features corresponding to the interactive voice data and the image emotion data corresponding to the object image data.
[0121] The voice interaction module 503 is configured to perform a target soothing operation based on the emotion recognition data and the emotion soothing strategy corresponding to the emotion recognition data, and obtain an emotion soothing result.
[0122] In one possible implementation of this embodiment, the voice interaction device determines the emotion recognition data of the target object based on the interactive voice features corresponding to the interactive voice data and the image emotion data corresponding to the object image data, including:
[0123] Perform emotion classification on the interactive voice features corresponding to the interactive voice data to obtain the voice emotion data corresponding to the interactive voice features;
[0124] Perform facial expression recognition on the object image features corresponding to the object image data to obtain image emotion data corresponding to the object image data;
[0125] The target object's emotion recognition data is obtained by matching the voice emotion data and the image emotion data.
[0126] In one possible implementation of this embodiment, the voice interaction device performs emotion classification on the interactive voice features corresponding to the interactive voice data to obtain voice emotion data corresponding to the interactive voice features, including:
[0127] The interactive voice data is subjected to voice feature extraction to obtain the interactive voice features of the interactive voice data.
[0128] The interactive speech features are initially classified into emotions to obtain the initial emotion features of the interactive speech features;
[0129] If the initial emotional feature is a negative emotional feature, then the initial emotional feature is subjected to secondary emotional classification according to the preset speech recognition model to obtain the speech emotion data corresponding to the interactive speech feature.
[0130] In one possible implementation of this embodiment, the voice interaction device performs facial expression recognition on the object image features corresponding to the object image data to obtain image emotion data corresponding to the object image data, including:
[0131] Facial features are extracted from the object image data to obtain the object image features corresponding to the object image data;
[0132] The target convolution module, stochastic gradient descent module, and target feature extractor are used to perform facial expression recognition on the object image features to obtain the image emotion data corresponding to the object image data.
[0133] In one possible implementation of this embodiment, the voice interaction device matches the voice emotion data and the image emotion data to obtain the emotion recognition data of the target object, including:
[0134] Obtain the voice emotion level corresponding to the voice emotion data, and obtain the image emotion level corresponding to the image emotion data;
[0135] If the difference between the voice emotion level and the image emotion level is less than or equal to a preset first difference threshold, then the voice emotion level or the image emotion level with the larger level value is used as the emotion recognition data of the target object.
[0136] If the difference between the voice emotion level and the image emotion level is greater than or equal to a preset second difference threshold, then the voice recognition model is used to perform secondary recognition on the interactive voice data to obtain the emotion recognition data of the target object.
[0137] In one possible implementation of this embodiment, the voice interaction device uses a speech recognition model to perform secondary recognition on the interactive voice data and the interactive voice data to obtain the emotion recognition data of the target object, including:
[0138] The interactive voice data and the interactive voice data are re-identified using a speech recognition model to obtain an updated emotion level, which includes an updated voice emotion level and / or an updated image emotion level.
[0139] The average value of the voice emotion level, image emotion level, updated voice emotion level, and / or updated image emotion level is calculated to obtain the emotion recognition data of the target object.
[0140] In one possible implementation of this embodiment, the voice interaction device performs a target soothing operation based on the emotion recognition data and the corresponding emotion soothing strategy to obtain an emotion soothing result, including:
[0141] Obtain the target emotion level from the emotion recognition data, and generate an emotion soothing test strategy corresponding to the target emotion level;
[0142] Based on the stated emotion-soothing strategy, corresponding emotion-soothing content is output, and the emotion-soothing result is obtained.
[0143] In one possible implementation of this embodiment, the voice interaction device outputs corresponding emotional soothing content based on the emotional soothing strategy to obtain an emotional soothing result, including:
[0144] Access the cloud database to obtain the emotion soothing strategy and the emotion soothing content corresponding to the interactive voice data;
[0145] The emotional comforting content is output to the target object, and the target object's emotional update data is collected;
[0146] By comparing the emotion update data and the emotion recognition data, the emotion soothing result of the target object is obtained.
[0147] In this embodiment, the voice interaction device collects the interactive voice data and image data of the target object in response to a voice soothing request triggered by the target object. Emotional data is used to determine the target object's emotion recognition data based on the interactive voice features corresponding to the interactive voice data and the image emotion data corresponding to the object image data. A target soothing operation is then performed based on the emotion recognition data and the corresponding emotion soothing strategy to obtain the emotion soothing result. This approach achieves the goal of acquiring the target object's interactive voice data and object image data during voice interaction, performing multi-dimensional emotion recognition based on the interactive voice data and object image data, thereby determining the target object's emotion recognition data during voice interaction, and performing corresponding soothing operations based on the emotion soothing strategy corresponding to the emotion mouse data. This reduces the probability of users experiencing anger during voice interaction, improving the user's voice interaction experience and satisfaction.
[0148] This invention also provides a voice interaction device, such as... Figure 6 As shown, Figure 6 This is a schematic diagram of an embodiment of the voice interaction device provided in this application.
[0149] The voice interaction device integrates any one of the voice interaction devices provided in the embodiments of the present invention, and the voice interaction device includes:
[0150] One or more processors;
[0151] Memory; and
[0152] One or more applications, wherein the one or more applications are stored in the memory and configured by the processor to execute the steps of the voice interaction method in any of the embodiments described above.
[0153] Specifically, a voice interaction device may include components such as a processor 601 with one or more processing cores, a memory 602 with one or more computer-readable storage media, a power supply 603, and an input unit 604. Those skilled in the art will understand that... Figure 6 The structure of the voice interaction device shown does not constitute a limitation on the voice interaction device. It may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0154] The processor 601 is the control center of the voice interaction device. It connects various parts of the device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 602, and by calling data stored in the memory 602, thereby providing overall monitoring of the voice interaction device. Optionally, the processor 601 may include one or more processing cores; preferably, the processor 601 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 601.
[0155] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and data processing by running the software programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created based on the use of the voice interaction device, etc. In addition, the memory 602 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 602 may also include a memory controller to provide the processor 601 with access to the memory 602.
[0156] The voice interaction device also includes a power supply 603 that supplies power to the various components. Preferably, the power supply 603 can be logically connected to the processor 601 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 603 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0157] The voice interaction device may also include an input unit 604, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0158] Although not shown, the voice interaction device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 601 in the voice interaction device loads the executable files corresponding to the processes of one or more applications into the memory 602 according to the following instructions, and the processor 601 runs the applications stored in the memory 602 to realize various functions, as follows:
[0159] In response to a voice soothing request triggered by the target object, the system collects the interactive voice data and object image data of the target object.
[0160] The emotion data determines the emotion recognition data of the target object based on the interactive voice features corresponding to the interactive voice data and the image emotion data corresponding to the object image data.
[0161] The target soothing operation is performed based on the emotion recognition data and the corresponding emotion soothing strategy to obtain the emotion soothing result.
[0162] Therefore, embodiments of the present invention provide a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc. A computer program is stored thereon, which is loaded by a processor to execute the steps in any of the voice interaction methods provided in the embodiments of the present invention. For example, the computer program loaded by the processor can execute the following steps:
[0163] In response to a voice soothing request triggered by the target object, the system collects the interactive voice data and object image data of the target object.
[0164] The emotion data determines the emotion recognition data of the target object based on the interactive voice features corresponding to the interactive voice data and the image emotion data corresponding to the object image data.
[0165] The target soothing operation is performed based on the emotion recognition data and the corresponding emotion soothing strategy to obtain the emotion soothing result.
[0166] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.
[0167] In practice, each of the above units or structures can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units or structures, please refer to the previous method embodiments, which will not be repeated here.
[0168] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0169] The above provides a detailed description of a voice interaction method provided by the embodiments of this application. Specific embodiments have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A voice interaction method, characterized in that, The voice interaction method includes: In response to a voice soothing request triggered by the target object, the system collects the interactive voice data and object image data of the target object. Emotion classification is performed on the interactive voice features corresponding to the interactive voice data to obtain voice emotion data corresponding to the interactive voice features; facial expression recognition is performed on the object image features corresponding to the object image data to obtain image emotion data corresponding to the object image data; the voice emotion level corresponding to the voice emotion data is obtained, and the image emotion level corresponding to the image emotion data is obtained. If the difference between the voice emotion level and the image emotion level is less than or equal to a preset first difference threshold, then the voice emotion level or image emotion level with the larger value is used as the emotion recognition data of the target object; and / or, if the difference between the voice emotion level and the image emotion level is greater than or equal to a preset second difference threshold, then the voice recognition model is used to perform secondary recognition on the interactive voice data and the interactive voice data to obtain the emotion recognition data of the target object; The target soothing operation is performed based on the emotion recognition data and the corresponding emotion soothing strategy to obtain the emotion soothing result.
2. The voice interaction method according to claim 1, characterized in that, The step of performing emotion classification on the interactive voice features corresponding to the interactive voice data to obtain the voice emotion data corresponding to the interactive voice features includes: The interactive voice data is subjected to voice feature extraction to obtain the interactive voice features of the interactive voice data. The interactive speech features are initially classified into emotions to obtain the initial emotion features of the interactive speech features; If the initial emotional feature is a negative emotional feature, then the initial emotional feature is subjected to secondary emotional classification according to the preset speech recognition model to obtain the speech emotion data corresponding to the interactive speech feature.
3. The voice interaction method according to claim 1, characterized in that, The step of performing facial expression recognition on the object image features corresponding to the object image data to obtain image emotion data corresponding to the object image data includes: Facial features are extracted from the object image data to obtain the object image features corresponding to the object image data; The target convolution module, stochastic gradient descent module, and target feature extractor are used to perform facial expression recognition on the object image features to obtain the image emotion data corresponding to the object image data.
4. The voice interaction method according to claim 1, characterized in that, The step of using a speech recognition model to perform secondary recognition on the interactive speech data and the interactive speech data to obtain the emotion recognition data of the target object includes: The interactive voice data and the interactive voice data are re-identified using a speech recognition model to obtain an updated emotion level, which includes an updated voice emotion level and / or an updated image emotion level. The average value of the voice emotion level, image emotion level, updated voice emotion level, and / or updated image emotion level is calculated to obtain the emotion recognition data of the target object.
5. The voice interaction method according to any one of claims 1-4, characterized in that, The step of performing a target soothing operation based on the emotion recognition data and the corresponding emotion soothing strategy to obtain an emotion soothing result includes: Obtain the target emotion level from the emotion recognition data, and generate an emotion soothing test strategy corresponding to the target emotion level; Based on the stated emotion-soothing strategy, corresponding emotion-soothing content is output, and the emotion-soothing result is obtained.
6. The voice interaction method according to claim 5, characterized in that, The process of outputting corresponding emotional soothing content based on the emotional soothing strategy to obtain emotional soothing results includes: Access the cloud database to obtain the emotion soothing strategy and the emotion soothing content corresponding to the interactive voice data; The emotional comforting content is output to the target object, and the target object's emotional update data is collected; By comparing the emotion update data and the emotion recognition data, the emotion soothing result of the target object is obtained.
7. A voice interaction device, characterized in that, The voice interaction device includes: The data acquisition module is configured to respond to a voice reassurance request triggered by the target object and acquire the interactive voice data and object image data of the target object; The emotion recognition module is configured as follows: The interactive voice features corresponding to the interactive voice data are classified for emotion to obtain voice emotion data corresponding to the interactive voice features. Expression recognition is performed on the object image features corresponding to the object image data to obtain image emotion data corresponding to the object image data. The voice emotion level corresponding to the voice emotion data and the image emotion level corresponding to the image emotion data are obtained. If the difference between the voice emotion level and the image emotion level is less than or equal to a preset first difference threshold, the voice emotion level or image emotion level with the larger value is used as the emotion recognition data of the target object. And / or, if the difference between the voice emotion level and the image emotion level is greater than or equal to a preset second difference threshold, a secondary recognition is performed on the interactive voice data using a voice recognition model to obtain the emotion recognition data of the target object. The voice interaction module is configured to perform a target soothing operation based on the emotion recognition data and the corresponding emotion soothing strategy, and obtain an emotion soothing result.
8. A voice interaction device, characterized in that, The voice interaction device includes: One or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the steps of the voice interaction method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Voice interaction method and device, equipment and storage medium
CN109036405A