Voice interaction method and device and electronic equipment

By performing emotional analysis on the user's voice commands and images, a voice style adapted to the target emotional state is generated, solving the cumbersome problem of voice style switching and achieving real-time emotional response and negative emotion relief.

CN120656456APending Publication Date: 2025-09-16GUANGZHOU AUTOMOBILE GROUP CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511128783.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In the existing technology, the switching operation of voice styles in the vehicle system is cumbersome and cannot be achieved in real time, which cannot meet the user's real-time emotional change needs.

Method used

By obtaining the user's voice commands and user images, emotion classification is performed separately, and combined with the speech generation model to generate a voice style that adapts to the target emotional state to alleviate negative emotions.

Benefits of technology

It realizes real-time switching of voice styles, accurately adjusts voice responses according to the user's emotional state, and effectively alleviates the user's negative emotions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656456A_ABST
    Figure CN120656456A_ABST
Patent Text Reader

Abstract

The invention relates to a voice interaction method and apparatus, and an electronic device. The method comprises the steps of obtaining a voice instruction and a user image of a user; performing emotion classification on the voice instruction to obtain a first emotion classification result; performing emotion classification on the user image to obtain a second emotion classification result; determining a target emotion state of the user based on the first emotion classification result and the second emotion classification result; a voice generation model generates reply voice in the target voice style according to the target emotional state and a reply text for the voice instruction; wherein the target voice style is the voice style matched with the target emotional state, namely the voice style for relieving the negative emotion when the target emotional state indicates the negative emotion; according to the method provided by the invention, real-time switching of voice styles can be realized according to the state of the user; and the voice style considers the emotional state of the user and the text content of the reply text at the same time, so that the voice style of the reply voice is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a voice interaction method, device, and electronic device. Background Art

[0002] With the development of vehicle intelligence, voice interaction has gradually become an important means for car users to interact with vehicles.

[0003] In related technologies, users can choose a preset voice style in the vehicle system or upload their own favorite voice styles for voice interaction, but the operation is relatively cumbersome and real-time switching of voice styles cannot be achieved. Summary of the Invention

[0004] In view of this, the embodiments of the present application propose a voice interaction method, device and electronic device, which can automatically switch the voice style according to the user's real-time emotions.

[0005] The embodiments of the present application are implemented using the following technical solutions: In a first aspect, an embodiment of the present application provides a voice interaction method, comprising: obtaining a user's voice command and a user image; performing emotion classification on the voice command to obtain a first emotion classification result; performing emotion classification on the user image to obtain a second emotion classification result; determining the user's target emotional state based on the first emotion classification result and the second emotion classification result; generating a reply voice in a target voice style by a voice generation model based on the target emotional state and a reply text to the voice command; wherein the target voice style is a voice style adapted to the target emotional state; and the voice style adapted to the target emotional state is used to alleviate the negative emotion when the target emotional state indicates a negative emotion.

[0006] In the second aspect, an embodiment of the present application provides a voice interaction device, which includes: an acquisition module for acquiring a user's voice instructions and user images; a first emotion classification module for performing emotion classification on the voice instructions to obtain a first emotion classification result; a second emotion classification module for performing emotion classification on the user image to obtain a second emotion classification result; an emotion determination module for determining the target emotional state of the user based on the first emotion classification result and the second emotion classification result; an output module for generating a reply voice in a target voice style according to the target emotional state and the reply text to the voice instructions by a voice generation model; wherein the target voice style is a voice style adapted to the target emotional state; and the voice style adapted to the target emotional state is used to alleviate the negative emotion when the target emotional state indicates a negative emotion.

[0007] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor; a memory, wherein computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the processor, the above-mentioned method is implemented.

[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores computer-readable instructions. When the computer-readable instructions are executed by a processor, the above method is implemented.

[0009] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising computer instructions, which implement the above method when executed by a processor.

[0010] The present application provides a voice interaction method, device and electronic device, the method comprising: obtaining a user's voice command and user image; performing emotion classification on the voice command to obtain a first emotion classification result; performing emotion classification on the user image to obtain a second emotion classification result; determining the user's target emotional state based on the first emotion classification result and the second emotion classification result; it is understandable that when a user expresses emotions, not only will there be ups and downs in voice expression, but also changes in some body movements, such as changes in facial expressions, changes in behavior, etc. Therefore, emotion analysis can be performed on the voice command and the user image respectively to obtain emotional expressions in multiple dimensions, and then the target emotional state is determined based on the emotional expressions in multiple dimensions, so that the target emotional state of the user obtained is more accurate; finally, the speech generation model is used to generate the target emotional state of the user. According to the target emotional state and the reply text to the voice command, a reply voice in a target voice style is generated; wherein, the target voice style is a voice style adapted to the target emotional state, and the voice style adapted to the target emotional state is used to alleviate negative emotions when the target emotional state indicates negative emotions; through the above-mentioned voice interaction method, the adapted voice style can be switched in real time according to the user's state for voice reply; at the same time, since the target voice style of the reply voice is determined based on the target emotional state and the reply text to the voice command, the target voice style takes into account both the user's target emotional state and the emotions that may be contained in the text content of the reply text itself, thereby making the target voice style of the reply voice more accurate and can timely alleviate the user's negative emotions when the user's target emotional state is negative emotions.

[0011] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0013] Figure 1 A flow chart of a voice interaction method according to an embodiment of the present application is shown.

[0014] Figure 2 A schematic diagram of the training process of the speech analysis model involved in an embodiment of the present application is shown.

[0015] Figure 3 A schematic diagram of the training process of the image analysis model involved in an embodiment of the present application is shown.

[0016] Figure 4 An embodiment of the present application provides Figure 3 Flow chart of step S320 in FIG.

[0017] Figure 5 An embodiment of the present application provides Figure 1 Flow chart of step S140 in FIG.

[0018] Figure 6 An embodiment of the present application provides Figure 1 Another flowchart of step S140 in FIG.

[0019] Figure 7 A schematic diagram of the training process of the speech generation model involved in an embodiment of the present application is shown.

[0020] Figure 8 An embodiment of the present application provides Figure 7 Flow chart of step S620 in FIG.

[0021] Figure 9 An application scenario diagram involved in an embodiment of the present application is shown.

[0022] Figure 10 A schematic diagram of a voice interaction device provided in an embodiment of the present application is shown.

[0023] Figure 11 A schematic diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0024] In order to make the technical problems, technical solutions and beneficial effects solved by this application more clearly understood, this application is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0025] With the development of intelligent vehicles, voice interaction has gradually become an important means for car users to interact with their vehicles. In related technologies, users can select preset voice styles in the vehicle system or upload their own preferred voice styles for voice interaction. However, this operation is cumbersome and cannot achieve real-time switching of voice styles.

[0026] An embodiment of the present application provides a voice interaction method, which includes: obtaining a user's voice command and a user image; performing emotion classification on the voice command to obtain a first emotion classification result; performing emotion classification on the user image to obtain a second emotion classification result; determining the user's target emotional state based on the first emotion classification result and the second emotion classification result; finally, a voice generation model generates a reply voice in a target voice style according to the target emotional state and a reply text to the voice command; wherein the target voice style is a voice style adapted to the target emotional state; the voice style adapted to the target emotional state is used to alleviate negative emotions when the target emotional state indicates negative emotions.

[0027] Through the method provided in the embodiment of the present application, the adaptive voice style can be switched in real time according to the user's emotional state for voice reply; and because when the user expresses emotions, there will be not only ups and downs in the voice expression, but also changes in some body movements, such as changes in expression, changes in behavior, etc. Therefore, the method provided in the present application obtains emotional expressions in multiple dimensions by performing emotional analysis on voice commands and user images respectively, and then determines the target emotional state according to the emotional expressions in multiple dimensions, so that the target emotional state of the user is more accurate; at the same time, because the target voice style of the reply voice is determined based on the target emotional state and the reply text to the voice command, the target voice style takes into account both the target emotional state of the user and the emotions that may be contained in the text content of the reply text itself, thereby making the target voice style of the reply voice more accurate and, when the user's target emotional state is negative, it can alleviate the user's negative emotions.

[0028] See also Figure 1 , Figure 1 The flowchart of the voice interaction method provided in the embodiment of the present application is shown. The voice interaction method includes steps S110-S150: S110: Acquire a user's voice command and user image.

[0029] Here, a user refers to a person who is using a voice interaction service or a voice interaction device; the voice interaction service or the voice interaction device can provide a voice response to the user's voice commands.

[0030] In some embodiments, the voice interaction service is, for example, a voice assistant service provided by a vehicle, mobile phone, computer, or other electronic device with voice interaction functionality; a voice interaction device is, for example, a smart speaker, a smart interactive screen, etc.

[0031] The user's voice command refers to a command or request issued by the user through voice; in some embodiments, the user's voice command can be collected by a microphone; further, in order to facilitate the subsequent input of the voice analysis model for emotion classification, the voice commands collected by the microphone can be further processed, for example, the sound waves can be converted into analog signals, and the analog signals can be converted into digital formats for signal enhancement and coding compression.

[0032] The user image of the user refers to the user image collected when the user issues a voice command; the collected user image can be one or more; in some implementations, the user image of the user can be collected through a camera.

[0033] In some embodiments, the user can wake up the voice interaction service or start the voice interaction device through a specific wake-up word. At the same time, when the voice interaction service is woken up or the voice interaction device is started, the voice interaction service or voice interaction device can call the microphone to start receiving the user's voice commands, and call the camera to capture the user's image.

[0034] S120: Emotionally classify the voice command to obtain a first emotion classification result.

[0035] The first emotion classification result includes at least a first emotion category; further, the types of emotion categories are pre-set, for example, the emotion categories can be divided into happiness, sadness, calmness, anger, anxiety, depression, excitement, etc.

[0036] In some embodiments, the first emotion classification result may also include a first confidence level of the first emotion category and a first emotion level under the first emotion category; wherein the first confidence level is used to describe the credibility of the first emotion category, and the higher the first confidence level, the higher the reliability of the first emotion category; the first emotion level under the first emotion category is used to describe the emotional intensity of the first emotion category.

[0037] For ease of understanding, illustratively, the first emotion category is "happy", corresponding to three first emotion levels, specifically, the first level "somewhat happy", the second level "normally happy" and the third level "very happy".

[0038] In some embodiments, emotion analysis can be performed based on the acoustic features of the voice command; for example, the timbre characteristics of the voice command can be analyzed using the Mel-Frequency Cepstral Coefficients (MFCC, used to characterize the spectral envelope of sound) of the voice command; the changes in timbre and pitch can be captured using the energy distribution of the voice command, thereby obtaining a first emotion classification result based on the timbre characteristics and the changes in timbre and pitch.

[0039] In other embodiments, a neural network model may be used to extract deeper features from voice instructions, thereby realizing sentiment analysis of the voice instructions. Specifically, step S120 may include: the voice analysis model performs sentiment classification on the voice instructions to obtain a first sentiment classification result.

[0040] Among them, the training process of the speech analysis model is as follows Figure 2 As shown, Figure 2 The figure shows a training process diagram of the speech analysis model involved in the embodiment of the present application. The training process of the speech analysis model includes S210-S250: S210: Acquire training speech and first emotion category labeling information of the training speech.

[0041] In some embodiments, the training speech may be a collected voice instruction or other voice content; the first emotion category labeling information of the training speech includes at least the emotion category of the training speech, and on this basis, may also include the emotion level of the emotion category.

[0042] S220: The speech analysis model performs feature extraction on the training speech to obtain speech features.

[0043] Among them, speech features may include acoustic features of voice commands, such as Mel-frequency cepstral coefficients (MFCC, which can characterize the spectral envelope of sound and reflect timbre characteristics); Linear Predictive Coding (LPC, which can efficiently represent voice signals at low bit rates and remove redundant information in voice signals); Harmonic-to-Noise Ratio (HNR, which can measure the proportion of harmonic components in sound and help distinguish different timbres); fundamental frequency (F0, which can be estimated from each segment by an automatic fundamental frequency detection algorithm); energy distribution (which can capture changes in timbre and pitch by analyzing the energy distribution of different frequency bands on a spectrogram); and other audio features of voice commands, such as timbre features, pitch features, etc.

[0044] In some embodiments, the speech analysis model can be a model based on the Transformer architecture; further, the Transformer architecture adopts an encoder-decoder structure, wherein there are multiple encoders, each of which can include a multi-head self-attention layer and a feedforward neural network. The audio features of the voice instructions can be proposed layer by layer through the multiple encoders. At the same time, the multi-head self-attention layer can capture the long-distance dependencies in the voice instructions, and can also retain sequence information by adding position encoding to help the model understand the time dimension of the voice instructions; similarly, there are multiple decoders, each of which can include a masked multi-head self-attention layer, an interactive attention layer and a feedforward neural network. The extracted features can be decoded layer by layer through the multiple decoders to obtain the target output, that is, the first predicted emotion classification result. At the same time, the decoder can force the model to learn a more robust feature representation by introducing a masking mechanism (that is, randomly blocking a part of the input features).

[0045] S230 , performing emotion classification based on the speech features by the speech analysis model to obtain a first predicted emotion classification result.

[0046] The first predicted emotion classification result at least includes the predicted emotion category of the training speech, and on this basis, may also include the predicted emotion level of the predicted emotion category.

[0047] In some embodiments, the content included in the first predicted emotion classification result is consistent with the content included in the first emotion category labeling information; for example, if the first emotion category labeling information only includes the emotion category of the training speech, then the first predicted emotion classification result also only includes the predicted emotion category of the training speech.

[0048] S240: Calculate a second loss based on the first predicted emotion classification result and the first emotion category labeling information.

[0049] The second loss can be calculated by a loss function; the loss function may be, for example, a cross entropy loss, a KL divergence loss, or the like.

[0050] S250. Adjust the model parameters of the speech analysis model based on the second loss until the second training end condition is reached.

[0051] Specifically, the model parameters of the speech analysis model are adjusted in a direction to minimize the second loss until the second training end condition is reached.

[0052] The second training end condition may be that the number of training iterations reaches a second threshold, or that the second loss is less than a second loss threshold.

[0053] Through the above training of the speech analysis model, the trained speech analysis model can be used to perform emotion classification on the voice instruction, thereby obtaining a first emotion classification result of the voice instruction.

[0054] In some embodiments, the voice command may be pre-processed before executing step S120; for example, the user associated with the voice command may be anonymized to remove personal identity information and protect user privacy; for example, for longer voice commands, the continuous voice stream may be segmented into short segments (such as per second or per sentence) to better capture instantaneous changes; digital signal processing technology may also be used to remove background noise and enhance human voices.

[0055] S130: Emotionally classify the user image to obtain a second emotion classification result.

[0056] The second emotion classification result includes at least a second emotion category; wherein the type of emotion category is pre-set, for example, the emotion category can be divided into happiness, sadness, calmness, anger, anxiety, depression, excitement, etc.

[0057] In some embodiments, the second emotion classification result may also include a second confidence level of the second emotion category, and a second emotion level under the second emotion category; wherein the second confidence level is used to describe the credibility of the second emotion category, and the higher the second confidence level, the higher the reliability of the second emotion category; the second emotion level under the second emotion category is used to describe the emotional intensity of the second emotion category.

[0058] In some implementations, a neural network model may be used to perform emotion classification on the user image. Specifically, step S130 may include: performing emotion classification on the user image using an image analysis model to obtain a second emotion classification result.

[0059] Among them, the training process of the image analysis model is as follows Figure 3 As shown, Figure 3 The following is a schematic diagram of the training process of the image analysis model involved in the embodiment of the present application. The training process of the image analysis model includes steps S310-S350: S310: Acquire a training image and second emotion category labeling information of the training image.

[0060] Among them, the training image refers to the user image in the voice interaction scenario; for example, if the voice interaction scenario is voice interaction with a vehicle voice assistant, the training image is the image of the vehicle user, such as the image of the driver, the image of the vehicle passenger, etc.; if the voice interaction scenario is voice interaction with a mobile phone assistant, the training image is the image of the mobile phone user, such as the image of the user captured by the front camera of the mobile phone.

[0061] In some embodiments, in order to ensure the consistency of the training images and facilitate the subsequent input of the image analysis model for emotion classification, after obtaining the training images, the obtained training images can also be preprocessed; preprocessing such as image cropping, image scaling, normalization and other operations.

[0062] The second emotion category labeling information of the training image at least includes the emotion category of the training image, and on this basis, may also include the emotion level of the emotion category.

[0063] In some implementations, the second emotion category labeling information may be manually labeled by a tester.

[0064] S320: The image analysis model performs feature extraction on the training image to obtain image features.

[0065] It can be understood that the user expressions, user behaviors and user clothing included in the training images can express the user's emotions to a certain extent; therefore, a variety of features that can express user emotions can be extracted from the training images, such as user expression features, user clothing features and user behavior features.

[0066] In some embodiments, the image feature may be a single feature, such as one of a user expression feature, a user clothing feature, and a user behavior feature; the image feature may also be a fusion feature, such as a feature obtained by fusing a user expression feature, a user clothing feature, and a user behavior feature.

[0067] In some embodiments, when the image features are fused features, such as the image features are features obtained by fusing user expression features, user clothing features, and user behavior features, the image analysis model can be a multimodal fusion model based on the transformer architecture; wherein, the description of the transformer architecture can refer to the specific description in step S220 in the aforementioned embodiment, which will not be repeated here; the difference is that the image analysis model can include multiple different feature extraction modules for extracting features from different angles, and then perform feature fusion (such as feature splicing, feature alignment, etc.) on the extracted features from multiple different angles to obtain the final image features, so that the model can automatically pay attention to the correlation and importance between different features, thereby making more accurate emotional judgments.

[0068] In some embodiments, see Figure 4 , Figure 4 The embodiment of the present application provides Figure 3 Schematic diagram of the process of step S320, step S320 includes steps S410-S440: S410: Extract user expression features from the training image using an image analysis model to obtain user expression features.

[0069] It is understandable that user expressions can reflect the user's emotional state; for example, an expression with upturned corners of the mouth can express happiness, an expression with drooping corners of the mouth or furrowed eyebrows can express nervousness, etc.

[0070] S420: Extract user clothing features from the training image using the image analysis model to obtain user clothing features.

[0071] It is understandable that a user's clothing can also reflect the user's emotional state to a certain extent; for example, bright-colored clothing can express excitement or positivity, formal clothing can express seriousness or steadiness, and cartoon clothing can express optimism or relaxation.

[0072] S430: Extract user behavior features from the training image using the image analysis model to obtain user behavior features.

[0073] It is understandable that user behavior is also one of the ways for users to express their emotions. For example, scratching the head can express anxiety, clenching hands can express tension, and leaning the body can express relaxation.

[0074] S440: The image analysis model performs feature fusion on the user's facial expression features, the user's clothing features, and the user's behavior features to obtain image features.

[0075] In some embodiments, feature fusion can be feature splicing or feature alignment, and the specific fusion method is not limited here.

[0076] In the above embodiment, image features are obtained by fusing user expression features, user clothing features, and user behavior features, so that the obtained image features can include the associations between different features, thereby more comprehensively describing the user's emotional state.

[0077] S330: The image analysis model performs emotion classification based on the image features to obtain a second predicted emotion classification result.

[0078] The second predicted emotion classification result at least includes the predicted emotion category of the training image, and on this basis, may also include the predicted emotion level of the predicted emotion category.

[0079] In some embodiments, the content included in the second predicted emotion classification result is consistent with the content included in the second emotion category labeling information; for example, if the second emotion category labeling information only includes the emotion category of the training speech, then the second predicted emotion classification result also only includes the predicted emotion category of the training speech.

[0080] S340: Calculate a third loss based on the second predicted emotion classification result and the second emotion category labeling information.

[0081] Among them, the third loss can be calculated by a loss function; the loss function can be, for example, cross entropy loss, KL divergence loss, etc.

[0082] S350. Adjust the model parameters of the image analysis model based on the third loss until a third training end condition is reached.

[0083] Specifically, the model parameters of the image analysis are adjusted in a direction of minimizing the third loss until the third training end condition is reached.

[0084] The third training end condition may be that the number of training iterations reaches a third threshold, or that the third loss is less than a third loss threshold.

[0085] Through the above training of the image analysis model, the trained image analysis model can be used to perform emotion classification based on features of multiple different angles contained in the training image, thereby making the second emotion classification result of the training image more accurate.

[0086] S140 : Determine the target emotional state of the user based on the first emotion classification result and the second emotion classification result.

[0087] In some embodiments, the first emotion classification result includes a first emotion category and a first emotion level under the first emotion category; the second emotion classification result includes a second emotion category and a second emotion level under the second emotion category; see Figure 5 , Figure 5 The embodiment of the present application provides Figure 1 Schematic diagram of the process of step S140, step S140 includes steps S510-S530: S510: If the first emotion category and the second emotion category are the same, use the first emotion category or the second emotion category as the target emotion category.

[0088] It should be noted that the first emotion category and the second emotion category being the same does not necessarily require the first emotion level and the second emotion level to be the same; they may be the same emotion category but have different emotion levels.

[0089] S520: Determine a target emotion level under a target emotion category based on the first emotion level and the second emotion level.

[0090] In some embodiments, the average of the first emotion level and the second emotion level may be used as the target emotion level under the target emotion category; or the first emotion level and the second emotion level may be weightedly summed and the weighted sum result may be used as the target emotion level under the target emotion category.

[0091] Considering that the emotion level is usually set as an integer, after calculating the mean or weighted sum result, the mean or weighted sum result can be rounded, that is, the integer closest to the mean or weighted sum result is taken as the target emotion level.

[0092] S530: Use the target emotion category and the target emotion level as the user's target emotional state.

[0093] In the above embodiment, when the first emotion category and the second emotion category are the same, it means that both emotion classification results have a certain reliability. Therefore, the two emotion classification results are comprehensively considered to determine the user's target emotional state, so that the determined target emotional state of the user is more accurate.

[0094] In some embodiments, based on the first emotion classification result including the first emotion category and the first emotion level under the first emotion category, the first emotion classification result also includes a first confidence level for the first emotion category; and based on the second emotion classification result including the second emotion category and the second emotion level under the second emotion category, the second emotion classification result also includes a second confidence level for the second emotion category; see Figure 6 , Figure 6 The embodiment of the present application provides Figure 1 Another flow chart of step S140, step S140 may further include steps S540-S550: S540: If the first emotion category and the second emotion category are different, determine the maximum value of the first confidence level and the second confidence level as the target confidence level.

[0095] S550: Taking the emotion category and emotion level included in the emotion classification result corresponding to the target confidence as the target emotional state of the user.

[0096] It can be understood that if the first emotion category and the second emotion category are different, it means that there is a disagreement between the two emotion classification results, that is, at least one of the emotion classification results is unreliable. In this case, the more reliable emotion classification result is preferred through the target confidence, so as to avoid the unreliable emotion classification result affecting the accuracy of the target emotional state.

[0097] S150. The speech generation model generates a reply speech in a target speech style according to the target emotional state and the reply text to the voice command; wherein the target speech style is a speech style adapted to the target emotional state; the speech style adapted to the target emotional state is used to alleviate negative emotions when the target emotional state indicates negative emotions.

[0098] In some embodiments, the reply text for the voice command can be generated by other text generation models based on the instruction text contained in the voice command; it can also be obtained by querying a preset corpus based on the instruction text contained in the voice command; wherein the preset corpus can contain reply corpus text corresponding to various voice commands.

[0099] The speech generation model generates a reply speech in a target voice style according to the target emotional state and the reply text to the voice command. This can be broken down into two steps. The speech generation model determines the target voice style according to the target emotional state and the reply text to the voice command. Then, the speech generation model generates a reply speech in the target voice style according to the target voice style and the reply text to the voice command.

[0100] The target voice style adapts to the target emotional state, which means that the target voice style is consistent with the communication style under the target emotional state. For example, if the target emotional state is happy, the overall communication style should be cheerful and joyful. Therefore, the adapted target voice style can be playful, excited, etc., rather than an overly calm or serious voice style.

[0101] A voice style that adapts to the target emotion means that the target voice style has a positive effect on the target emotion indicated by the target emotional state; specifically, a voice style that adapts to the target emotion is a voice style that alleviates the current negative emotion when the target emotional state indicates a negative emotion; for example, a "comforting" or "soothing" voice style can effectively alleviate the "anxiety" emotion. For the "anxiety" emotion, the "comforting" or "soothing" voice style is the voice style that adapts to the "anxiety" emotion; for another example, a "enthusiastic" or "encouraging" voice style can effectively alleviate the "depression" emotion. For the "depression" emotion, the "enthusiastic" or "encouraging" voice style is the voice style that adapts to the "depression" emotion.

[0102] In other embodiments, the voice style adapted to the target emotion may also be a voice style that does not destroy the current non-negative emotion when the target emotional state indicates a non-negative emotion; for example, for the emotion of "happy", voice styles such as "playful", "excited" or "funny" are not likely to destroy the user's "happy" emotion. Therefore, for the emotion of "happy", voice styles such as "playful", "excited" or "funny" are the voice styles adapted to the emotion of "happy".

[0103] In some implementations, the reply text included in the reply voice may be consistent with the reply text to the voice instruction, that is, the reply voice is the reply text in the target voice style.

[0104] In other embodiments, the reply text included in the reply voice can be determined based on the reply text to the voice command and the target voice style, so that the reply text included in the reply voice is more consistent with the target voice style. For example, the reply text to the voice command can be "The temperature today is 35 degrees Celsius." When the target voice style is "calm," the reply text can also be "The temperature today is 35 degrees Celsius. Please pay attention to sun protection." When the target voice style is "playful," the reply text can also be "It's 35 degrees Celsius today. Remember to protect yourself from the sun and don't be lazy."

[0105] The method provided by the embodiment of the present application obtains the user's voice command and user image; performs emotion classification on the voice command to obtain a first emotion classification result; performs emotion classification on the user image to obtain a second emotion classification result; determines the user's target emotional state based on the first emotion classification result and the second emotion classification result; finally, the speech generation model generates a reply speech in a target voice style according to the target emotional state and the reply text to the voice command; wherein the target voice style is a voice style adapted to the target emotional state, and the voice style adapted to the target emotional state is used to alleviate negative emotions when the target emotional state indicates negative emotions; through the method provided by the embodiment of the present application, the adapted voice style can be switched in real time according to the user's state for voice reply; and because the user expresses emotions , not only will there be fluctuations in voice expression, but it may also be reflected in changes in some body movements, such as changes in facial expressions, changes in behavior, etc. Therefore, the method provided in this application obtains emotional expressions in multiple dimensions by performing emotional analysis on voice commands and user images respectively, and then determines the target emotional state according to the emotional expressions in multiple dimensions, so that the target emotional state of the user is more accurate; at the same time, since the target voice style of the reply voice is determined based on the target emotional state and the reply text to the voice command, the target voice style takes into account the target emotional state of the user and the emotions that may be contained in the text content of the reply text itself, thereby making the target voice style of the reply voice more accurate and when the user's target emotional state is negative, the user's negative emotions can be alleviated in time.

[0106] See also Figure 7 , Figure 7 A schematic diagram of the training process of the speech generation model involved in the embodiment of the present application is shown. The training process of the speech generation model includes steps S610-S640: S610: Acquire a sample speech, a sample text obtained by converting the sample speech into text, and a reference emotional state annotated for the sample speech; wherein the speech style of the sample speech is a speech style that alleviates the negative emotion indicated by the reference emotional state.

[0107] Among them, the sample voice can come from different channels such as professional voice recording, public voice datasets, etc.

[0108] Of course, the reference emotional state annotated for the sample speech may also be an emotional state indicating a non-negative emotion. In this case, the speech style of the sample speech is a speech style that does not destroy the non-negative emotion indicated by the reference emotional state.

[0109] In some embodiments, the sample voices obtained may be related to the voice interaction scenarios of actual applications. For example, taking the voice interaction scenario of car users as an example, the collected sample voices may be voice samples covering a variety of positive voice styles and corresponding to different driver emotional scenarios; for example, for drivers with "anxious" emotions, sample voices with soothing and relaxing tones are collected; for drivers with "depressed" emotions, enthusiastic and encouraging sample voices are collected, etc.

[0110] In some embodiments, the collected sample language can be preprocessed to facilitate use as input for a subsequent speech generation model; the preprocessing can include unifying the audio format (such as converting it into a common WAV format, etc.), adjusting the audio sampling rate (such as setting it to 16kHz or other suitable values), unifying the number of channels (usually mono or dual channels, unified according to needs), and other parameters to ensure the consistency of the audio data; the acoustic features of the audio can also be extracted, such as extracting Mel spectrum features, as auxiliary input for the subsequent model.

[0111] In some embodiments, the sample speech may be converted into text, and the text obtained by the text conversion may be directly used as the sample text; in other embodiments, the sample speech may be converted into text, and then the text obtained by the text conversion may be subjected to a series of text processing, such as removing modal particles, removing repeated words, and structuring processing, and the processed text may be used as the sample text.

[0112] S620: The speech generation model generates speech based on the sample text and the reference emotional state to obtain predicted speech in a predicted speech style.

[0113] It should be noted that, for pure text, although it lacks information such as timbre, pitch, and rhythm, the content of the text can also directly or indirectly reflect the expression of emotions to a certain extent; for example, the text "The weather is really nice today" indirectly expresses positive and optimistic emotions; the text "I am very happy to serve you" directly expresses happy emotions.

[0114] In some embodiments, the speech generation model can be a TTS (Text-to-Speech) model based on the transformer architecture; wherein, the TTS model can effectively capture long-distance dependencies in the text and the association between different modalities (such as emotion categories and text) by virtue of the multi-head attention mechanism, which helps to generate response speech that meets specific requirements.

[0115] In some embodiments, the sample text and the reference emotional state may be embedded separately to obtain a first embedding vector for the sample text and a second embedding vector for the reference emotional state. For example, for the sample text, word embedding may be used to obtain a first embedding vector. For the reference emotional state, a special embedding module may be created to convert the reference emotional state into a second embedding vector. A simple one-hot encoding method may be used followed by a linear layer, or a learnable embedding matrix may be used so that different reference emotional states can be input into the model in the form of suitable vectors, thereby facilitating the model to learn the speech style characteristics corresponding to different emotions. Thereafter, the first embedding vector and the second embedding vector are fused, such as by concatenation or addition, to obtain a fused vector. The speech generation model generates speech based on the fused vector to obtain predicted speech under the predicted speech style.

[0116] In other embodiments, see Figure 8 , Figure 8 The embodiment of the present application provides Figure 7 Schematic diagram of the process of step S620, step S620 includes steps S710-S740: S710: The speech generation model performs feature extraction on the sample text to obtain text content features.

[0117] Among them, the text content features include both the semantic content of the sample text and the emotions expressed by the sample text.

[0118] S720 , performing feature extraction on the reference emotion classification result by the speech generation model to obtain emotion features.

[0119] Among them, the emotional features represent the emotions expressed by the reference emotion classification results.

[0120] S730: The speech generation model performs feature fusion on the text content features and the emotion features to obtain intermediate features.

[0121] It is understandable that the intermediate features obtained contain both the emotions contained in the sample text and the emotions of the reference emotion classification results; for example, the speech generation model under the Transformer architecture can use multiple encoding and decoding layers to finely set hyperparameters such as the number of heads of the multi-head attention mechanism and the hidden layer dimension in each layer to ensure that the model can efficiently extract, convert and fuse the input information, and gradually construct intermediate features.

[0122] S740: The speech generation model generates speech based on the intermediate features to obtain predicted speech in a predicted speech style.

[0123] Specifically, the speech generation model can determine the acoustic features required for speech synthesis based on the intermediate features, usually Mel-spectrogram, etc. Then, with the help of a vocoder, these acoustic features are converted into actual speech audio to obtain the predicted speech in the predicted speech style.

[0124] In the embodiment of the present application, the speech generation model comprehensively considers the text content and the recognized emotional state when generating speech, so that the predicted speech style of the obtained predicted speech is more accurate.

[0125] S630: Calculate a first loss based on the speech features of the predicted speech and the speech features of the sample speech.

[0126] In some embodiments, the speech feature may be a Mel-spectrogram feature; it is understandable that by calculating the first loss using the speech feature, the speech style of the generated predicted speech may be more closely aligned with the speech style of the sample speech at the acoustic level.

[0127] The first loss can be calculated by a loss function, such as cross entropy loss, KL divergence loss, etc.

[0128] S640: Adjust model parameters of the speech generation model based on the first loss until a first training end condition is reached.

[0129] Specifically, the model parameters of the speech analysis model are adjusted in a direction to minimize the first loss until the first training end condition is reached.

[0130] The first training end condition may be that the number of training iterations reaches a first threshold, or that the first loss is less than a first loss threshold.

[0131] During the training process, you can select a suitable optimizer and, in combination with the set loss function, conduct multiple rounds of model training. By closely monitoring changes in indicators such as the loss value and acoustic feature similarity on the validation set, timely adjust hyperparameters such as the learning rate to prevent overfitting and ensure that the model can stably and effectively learn to generate reply speech in a positive voice style based on different emotions and different reply texts. If the speech generation model overfits, you can improve model performance by adding data augmentation methods (such as slight changes in the duration and pitch of the sample speech, synonym replacement of the text, etc.), adjusting model hyperparameters (such as reducing the number of layers, adjusting the number of heads, reducing the hidden layer dimension, etc.), etc.

[0132] Through the above-mentioned training of the speech generation model, the trained speech generation model can be used to comprehensively consider the text content and the recognized emotional state when generating speech, so that the predicted speech style of the predicted speech is more in line with the user's emotions, and at the same time, the text content of the predicted speech is more in line with the current speech style.

[0133] For easier understanding, see Figure 9 , Figure 9 An application scenario diagram involved in the present application is shown, which is applied to a target vehicle. The target vehicle includes an infotainment domain controller (IDC), a driver monitoring system (DMS), and a microphone; a speech analysis model, an image analysis module, and a speech generation model are deployed in the target vehicle.

[0134] Among them, the infotainment domain refers to the central computing platform in the car specifically used to process and manage infotainment system data, and is used to receive the user's target emotional state and receive the response voice in the method provided in this application.

[0135] A driver monitoring system refers to a system technology installed in a vehicle that monitors the driver's status in real time through cameras and sensors, and is used to provide driver images in the method provided in this application.

[0136] Specifically, the microphone can collect the user's voice instructions, and then the collected voice instructions are processed into voice signals and input into the voice analysis model to obtain a first emotion classification result.

[0137] At the same time, the driver monitoring system can collect driver images in real time through the camera; then, the collected driver images are input into the image analysis model to obtain the second emotion classification result.

[0138] Afterwards, the first emotion classification result and the second emotion classification result are respectively input into the infotainment domain, which determines the user's target emotional state based on the first emotion classification result and the second emotion classification result, and inputs the user's target emotional state into the speech generation model.

[0139] Finally, the speech generation model receives the user's target emotional state and the response text to the voice command, generates the reply speech in the target voice style, and sends it to the infotainment domain for playback.

[0140] In some embodiments, before sending the reply voice to the infotainment domain, the generated reply voice can be formatted to ensure that its format meets the requirements of storage and use in the infotainment domain; at the same time, corresponding metadata can be added to the audio file, such as emotion category labels, generation time, etc., to facilitate subsequent retrieval, management and use in the infotainment domain.

[0141] In some implementations, when the reply voice is sent to the infotainment domain, encrypted transmission or other means may be used to ensure transmission, thereby ensuring user privacy.

[0142] In the above application scenario, a reply voice in the target voice style can be generated according to the driver's emotional state and the emotions contained in the text content of the reply voice itself, so that when the user's target emotional state is negative, the user's negative emotions can be alleviated in time, thereby improving the safety and comfort of the driving experience.

[0143] In some embodiments, see Figure 10 , Figure 10 A schematic diagram of a voice interaction device provided in an embodiment of the present application is provided. The voice interaction device 800 includes: The acquisition module 810 is used to acquire the user's voice command and user image.

[0144] The first emotion classification module 820 is used to perform emotion classification on the voice instruction to obtain a first emotion classification result.

[0145] The second emotion classification module 830 is used to perform emotion classification on the user image to obtain a second emotion classification result.

[0146] The emotion determination module 840 is configured to determine a target emotional state of the user based on the first emotion classification result and the second emotion classification result.

[0147] The output module 850 is used to generate a reply speech in a target speech style according to the target emotional state and the reply text to the voice instruction by the speech generation model; wherein the target speech style is a speech style adapted to the target emotional state; the speech style adapted to the target emotional state is used to alleviate the negative emotion when the target emotional state indicates a negative emotion.

[0148] In some embodiments, the voice interaction device 800 also includes a voice generation model training module, which is used to obtain sample voice, sample text obtained by converting the sample voice into text, and a reference emotional state annotated for the sample voice; wherein the voice style of the sample voice is a voice style that alleviates the negative emotion indicated by the reference emotional state; the voice generation model generates voice according to the sample text and the reference emotional state to obtain predicted voice under the predicted voice style; the first loss is calculated based on the voice features of the predicted voice and the voice features of the sample voice; and the model parameters of the voice generation model are adjusted based on the first loss until the first training end condition is reached.

[0149] In some embodiments, the speech analysis model training module is specifically used to perform feature extraction on the sample text by the speech generation model to obtain text content features; and, perform feature extraction on the reference emotion classification results by the speech generation model to obtain emotion features; perform feature fusion on the text content features and emotion features by the speech generation model to obtain intermediate features; and perform speech generation by the speech generation model based on the intermediate features to obtain predicted speech under the predicted speech style.

[0150] In some embodiments, the first emotion classification result includes a first emotion category and a first emotion level under the first emotion category; the second emotion classification result includes a second emotion category and a second emotion level under the second emotion category; the emotion determination module 840 is specifically used to use the first emotion category or the second emotion category as the target emotion category if the first emotion category and the second emotion category are the same; based on the first emotion level and the second emotion level, determine the target emotion level under the target emotion category; and use the target emotion category and the target emotion level as the user's target emotional state.

[0151] In some embodiments, the first emotion classification result also includes a first confidence level of the first emotion category, and the second emotion classification result also includes a second confidence level of the second emotion category; the emotion determination module 840 is also used to determine the maximum value of the first confidence level and the second confidence level as the target confidence level if the first emotion category and the second emotion category are different; and use the emotion category and emotion level contained in the emotion classification result corresponding to the target confidence level as the user's target emotional state.

[0152] In some embodiments, the first emotion classification module 820 is specifically used to perform emotion classification on the voice instructions by the speech analysis model to obtain a first emotion classification result; the voice interaction device 800 also includes a speech analysis model training module, which is used to obtain training speech and first emotion category labeling information of the training speech; the speech analysis model performs feature extraction on the training speech to obtain speech features; and the speech analysis model performs emotion classification based on the speech features to obtain a first predicted emotion classification result; the second loss is calculated based on the first predicted emotion classification result and the first emotion category labeling information; and the model parameters of the speech analysis model are adjusted based on the second loss until the second training end condition is reached.

[0153] In some embodiments, the second emotion classification module 830 is specifically used to perform emotion classification on user images by an image analysis model to obtain a second emotion classification result; the voice interaction device 800 also includes an image analysis model training module, which is used to obtain training images and second emotion category labeling information of the training images; the image analysis model performs feature extraction on the training images to obtain image features; and the image analysis model performs emotion classification based on the image features to obtain a second predicted emotion classification result; the third loss is calculated based on the second predicted emotion classification result and the second emotion category labeling information; and the model parameters of the image analysis model are adjusted based on the third loss until the third training end condition is reached.

[0154] In some embodiments, the image analysis model training module is also used to extract user expression features from training images by the image analysis model to obtain user expression features; extract user clothing features from training images by the image analysis model to obtain user clothing features; extract user behavior features from training images by the image analysis model to obtain user behavior features; and perform feature fusion of user expression features, user clothing features, and user behavior features by the image analysis model to obtain image features.

[0155] In some implementations, according to the voice interaction method provided in the above embodiment, the present application also provides an electronic device, such as Figure 11 , Figure 11 A structural block diagram of an electronic device provided in an embodiment of the present application is given. The electronic device 900 includes a processor 910; a memory 920; and computer-readable instructions are stored on the memory 920. When the computer-readable instructions are executed by the processor 910, the above method is implemented.

[0156] The electronic device 900 may be a terminal device, and the terminal device may be a vehicle-mounted terminal, a desktop computer, a laptop computer, etc.

[0157] The processor 910 may include one or more processing cores. Using various interfaces and circuits, the processor 910 connects various components within the wearable device. It executes instructions, programs, code sets, or instruction sets stored in the memory 920, as well as accesses data stored in the memory 920, to perform various functions of the wearable device and process data. Optionally, the processor 910 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 910 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing displayed content; and the modem handles wireless communications. It is understood that the modem may also be implemented independently of the processor via a separate communications chip.

[0158] The memory 920 may include random access memory (RAM) or read-only memory (ROM). The memory 920 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 920 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data generated by the electronic device during use.

[0159] In some embodiments, the present application further provides a computer-readable storage medium having computer-readable instructions stored thereon. When the computer-readable instructions are executed by the processor 910, the above method is implemented.

[0160] The computer-readable storage medium may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Alternatively, the computer-readable storage medium includes a non-transitory computer-readable storage medium. The computer-readable storage medium has storage space for computer-readable instructions for executing any of the method steps described above. These computer-readable instructions can be read from or written to one or more computer program products. The computer-readable instructions may be compressed in a suitable format.

[0161] In particular, according to embodiments of the present application, the processes described above may be implemented as computer software programs. For example, embodiments of the present application include a computer program product comprising computer instructions. When the computer instructions are executed by a central processing unit (CPU), the various functions defined in the system of the present application are performed.

[0162] In this application, "a plurality" refers to two or more than two. In this application, unless otherwise expressly defined, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, or it can be internal communication between two components. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to the specific circumstances.

[0163] The terms "first," "second," "third," "fourth," etc. (if any) in this application are used to distinguish similar objects and are not necessarily used to describe a particular sequential order.

[0164] The term "and / or" in this application simply describes an association between related objects, indicating that three possible relationships exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this application generally indicates that the related objects are in an "or" relationship.

[0165] Unless otherwise specified, all steps of this application may be performed sequentially or randomly. For example, "a method includes steps A and B" means that the method may include steps A and B performed sequentially, or steps B and A performed sequentially. For example, "a method may also include step C" means that step C may be added to the method in any order. For example, the method may include steps A, B, and C, or steps A, C, and B, or steps C, A, and B, etc.

[0166] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A voice interaction method, characterized in that: include: Obtain user's voice commands and user images; Performing emotion classification on the voice command to obtain a first emotion classification result; Performing emotion classification on the user image to obtain a second emotion classification result; determining a target emotional state of the user based on the first emotion classification result and the second emotion classification result; The speech generation model generates a reply speech in a target speech style according to the target emotional state and the reply text to the speech instruction; wherein the target speech style is a speech style adapted to the target emotional state; the speech style adapted to the target emotional state is used to alleviate the negative emotion when the target emotional state indicates a negative emotion.

2. The method according to claim 1, characterized in that The training process of the speech generation model includes: Acquiring a sample speech, converting the sample speech into text to obtain a sample text, and a reference emotional state annotated for the sample speech; wherein the speech style of the sample speech is a speech style that alleviates the negative emotion indicated by the reference emotional state; The speech generation model generates speech based on the sample text and the reference emotional state to obtain predicted speech in a predicted speech style; Calculating a first loss based on the speech features of the predicted speech and the speech features of the sample speech; Adjust the model parameters of the speech generation model based on the first loss until a first training end condition is reached.

3. The method according to claim 2, characterized in that The speech generation model generates speech based on the sample text and the reference emotion classification result to obtain predicted speech in a predicted speech style, including: The speech generation model performs feature extraction on the sample text to obtain text content features; and The speech generation model performs feature extraction on the reference emotion classification result to obtain emotion features; The speech generation model performs feature fusion on the text content feature and the emotion feature to obtain an intermediate feature; The speech generation model performs speech generation based on the intermediate features to obtain predicted speech in a predicted speech style.

4. The method according to claim 1, wherein The first emotion classification result includes a first emotion category and a first emotion level under the first emotion category; the second emotion classification result includes a second emotion category and a second emotion level under the second emotion category; The determining the target emotional state of the user based on the first emotion classification result and the second emotion classification result includes: If the first emotion category and the second emotion category are the same, taking the first emotion category or the second emotion category as the target emotion category; determining a target emotion level under the target emotion category based on the first emotion level and the second emotion level; The target emotion category and the target emotion level are used as the target emotional state of the user.

5. The method according to claim 4, characterized in that The first emotion classification result further includes a first confidence score of the first emotion category, and the second emotion classification result further includes a second confidence score of the second emotion category; The determining the target emotional state of the user based on the first emotion classification result and the second emotion classification result includes: If the first emotion category and the second emotion category are different, determining the maximum value of the first confidence level and the second confidence level as a target confidence level; The emotion category and emotion level included in the emotion classification result corresponding to the target confidence are used as the target emotional state of the user.

6. The method according to claim 1, characterized in that The performing emotion classification on the voice instruction to obtain a first emotion classification result includes: Performing emotion classification on the voice command using a voice analysis model to obtain a first emotion classification result; The training process of the speech analysis model includes: Acquire training speech and first emotion category labeling information of the training speech; Extracting features from the training speech using the speech analysis model to obtain speech features; and The speech analysis model performs emotion classification based on the speech features to obtain a first predicted emotion classification result; Calculating a second loss based on the first predicted emotion classification result and the first emotion category labeling information; The model parameters of the speech analysis model are adjusted based on the second loss until a second training end condition is reached.

7. The method according to claim 1, characterized in that The performing emotion classification on the user image to obtain a second emotion classification result includes: Performing emotion classification on the user image using an image analysis model to obtain a second emotion classification result; The training process of the image analysis model includes: Acquire a training image and second emotion category labeling information of the training image; Extracting features from the training image using the image analysis model to obtain image features; and The image analysis model performs emotion classification based on the image features to obtain a second predicted emotion classification result; Calculating a third loss based on the second predicted emotion classification result and the second emotion category labeling information; The model parameters of the image analysis model are adjusted based on the third loss until a third training end condition is reached.

8. The method according to claim 7, characterized in that The step of extracting features from the training image by the image analysis model to obtain image features includes: Extracting user expression features from the training image using the image analysis model to obtain user expression features; Extracting user clothing features from the training image using the image analysis model to obtain user clothing features; Extracting user behavior features from the training image using the image analysis model to obtain user behavior features; The image analysis model performs feature fusion on the user's facial expression features, the user's clothing features, and the user's behavior features to obtain image features.

9. A voice interaction device, characterized in that: include: An acquisition module, used to acquire the user's voice commands and user images; A first emotion classification module is used to perform emotion classification on the voice instruction to obtain a first emotion classification result; A second emotion classification module, configured to perform emotion classification on the user image to obtain a second emotion classification result; an emotion determination module, configured to determine a target emotional state of the user based on the first emotion classification result and the second emotion classification result; An output module is used to generate a reply speech in a target speech style according to the target emotional state and the reply text to the voice instruction by a speech generation model; wherein the target speech style is a speech style adapted to the target emotional state; and the speech style adapted to the target emotional state is used to alleviate the negative emotion when the target emotional state indicates a negative emotion.

10. An electronic device, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 8 is implemented.

11. A computer-readable storage medium, characterized in that Computer-readable instructions are stored thereon, and when the computer-readable instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Voice emotion interaction method, computer equipment and computer readable storage medium

    CN110085221A

  • Method and device for processing voice

    CN111883127A

  • Voice conversation method and device, electronic equipment and readable storage medium

    CN117116260A

  • Call answering method and device, electronic equipment and storage medium

    CN117612568A

  • Voice interaction method and device, equipment and storage medium

    CN119673160A