Voice interaction methods, devices, electronic devices and computer-readable storage media

By identifying and integrating the acoustic features and emotion types of user voice information, target voice information is generated, which solves the problem of insufficient user emotion recognition in intelligent voice interaction systems and improves user experience.

CN119446137BActive Publication Date: 2025-10-28TCL AIR CONDITIONER ZHONGSHAN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411377962.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-10-28
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

Existing intelligent voice interaction systems lack the ability to recognize users' emotional states and provide corresponding emotional feedback, resulting in an inadequate user experience.

Method used

By identifying the acoustic features and emotional type information of user voice information, the target voice information is generated through fusion to improve the matching and adaptability of voice interaction, including the use of emotion classification models and voice parameter fusion technology.

Benefits of technology

It improves the user experience by generating target speech that takes into account the user's emotions, thereby enhancing the emotional resonance and adaptability of voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119446137B_ABST
    Figure CN119446137B_ABST
Patent Text Reader

Abstract

This application discloses a voice interaction method, device, electronic device, and computer-readable storage medium. The method includes: in response to receiving current voice information, determining a response text for the current voice information; identifying acoustic feature information corresponding to the current voice information and identifying emotion type information corresponding to the current voice information; fusing the voice parameters corresponding to the emotion type information with the acoustic feature information to obtain fused feature information; synthesizing target voice information based on the fused feature information and the response text, and broadcasting the target voice information. This ensures that the generated target voice takes into account the user's emotions and acoustic features, thereby improving the user experience when interacting with the user based on the target voice.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing, specifically to a speech interaction method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] The development of intelligent voice interaction has greatly improved the convenience of daily life. It allows for feedback to user input through voice responses, enabling interaction with users via voice.

[0003] However, most current intelligent voice responses are still mechanical, lacking the ability to recognize users' emotional states and provide corresponding emotional feedback, which means the user experience still needs improvement. Summary of the Invention

[0004] This application provides a voice interaction method, device, electronic device, and computer-readable storage medium that can take user emotions into account during voice interaction processing, thereby improving user experience.

[0005] In a first aspect, embodiments of this application provide a voice interaction method, the method comprising:

[0006] In response to receiving current voice information, determine the response text for the current voice information;

[0007] Identify the acoustic feature information corresponding to the current speech information, and identify the emotion type information corresponding to the current speech information;

[0008] The speech parameters corresponding to the emotion type information are fused with the acoustic feature information to obtain fused feature information;

[0009] Based on the fused feature information and the response text, target speech information is synthesized, and the target speech corresponding to the target speech information is broadcast.

[0010] Secondly, embodiments of this application also provide a voice interaction device, the device comprising:

[0011] The response module is used to determine the response text for the received current voice information in response to the current voice information.

[0012] The recognition module is used to recognize the acoustic feature information corresponding to the current voice information and the emotion type information corresponding to the current voice information;

[0013] The fusion module is used to fuse the speech parameters corresponding to the emotion type information with the acoustic feature information to obtain fused feature information;

[0014] The synthesis module is used to synthesize target speech information based on the fused feature information and the response text, and to broadcast the target speech corresponding to the target speech information.

[0015] Thirdly, embodiments of this application also provide an electronic device, which includes an air conditioner. The air conditioner includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the above-described voice interaction method.

[0016] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in the aforementioned voice interaction method.

[0017] Fifthly, embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described in embodiments of this application.

[0018] In this embodiment of the application, in response to receiving current voice information, the system determines the response text for the current voice information, identifies the acoustic feature information corresponding to the current voice information, identifies the emotion type information corresponding to the current voice information, fuses the voice parameters corresponding to the emotion type information with the acoustic feature information to obtain fused feature information, synthesizes target voice information based on the fused feature information and the response text, and broadcasts the target voice corresponding to the target voice information.

[0019] Specifically, by combining the user's emotional type to generate target speech based on the response text, the generated target speech takes the user's emotions into account, thereby improving the user experience when interacting with the user based on this target speech. Furthermore, by combining the acoustic features of the user's current voice input to generate the target speech, the matching degree between the target speech and the user's current voice information is improved, enhancing the adaptability of the target speech to the user. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1This is a schematic diagram of a scenario for the voice interaction method provided in an embodiment of this application;

[0022] Figure 2 This is a flowchart illustrating the voice interaction method provided in an embodiment of this application;

[0023] Figure 3 This is a schematic diagram of the structure of the voice interaction device provided in the embodiments of this application;

[0024] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0025] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] This application provides a voice interaction method, apparatus, electronic device, and computer-readable storage medium. Specifically, this application provides a voice interaction apparatus suitable for electronic devices, used to process the voice output by the device during voice interaction to obtain target voice that matches the user's emotion type and acoustic characteristics. The electronic device includes a terminal device or a server. The terminal device includes, but is not limited to, devices with voice interaction capabilities such as air conditioners, televisions, refrigerators, washing machines, and stereos. For example, a voice module can be built into or integrated into an air conditioner to process and provide feedback on the user's input voice, so as to interact with the user through voice. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The server can be directly or indirectly connected through wired or wireless communication.

[0027] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating a scenario in which a terminal device, according to an embodiment of this application, executes the voice interaction method. The specific execution process of the voice interaction method by the terminal device is as follows:

[0028] In response to receiving current voice information, terminal device 10 determines the response text for the current voice information, identifies the acoustic feature information corresponding to the current voice information, and identifies the emotion type information corresponding to the current voice information. It then fuses the voice parameters corresponding to the emotion type information with the acoustic feature information to obtain fused feature information. Based on the fused feature information and the response text, it synthesizes target voice information and broadcasts the target voice information.

[0029] In this embodiment of the application, the current voice information is the user's current voice information, such as the wake word, question text, or operation command input by the user via voice.

[0030] It is understood that the terminal device in this application embodiment generates target speech based on the response text by combining the user's emotional type, thus taking the user's emotions into account in the generated target speech, thereby improving the user experience when interacting with the user based on the target speech. Specifically, by combining the acoustic features of the user's current voice input information to generate the target speech, the matching degree between the target speech and the user's current voice information is improved, and the adaptability between the target speech and the user is enhanced.

[0031] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the priority of the embodiments.

[0032] Please see Figure 2 , Figure 2 This is a flowchart illustrating the voice interaction method provided in an embodiment of this application. Although a logical sequence is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown in the figures.

[0033] 101. In response to receiving current voice information, determine the response text for the current voice information.

[0034] It should be noted that the response text is the text of the voice content of the current voice information, used to respond to the voice content of the current voice information. For example, the response text is the output text of a simulated customer service dialogue.

[0035] Understandably, the response text can be retrieved from a text library. For example, based on the speech content corresponding to the current speech information, the response text corresponding to that speech content can be matched from the text library. The text library pre-stores the response texts corresponding to each question text, and the matching of response texts is based on the matching results between the speech content and the question texts.

[0036] In some feasible solutions, the response text can also be obtained by using a large model (LLM). For example, the speech content corresponding to the current speech information can be input into the large model, and the large model can output the response text for the speech content.

[0037] Since voice interaction is usually achieved through multiple dialogues, the context of the dialogue can also be considered when determining the response text.

[0038] 102. Identify the acoustic feature information corresponding to the current voice information, and identify the emotion type information corresponding to the current voice information.

[0039] It should be noted that the acoustic feature information includes scoring parameter information for the acoustic features. The acoustic features include pitch, volume, speech rate, timbre, etc. The scoring parameter information refers to the score or value of the current speech information for each acoustic feature, such as a volume of 80 decibels and a speech rate of 300 words per minute.

[0040] Among them, emotion type information refers to the type of emotion reflected in the user's voice, such as the user's emotion type including depressed, anxious, fearful, excited, or peaceful.

[0041] 103. The speech parameters corresponding to the emotion type information are fused with the acoustic feature information to obtain fused feature information.

[0042] It should be noted that each emotion type corresponds to different speech parameters. Furthermore, these speech parameters are scores or values ​​for each acoustic feature specific to each emotion type.

[0043] In this process, fusion feature information is obtained through fusion, which enables the comprehensive representation of a user's acoustic characteristics and emotional type.

[0044] 104. Synthesize target speech information based on the fused feature information and the response text, and broadcast the target speech corresponding to the target speech information.

[0045] In general scenarios, speech can be synthesized directly from the response text. However, since the speech generated solely from the response text lacks consideration for the user's emotions, the output speech is monotonous and fails to resonate with the user's emotions, resulting in a need to improve the user experience.

[0046] In this embodiment of the application, the target speech information corresponding to the response text is generated based on the fusion feature information. This makes the target speech information take into account the user's emotional type and the user's acoustic features, thereby improving the matching degree or adaptability between the target speech and the user's emotions and acoustic features, and improving the user experience.

[0047] Optionally, in some feasible solutions, the emotion type information corresponding to the current voice information can be identified and classified using an emotion classification model. That is, optionally, in some embodiments of this application, the step of "identifying the emotion type information corresponding to the current voice information" includes:

[0048] The emotion classification model outputs the emotion type information corresponding to the current voice information.

[0049] The emotion classification model is obtained by training the original classification model, and the base network of the original classification includes at least one of convolutional neural networks or long short-term memory networks.

[0050] It should be noted that the emotion classification model is a machine learning model. It uses machine learning to classify and identify the emotion type information corresponding to the current voice information, thereby improving recognition efficiency.

[0051] Among them, convolutional neural networks (CNN) or long short-term memory networks (LSTM) have shown remarkable performance in the classification field. Therefore, choosing to train an emotion classification model based on convolutional neural networks (CNN) or long short-term memory networks (LSTM) can help improve the accuracy of emotion classification.

[0052] Optionally, in some scenarios, to reduce the complexity of the model structure, features can be extracted from the current speech information first, and the extracted features can be used as input to the emotion classification model. This allows the model to directly process the extracted features without performing feature extraction or other tasks, thus reducing the model's complexity. Specifically, in some embodiments of this application, the step "outputting the emotion type information corresponding to the current speech information through the emotion classification model" includes:

[0053] Speech features are extracted from the current speech information to obtain first speech feature information;

[0054] The first speech feature information is input into the emotion classification model, so that the emotion classification model outputs the emotion type information corresponding to the current speech information.

[0055] It should be noted that the first speech feature information reflects the characteristics of the current speech information. For example, the first speech feature information includes pitch, volume, speech rate, timbre, etc. In this embodiment of the application, the speech feature extraction can be achieved through speech processing technology, for example, by using the Mel-frequency cepstral coefficient (MFCC) algorithm to extract the features of the current speech information to obtain the first speech feature information.

[0056] The process involves inputting the first speech feature information into the emotion classification model, which then outputs the confidence level of the first speech feature information for each emotion type. Based on the confidence level, the emotion type corresponding to the highest confidence level is selected as the emotion type information corresponding to the current speech information.

[0057] It is understood that the emotion classification model can be trained using sample data. For example, the original classification model can be trained using a general or publicly available emotion speech dataset to obtain the emotion classification model. That is, optionally, in some embodiments of this application, before the step "inputting the first speech feature information into the emotion classification model to output the emotion type information corresponding to the current speech information through the emotion classification model", the method further includes:

[0058] Obtain general sample information, which includes general sample speech information and sample emotion labels corresponding to the general sample speech information;

[0059] An emotion classification model is obtained by training the original classification model using the general sample speech information and the corresponding sample emotion labels.

[0060] Among them, the sample emotion label is the sample emotion classification result, such as the emotion labels such as low mood and excitement corresponding to general sample voice information.

[0061] Among them, the general sample information is the samples in the publicly available emotion speech dataset. The original classification model is trained using the general samples, so that the trained emotion classification model has a general emotion classification ability and can be applied to general emotion classification tasks.

[0062] Optionally, relevant speech data can be collected and labeled according to specific application scenarios, and this data can be used to fine-tune the model. The goal of fine-tuning is to enable the model to better understand and recognize emotional expressions in specific application scenarios.

[0063] Optionally, in addition to training the original classification model using publicly available or general sample data, to better suit the emotion classification of the target user, personalized voice information corresponding to different emotion types of the target user can be collected. Based on this personalized voice information for different emotion types, the emotion classification model trained on general samples can be retrained or optimized to improve the accuracy of the model in classifying the target user's emotions. That is, optionally, in some embodiments of this application, after the step "training the original classification model using the general sample voice information and the corresponding sample emotion labels to obtain an emotion classification model," the method further includes:

[0064] Obtain personalized voice information corresponding to different emotion types of target users;

[0065] The personalized voice information is subjected to feature extraction using a feature extraction model to obtain the second voice feature information corresponding to the personalized voice information.

[0066] The emotion classification model is optimized by using the second voice feature information of each emotion type of the target user to obtain the optimized emotion classification model.

[0067] The step of inputting the first speech feature information into the emotion classification model, so as to output the emotion type information corresponding to the current speech information through the emotion classification model, includes:

[0068] The first speech feature information is input into the optimized emotion classification model, so that the optimized emotion classification model outputs the emotion type information corresponding to the current speech information.

[0069] It is understandable that by combining the personalized voice information of the target user in different emotion types to optimize the model, the accuracy of the optimized emotion classification model in analyzing the target user's emotion type task can be improved.

[0070] Optionally, to improve the accuracy of the extraction of the second speech feature information, the personalized speech information can be preprocessed first, and the preprocessed personalized speech information can be input into the feature extraction model to extract the second speech feature information.

[0071] Preprocessing includes noise reduction and normalization. Noise reduction involves identifying and filtering noise from personalized speech information. Normalization includes amplitude normalization or spectrum normalization. Amplitude normalization normalizes the amplitude of personalized speech information to a fixed range, such as the interval [-1, 1]. The specific calculation steps are: calculate the maximum value (the maximum absolute value of the personalized speech information), divide the value of each sampling point by the maximum value, and obtain the normalized signal, thus completing amplitude normalization. Spectrum normalization refers to normalizing the spectrum of personalized speech information to eliminate the influence of recording equipment and environment. The calculation steps include: processing the personalized speech information through Short Time Fourier Transform (STFT) to obtain a spectrum, normalizing each frequency channel in the spectrum to make its mean 0 and standard deviation 1, and inverse STFT processing (converting the normalized spectrum back to a time-series signal through STFT).

[0072] Optionally, to improve the accuracy of the second speech feature information extraction, long-distance dependencies between features are considered. That is, optionally, in some embodiments of this application, the feature extraction model is obtained through the following operations:

[0073] Feature extraction of sample audio information is performed using a convolutional neural network to generate preliminary feature representations;

[0074] The preliminary feature representations are context-encoded using a temporal feature model to capture long-distance dependencies between the preliminary feature representations;

[0075] Randomly mask out a portion of the sample audio information to determine the masked audio segment;

[0076] Based on the long-distance dependency, the masked audio segment is predicted to obtain the predicted audio segment;

[0077] A feature extraction model is obtained by training the base model based on the masked audio segment and the predicted audio segment.

[0078] For example, a convolutional neural network (CNN) is used to extract features from personalized speech information to obtain a preliminary feature representation. This preliminary feature representation is then context-encoded using an attention mechanism model (such as the Transformer model) or a temporal feature model (such as a Long Short-Term Memory (LSTM) network) to obtain long-distance dependencies. Then, self-supervised learning is used to predict the masked portion of the audio signal, and the model is trained based on the predicted audio signal and the masked signal. The base network is a network suitable for audio feature extraction, which includes, but is not limited to, convolutional neural networks, recurrent neural networks, autoencoders, or attention-based networks.

[0079] Specifically, the self-supervised learning in this scheme involves randomly masking audio data for a certain period of time in the input audio signal. The model needs to predict the masked part using contextual information. If the prediction is correct, the model deepens its understanding of the context; if it is incorrect, the model corrects the prediction. In this way, the model learns the rich features and structural information of the audio signal.

[0080] Optionally, since both speech parameters and acoustic feature information contain values ​​or scores for acoustic features, the fusion of speech parameters and acoustic feature information can be achieved based on weights. That is, optionally, in some embodiments of this application, the step "fusion of the speech parameters corresponding to the emotion type information with the acoustic feature information to obtain fused feature information" includes:

[0081] Determine the voice parameters corresponding to the emotion type information;

[0082] A first value for each acoustic feature is determined from the speech parameters, and a second value for each acoustic feature is determined from the acoustic feature information;

[0083] For each acoustic feature, the fusion value of the first and second values ​​of each acoustic feature is calculated according to the preset weights;

[0084] Fusion feature information is generated based on the fusion values ​​corresponding to each acoustic feature.

[0085] Understandably, weight-based fusion schemes ensure the accuracy and effectiveness of fused feature information.

[0086] Understandably, based on the output target speech, a feedback mechanism can be configured to receive user feedback, retrieve the user's satisfaction level, and feed it back to the emotion classification model and feature extraction model, and adjust the preset weights to optimize the output target speech.

[0087] In summary, this application embodiment generates target speech based on the response text by combining the user's emotional type, thus taking the user's emotions into account. This improves the user experience when interacting with the user based on the target speech. Specifically, by combining the acoustic features of the user's current voice input to generate the target speech, the matching degree between the target speech and the user's current voice information is improved, enhancing the adaptability of the target speech to the user.

[0088] By combining personalized voice information from users in different emotional states to optimize the model, the accuracy of the model in classifying the emotions of the target users can be improved.

[0089] To facilitate better implementation of the voice interaction method of this application, this application also provides a voice interaction device based on the above-described voice interaction method. The meanings of the terms used are the same as in the above-described voice interaction method, and specific implementation details can be found in the description of the method embodiments.

[0090] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of the voice interaction device provided in the embodiment of this application, wherein the voice interaction device may specifically be as follows:

[0091] Response module 201 is used to determine a response text for the current voice information in response to receiving the current voice information;

[0092] The recognition module 202 is used to recognize the acoustic feature information corresponding to the current voice information and the emotion type information corresponding to the current voice information;

[0093] The fusion module 203 is used to fuse the speech parameters corresponding to the emotion type information with the acoustic feature information to obtain fused feature information;

[0094] The synthesis module 204 is used to synthesize target speech information based on the fused feature information and the response text, and to broadcast the target speech corresponding to the target speech information.

[0095] Optionally, in some embodiments of this application, identifying the emotion type information corresponding to the current voice information includes:

[0096] The emotion classification model outputs the emotion type information corresponding to the current voice information.

[0097] The emotion classification model is obtained by training the original classification model.

[0098] Optionally, in some embodiments of this application, the step of outputting the emotion type information corresponding to the current voice information through the emotion classification model includes:

[0099] Speech features are extracted from the current speech information to obtain first speech feature information;

[0100] The first speech feature information is input into the emotion classification model, so that the emotion classification model outputs the emotion type information corresponding to the current speech information.

[0101] Optionally, in some embodiments of this application, before inputting the first speech feature information into the emotion classification model to output the emotion type information corresponding to the current speech information through the emotion classification model, the method further includes:

[0102] Obtain general sample information, which includes general sample speech information and sample emotion labels corresponding to the general sample speech information;

[0103] An emotion classification model is obtained by training the original classification model using the general sample speech information and the corresponding sample emotion labels.

[0104] Optionally, in some embodiments of this application, after training the original classification model to obtain an emotion classification model using the general sample speech information and the corresponding sample emotion labels, the method further includes:

[0105] Obtain personalized voice information corresponding to different emotion types of target users;

[0106] The personalized voice information is subjected to feature extraction using a feature extraction model to obtain the second voice feature information corresponding to the personalized voice information.

[0107] The emotion classification model is optimized by using the second voice feature information of each emotion type of the target user to obtain the optimized emotion classification model.

[0108] The step of inputting the first speech feature information into the emotion classification model, so as to output the emotion type information corresponding to the current speech information through the emotion classification model, includes:

[0109] The first speech feature information is input into the optimized emotion classification model, so that the optimized emotion classification model outputs the emotion type information corresponding to the current speech information.

[0110] Optionally, in some embodiments of this application,

[0111] The feature extraction model is obtained through the following operations:

[0112] Feature extraction of sample audio information is performed using a convolutional neural network to generate preliminary feature representations;

[0113] The preliminary feature representations are context-encoded using a temporal feature model to capture long-distance dependencies between the preliminary feature representations;

[0114] Randomly mask out a portion of the sample audio information to determine the masked audio segment;

[0115] Based on the long-distance dependency, the masked audio segment is predicted to obtain the predicted audio segment;

[0116] A feature extraction model is obtained by training the base model based on the masked audio segment and the predicted audio segment.

[0117] Optionally, in some embodiments of this application, fusing the speech parameters corresponding to the emotion type information with the acoustic feature information to obtain fused feature information includes:

[0118] Determine the voice parameters corresponding to the emotion type information;

[0119] A first value for each acoustic feature is determined from the speech parameters, and a second value for each acoustic feature is determined from the acoustic feature information;

[0120] For each acoustic feature, the fusion value of the first and second values ​​of each acoustic feature is calculated according to the preset weights;

[0121] Fusion feature information is generated based on the fusion values ​​corresponding to each acoustic feature.

[0122] In this embodiment, the response module 201 first determines the response text for the received current voice information in response to the current voice information. Then, the recognition module 202 recognizes the acoustic feature information and the emotion type information corresponding to the current voice information. Subsequently, the fusion module 203 fuses the voice parameters corresponding to the emotion type information with the acoustic feature information to obtain fused feature information. Then, the synthesis module 204 synthesizes target voice information based on the fused feature information and the response text, and broadcasts the target voice corresponding to the target voice information.

[0123] In this embodiment, the target speech generated based on the response text is combined with the user's emotional type. This ensures that the generated target speech takes the user's emotions into account, thereby improving the user experience when interacting with the user based on the target speech. Furthermore, by combining the acoustic features of the user's current voice input to generate the target speech, the matching degree between the target speech and the user's current voice information is improved, thus enhancing the adaptability of the target speech to the user.

[0124] In addition, this application also provides an electronic device, such as Figure 4 As shown, it illustrates the structural diagram of the electronic device involved in this application, specifically:

[0125] The electronic device may include components such as a processor 301 with one or more processing cores, a memory 302 with one or more computer-readable storage media, a power supply 303, and an input unit 304. Those skilled in the art will understand that... Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0126] The processor 301 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 302, and by calling data stored in the memory 302, thereby providing overall monitoring of the electronic device. Optionally, the processor 301 may include one or more processing cores; preferably, the processor 301 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 301.

[0127] The memory 302 can be used to store software programs and modules. The processor 301 executes various functional applications and data processing by running the software programs and modules stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 302 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 302 may also include a memory controller to provide the processor 301 with access to the memory 302.

[0128] The electronic device also includes a power supply 303 that supplies power to the various components. Preferably, the power supply 303 can be logically connected to the processor 301 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 303 may also include one or more DC or AC power supplies, recharging systems, power equipment debugging circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0129] The electronic device may also include an input unit 304, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0130] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 301 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 302 according to the following instructions, and the processor 301 runs the applications stored in the memory 302, thereby implementing the steps in any of the voice interaction methods provided in the embodiments of this application.

[0131] In this embodiment of the application, in response to receiving current voice information, the system determines the response text for the current voice information, identifies the acoustic feature information corresponding to the current voice information, identifies the emotion type information corresponding to the current voice information, fuses the voice parameters corresponding to the emotion type information with the acoustic feature information to obtain fused feature information, synthesizes target voice information based on the fused feature information and the response text, and broadcasts the target voice corresponding to the target voice information.

[0132] Specifically, by combining the user's emotional type to generate target speech based on the response text, the generated target speech takes the user's emotions into account, thereby improving the user experience when interacting with the user based on this target speech. Furthermore, by combining the acoustic features of the user's current voice input to generate the target speech, the matching degree between the target speech and the user's current voice information is improved, enhancing the adaptability of the target speech to the user.

[0133] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0134] It is understood that the electronic device in this application embodiment is applicable to an air conditioner. By executing the above method through the air conditioner, the target user can be interacted with via voice to achieve the above effect.

[0135] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0136] Therefore, this application provides a computer-readable storage medium storing a computer program that can be loaded by a processor to execute the steps of any of the voice interaction methods provided in this application.

[0137] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0138] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0139] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the voice interaction methods provided in this application, the beneficial effects that any of the voice interaction methods provided in this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.

[0140] The above provides a detailed description of a voice interaction method, apparatus, electronic device, and computer-readable storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A voice interaction method, characterized in that, The method includes: In response to receiving current voice information, determine the response text for the current voice information; Identify the acoustic feature information corresponding to the current speech information, and identify the emotion type information corresponding to the current speech information; The speech parameters corresponding to the emotion type information are fused with the acoustic feature information to obtain fused feature information; Based on the fused feature information and the response text, target speech information is synthesized, and the target speech corresponding to the target speech information is broadcast. The step of fusing the speech parameters corresponding to the emotion type information with the acoustic feature information to obtain fused feature information includes: Determine the voice parameters corresponding to the emotion type information; A first value for each acoustic feature is determined from the speech parameters, and a second value for each acoustic feature is determined from the acoustic feature information; For each acoustic feature, the fusion value of the first and second values ​​of each acoustic feature is calculated according to the preset weights; Fusion feature information is generated based on the fusion values ​​corresponding to each acoustic feature.

2. The voice interaction method according to claim 1, characterized in that, The identification of the emotion type information corresponding to the current voice information includes: The emotion classification model outputs the emotion type information corresponding to the current voice information. The emotion classification model is obtained by training the original classification model.

3. The voice interaction method according to claim 2, characterized in that, The step of outputting the emotion type information corresponding to the current voice information through the emotion classification model includes: Speech features are extracted from the current speech information to obtain first speech feature information; The first speech feature information is input into the emotion classification model, so that the emotion classification model outputs the emotion type information corresponding to the current speech information.

4. The voice interaction method according to claim 3, characterized in that, Before inputting the first speech feature information into the emotion classification model to output the emotion type information corresponding to the current speech information through the emotion classification model, the method further includes: Obtain general sample information, which includes general sample speech information and sample emotion labels corresponding to the general sample speech information; An emotion classification model is obtained by training the original classification model using the general sample speech information and the corresponding sample emotion labels.

5. The voice interaction method according to claim 4, characterized in that, After training the original classification model to obtain an emotion classification model using the general sample speech information and the corresponding sample emotion labels, the method further includes: Obtain personalized voice information corresponding to different emotion types of target users; The personalized voice information is subjected to feature extraction using a feature extraction model to obtain the second voice feature information corresponding to the personalized voice information. The emotion classification model is optimized by using the second voice feature information of each emotion type of the target user to obtain the optimized emotion classification model. The step of inputting the first speech feature information into the emotion classification model, so as to output the emotion type information corresponding to the current speech information through the emotion classification model, includes: The first speech feature information is input into the optimized emotion classification model, so that the optimized emotion classification model outputs the emotion type information corresponding to the current speech information.

6. The voice interaction method according to claim 5, characterized in that, The feature extraction model is obtained through the following operations: Feature extraction of sample audio information is performed using a convolutional neural network to generate preliminary feature representations; The preliminary feature representations are context-encoded using a temporal feature model to capture long-distance dependencies between the preliminary feature representations; Randomly mask out a portion of the sample audio information to determine the masked audio segment; Based on the long-distance dependency, the masked audio segment is predicted to obtain the predicted audio segment; A feature extraction model is obtained by training the base model based on the masked audio segment and the predicted audio segment.

7. A voice interaction device, characterized in that, The device includes: The response module is used to determine the response text for the received current voice information in response to the current voice information. The recognition module is used to recognize the acoustic feature information corresponding to the current voice information and the emotion type information corresponding to the current voice information; The fusion module is used to fuse the speech parameters corresponding to the emotion type information with the acoustic feature information to obtain fused feature information; The synthesis module is used to synthesize target speech information based on the fused feature information and the response text, and to broadcast the target speech corresponding to the target speech information; The step of fusing the speech parameters corresponding to the emotion type information with the acoustic feature information to obtain fused feature information includes: Determine the voice parameters corresponding to the emotion type information; A first value for each acoustic feature is determined from the speech parameters, and a second value for each acoustic feature is determined from the acoustic feature information; For each acoustic feature, the fusion value of the first and second values ​​of each acoustic feature is calculated according to the preset weights; Fusion feature information is generated based on the fusion values ​​corresponding to each acoustic feature.

8. An electronic device, characterized in that, The device includes an air conditioner, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the voice interaction method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the voice interaction method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Voice interaction method and device, electronic equipment and storage medium

    CN115329057A

  • Voice conversation method and device, electronic equipment and readable storage medium

    CN117116260A