Voice interaction method, server and computer readable storage medium
By analyzing the acoustic features of voice requests and using a neural network model to identify emotional states and generate adaptive voice messages, the problem of insufficient user emotion recognition in the vehicle cabin is solved, improving user experience and the immediacy and relevance of feedback.
Patent Information
- Application Number
- CN202511397230.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2025-12-02
AI Technical Summary
In the context of vehicle cabins, existing technologies cannot recognize emotional information in users' voice requests, resulting in a poor user experience. Furthermore, when processing complex requests, there are inference delays and waiting gaps, which can easily cause users to feel anxious.
By analyzing the acoustic features of voice requests, a pre-trained neural network model is used to identify emotional state identifiers and generate emotion-adaptive filler messages to fill the waiting period and provide emotional feedback, thereby improving the user experience.
By identifying users' emotional states and generating tailored responses, we can establish emotional connections, reduce waiting time, improve user satisfaction and trust, and ensure the timeliness and relevance of feedback.
Smart Images

Figure CN121053986A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice interaction, and in particular to a voice interaction method, a server, and a computer-readable storage medium. Background Technology
[0002] In related technologies, within a vehicle cockpit setting, when a user issues a voice request, a large language model can directly output inference results that meet the user's needs based on the text recognition results of the voice request, thereby generating corresponding vehicle control commands that satisfy the user's requirements. However, this approach ignores the user's emotional expression in the voice request, fails to recognize the user's emotional information and provide emotional feedback, resulting in a poor user experience. Summary of the Invention
[0003] This application provides a voice interaction method, a server, and a computer-readable storage medium.
[0004] This application provides a voice interaction method, the method comprising: Based on the acoustic features of the received voice request, an emotional state identifier is determined, wherein the acoustic features include at least one of pitch features, speech rate features, and energy features; Based on the emotional state identifier and the voice request, an emotion-adaptive backing message is generated to complete the voice interaction.
[0005] Thus, the server determines an emotional state identifier based on the acoustic features of the received voice request. These acoustic features include at least one of pitch, speech rate, and energy characteristics. Next, based on the emotional state identifier and the voice request, an emotion-adaptive backing message is generated to complete the voice interaction. In this way, by analyzing acoustic features, identifying emotional state identifiers, and accurately generating emotion-adaptive backing messages corresponding to the voice request, an emotional connection is established with the user, enhancing the user experience. Furthermore, by generating emotion-adaptive backing messages, the waiting period between the user issuing a command and the system returning the final result can be filled, creating an "instant response" experience and preventing user anxiety due to long waiting times, thereby increasing user satisfaction and trust in the system.
[0006] In some implementations, determining the emotional state identifier based on the acoustic features of the received voice request includes: Based on a pre-trained neural network model, the emotional state identifier is determined according to the acoustic features. The neural network model includes an input layer, a hidden layer, and an output layer. The input layer is used to determine the target acoustic feature signal from the acoustic features and input it to the hidden layer. The hidden layer implicitly captures emotional information based on the target acoustic feature signal and inputs it to the output layer. The output layer outputs the emotional state identifier based on the emotional information.
[0007] Thus, based on the pre-trained neural network model, emotional state identifiers are determined according to acoustic features. The neural network model includes an input layer, a hidden layer, and an output layer. The input layer determines the target acoustic feature signal from the acoustic features and inputs it to the hidden layer. The hidden layer implicitly captures emotional information based on the target acoustic feature signal and inputs it to the output layer. The output layer outputs the emotional state identifier based on the emotional information. In this way, the pre-trained neural network model can accurately identify different emotional state identifiers, providing a basis for subsequently generating appropriate emotionally-adaptive background dialogue.
[0008] In some implementations, generating emotion-adaptive background dialogue based on the emotion state identifier and the voice request includes: Natural language processing is performed on the voice request to determine the user's intent; Based on the emotional state identifier and the user intent, generate an emotion-adaptive background message.
[0009] In this way, natural language processing is performed on the voice request to determine the user's intent. Then, based on the emotion state identifier and the user's intent, emotion-adapted background dialogue is generated. By determining the user's intent and combining it with the emotion state identifier and the user's intent, the generated emotion-adapted background dialogue can be anchored to the user's actual needs, avoiding the generation of irrelevant content based solely on emotions, and ensuring that the emotion-adapted background dialogue is strongly relevant to the user's goals.
[0010] In some implementations, performing natural language processing on the voice request to determine the user's intent includes: Based on the voice request, a first voice request segment received at a first moment and a second voice request segment received at a second moment that follows the first voice request segment are determined sequentially, wherein the second moment is later than the first moment; Perform speech recognition processing on the first voice request segment to determine the first recognition result; Perform natural language understanding processing on the first recognition result to determine the first understanding result; The second voice request segment is subjected to speech recognition processing to determine the second recognition result; The second recognition result is processed by natural language understanding to determine the second understanding result; Based on the temporal relationship between the first understanding result and the second understanding result, the first understanding result and the second understanding result are logically fused to determine the user intent.
[0011] Thus, based on the voice request, the server sequentially determines the first voice request segment received at a first moment and the second voice request segment received at a second moment, which follows the first voice request segment, with the second moment being later than the first moment. Next, speech recognition processing is performed on the first voice request segment to determine the first recognition result. Then, natural language understanding processing is performed on the first recognition result to determine the first understanding result. Subsequently, speech recognition processing is performed on the second voice request segment to determine the second recognition result. Next, natural language understanding processing is performed on the second recognition result to determine the second understanding result. Finally, based on the temporal relationship between the first and second understanding results, logical fusion processing is performed on the first and second understanding results to determine the user's intent. In this way, by segmenting the voice request into multiple voice request segments and processing them in a streaming manner, the user does not need to wait for the model to complete speech recognition of the complete voice request before natural language understanding, which shortens the model response time and improves the user experience.
[0012] In some implementations, the emotion-adaptive pre-recorded message includes at least one of reassuring pre-recorded messages, empathetic pre-recorded messages, and lead-in pre-recorded messages. Generating the emotion-adaptive pre-recorded message based on the emotion state identifier and the user intent includes: When the emotional state identifier is a first emotional identifier and the user intent is a need-based user intent, the reassuring words are generated, wherein the first emotional identifier is used to indicate that the user is in a peaceful state with emotional fluctuations less than a preset emotional threshold; or When the emotional state identifier is a second emotional identifier and the user intent is a solution-oriented user intent, the empathic preamble is generated, wherein the second emotional identifier is used to indicate a negative state where the emotional fluctuation is greater than or equal to the preset emotional threshold; or When the emotional state identifier is the first emotional identifier and the user intent is the solution-type user intent, the lead-in preamble is generated.
[0013] Thus, when the emotional state identifier is the first emotional identifier and the user intent is a need-based user intent, a soothing interlude is generated. The first emotional identifier indicates that the user is in a peaceful state with emotional fluctuations below a preset emotional threshold. Alternatively, when the emotional state identifier is the second emotional identifier and the user intent is a solution-based user intent, an empathetic interlude is generated. The second emotional identifier indicates a negative state with emotional fluctuations greater than or equal to a preset emotional threshold. Or, when the emotional state identifier is the first emotional identifier and the user intent is a solution-based user intent, a leading interlude is generated. This ability to generate different emotion-adaptive interludes for different scenarios makes voice interaction more natural and human-like, improving user satisfaction with the voice interaction experience.
[0014] In some embodiments, the method further includes: After generating the emotion-adaptive background message, the emotion-adaptive background message is sent to the first target electronic device so that the emotion-adaptive background message can be broadcast through the first target electronic device.
[0015] Thus, after generating the emotion-adaptive backing message, it is sent to the first target electronic device for broadcast. This rapid broadcast of the generated backing message through the target electronic device quickly delivers feedback to the user during speech recognition, semantic understanding, and command distribution and execution. It fills the gap between the user's voice request and the return of the formal result, preventing user anxiety due to prolonged waiting, improving the user's perception of response efficiency, and ultimately enhancing user satisfaction and trust.
[0016] In some embodiments, the method further includes: Generate control commands based on the user's intent; The control command is sent to the second target electronic device so that the control command is executed by the second target electronic device.
[0017] Thus, control commands are generated based on the user's intent. These commands are then sent to a second target electronic device for execution. This method of generating control commands based on the user's intent determined through natural language processing ensures a high degree of alignment between the commands and the user's actual needs.
[0018] In some embodiments, the method further includes: Receive the execution result of the control command fed back by the second target electronic device; Based on the execution result, a corresponding execution result broadcast information is generated and sent to the first target electronic device so that the first target electronic device can broadcast the execution result broadcast information after the emotion-adaptive padding broadcast is completed.
[0019] In this way, the server receives the execution result of the control command from the second target electronic device. Then, based on the execution result, it generates corresponding execution result broadcast information and sends it to the first target electronic device. This allows the first target electronic device to broadcast the execution result information after the emotion-adaptive background message has finished. In this way, receiving the execution result of the control command and generating the corresponding broadcast information accurately feeds back the final processing result of the user's needs to the user, forming a complete chain of "instant response—function execution—result confirmation" with the previous emotion-adaptive background message, thus avoiding gaps or abrupt switches between emotion-adaptive background messages and execution result feedback.
[0020] This application provides a server that includes a processor and a memory. The memory stores a computer program that, when executed by the processor, implements the method described above.
[0021] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method described above.
[0022] Additional aspects and advantages of embodiments of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of embodiments of this application. Attached Figure Description
[0023] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, wherein: Figure 1 This is one of the flowcharts illustrating a voice interaction method according to certain embodiments of this application; Figure 2 This is a schematic diagram of a conventional voice interaction process in some embodiments of this application; Figure 3 This is a schematic diagram of the voice interaction process in some embodiments of this application; Figure 4 This is a second flowchart illustrating a voice interaction method according to certain embodiments of this application; Figure 5 This is the third flowchart illustrating a voice interaction method according to certain embodiments of this application; Figure 6This is the fourth flowchart illustrating a voice interaction method according to certain embodiments of this application; Figure 7 This is the fifth flowchart illustrating a voice interaction method according to certain embodiments of this application; Figure 8 This is a flowchart of a voice interaction method according to certain embodiments of this application, number six. Figure 9 This is the seventh flowchart illustrating a voice interaction method according to certain embodiments of this application; Figure 10 This is the eighth flowchart of a voice interaction method according to certain embodiments of this application. Detailed Implementation
[0024] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of this application, and should not be construed as limiting the embodiments of this application.
[0025] In the dynamic and attention-intensive scenario of the vehicle cabin, users' voice interactions are often accompanied by complex real-time states and emotional needs. Related technologies rely on large language models to directly output inference results based on the text recognition results of voice requests, and then generate corresponding vehicle control commands.
[0026] However, user voice requests within the vehicle cabin are not simply "instruction transmissions." The emotional information contained in the voice signal (such as faster speech when anxious, higher pitch when angry, or energy fluctuations when confused) is just as important as the instruction content. Related technologies typically rely on the transcribed text for emotion recognition when processing voice requests. This involves using explicit emotional keywords such as "annoyed," "angry," and "anxious," or relying on semantic analysis of the text to determine negative tendencies in the statements and infer the user's emotional state.
[0027] Furthermore, given the high demand for "immediacy" in cockpit scenarios, large language models often suffer from inference delays when handling complex requests (such as multi-condition location searches and function troubleshooting). The relevant technologies neither fill the waiting gaps with filler words nor alleviate waiting anxiety with emotion-adaptive feedback. After making an emotional request, users must wait in silence for the system to execute the command and return the result. This double deficiency of "ignoring emotions + waiting without feedback" can easily make users feel neglected, ultimately leading to a fragmented interactive experience, decreased user satisfaction and trust, and even abandoning the use of voice interaction functions.
[0028] Based on the above issues, please refer to Figure 1This application provides a voice interaction method, the method including: 01: Determine the emotional state identifier based on the acoustic characteristics of the received voice request; 02: Generate emotion-adaptive background dialogue based on emotion state identifiers and voice requests to complete voice interaction.
[0029] This application also provides a server, including a memory and a processor. The voice interaction method of this application can be implemented by the server of this application. Specifically, the memory stores a computer program, and the processor is used to determine an emotion state identifier based on the acoustic characteristics of the received voice request, and to generate emotion-adaptive background dialogue based on the emotion state identifier and the voice request to complete the voice interaction.
[0030] This application also provides a voice interaction device. The voice interaction method of this application can be implemented by the voice interaction device of this application. Specifically, the voice interaction device includes a determining module and a voice interaction module. The determining module is used to determine an emotion state identifier based on the acoustic characteristics of the received voice request. The voice interaction module is used to generate emotion-adaptive background dialogue based on the emotion state identifier and the voice request to complete the voice interaction.
[0031] Specifically, please refer to Figure 2 , Figure 2 This diagram illustrates the voice interaction process in related technologies. After a user issues a voice request, the intelligent cockpit system processes it in a streaming manner. This means that while receiving the voice request, the system performs voice recognition, semantic understanding, command distribution, and command execution based on its voice recognition, semantic understanding, command distribution, and command execution modules, finally providing feedback. Voice recognition refers to the intelligent cockpit system first recognizing the user's voice command and converting it into text. Semantic understanding refers to the intelligent cockpit system parsing the recognized text to understand the user's intended meaning and intent. Command distribution refers to the intelligent cockpit system sending the command to the appropriate service module for processing based on the understood intent. Command execution refers to the service module executing the command, such as querying information or controlling equipment. Feedback refers to the intelligent cockpit system providing the execution result to the user, usually in voice form. Each step is performed sequentially; the next step can only proceed after the previous one is completed. Therefore, due to the multiple steps involved, the user needs to wait a certain amount of time from issuing a voice request to receiving feedback. Furthermore, the intelligent cockpit system cannot recognize the user's emotional state and cannot provide personalized service or reassurance.
[0032] To address the aforementioned issues, this application proposes a voice interaction method. This method, based on voice perception, constructs an independent link that runs in parallel with the speech recognition, semantic understanding, and command distribution modules. Please refer to [link to relevant documentation]. Figure 3 This independent link includes a pre-defined dialogue generation module (including a preset neural network model) and a voice broadcasting module. The pre-defined dialogue generation module takes the user's voice request as the input signal and generates pre-defined dialogue content (hereinafter referred to as pre-defined dialogue content) that is highly related to the user's current voice content as its core output. Compared with related technologies, where intelligent cockpit systems only analyze user intent through the text recognition results of voice requests, ignoring the emotional information carried by acoustic features such as tone, speed, and energy in the voice, the voice interaction method proposed in this application accurately captures emotions through acoustic features, transforms abstract emotions into signals that the system can process, and thus generates emotion-adaptive pre-defined dialogue, thereby laying the foundation for emotional interaction.
[0033] Voice requests refer to instructions or questions issued to a vehicle's smart cockpit system via continuous voice signals input through acoustic sensors (such as microphone arrays). Examples include, "Find me the highest-rated seafood restaurant nearby, with parking available," "Why is the navigation stuck?" and "What are the three rings in traditional Chinese culture?"
[0034] Acoustic features refer to the features extracted from user-input voice requests that reflect the physical properties of speech and the user's expressive state. These include at least one of pitch features, speech rate features, and energy features. Pitch features refer to the highness (frequency) of the corresponding voice request; for example, a user's pitch tends to rise when angry and remain relatively stable when calm. Speech rate features refer to the speed of pronunciation (number of syllables per unit time) of the corresponding voice request; for example, a user's speech rate usually increases when anxious and may slow down when confused. Energy features refer to the strength (amplitude) of the corresponding voice request; for example, a user's voice has higher energy when emotionally agitated and more balanced energy when calm.
[0035] Emotional state identifiers refer to the distinctive information generated after acoustic feature analysis based on voice requests, used to characterize the user's current emotional type. Specifically, emotional state identifiers are a categorized and symbolic description of a user's emotions, providing a basis for judging the emotional dimension in subsequent emotion-adaptive background dialogue generation.
[0036] Emotion-adaptive pre-recorded dialogue refers to preparatory dialogue generated by combining emotional state identifiers with the specific content of the voice request, possessing emotional matching and content relevance. Specifically, emotion-adaptive pre-recorded dialogue matches the user's "emotional state" and "needs" when needed, and serves as parallel feedback in the voice interaction process of related technologies, filling the gaps in the user's waiting time, while also providing a connection for the subsequent formal result broadcast.
[0037] After a user issues a voice request, the intelligent cockpit system determines an emotional state identifier based on the acoustic characteristics of the received voice request. Subsequently, based on the emotional state identifier and the voice request, it generates an emotion-adaptive dialogue to complete the voice interaction.
[0038] In summary, in the voice interaction method and server provided in this application, the server determines an emotional state identifier based on the acoustic features of the received voice request, wherein the acoustic features include at least one of pitch features, speech rate features, and energy features. Then, based on the emotional state identifier and the voice request, an emotion-adaptive backing message is generated to complete the voice interaction. Thus, by analyzing the acoustic features, identifying the emotional state identifier, and accurately generating an emotion-adaptive backing message corresponding to the voice request, an emotional connection is established with the user, enhancing the user experience. Furthermore, by generating an emotion-adaptive backing message, the waiting period between the user issuing a command and the system returning the formal result can be filled, creating an "instant response" experience, avoiding user anxiety due to long waiting times, and thereby increasing user satisfaction and trust in the system.
[0039] Please see Figure 4 In some implementations, step 01 (determining an emotional state identifier based on the acoustic characteristics of the received voice request) includes: 011: Based on the pre-trained neural network model, determine the emotional state identifier according to acoustic features.
[0040] In some implementations, the determining module is also used to determine an emotional state identifier based on acoustic features, using a pre-trained neural network model.
[0041] In some implementations, the processor is also used to determine emotional state identifiers based on acoustic features from a pre-trained neural network model.
[0042] Specifically, a pre-trained neural network model refers to a learning model that has been trained on prior data and possesses emotion recognition capabilities, enabling it to stably process acoustic features and output emotion state identifiers. In detail, a pre-trained neural network model includes an input layer, hidden layers, and an output layer.
[0043] The input layer filters and extracts target acoustic feature signals from the original acoustic features. It should be noted that the target acoustic feature signals refer to the feature subset that plays a key role in emotion recognition, while the original acoustic features may contain redundant information unrelated to emotion (such as features of environmental noise interference). The input layer can use the judgment logic learned through pre-training to remove invalid information from the original features such as pitch, speech rate, and energy, and retain the feature signals that are most distinguishable for emotion classification (such as calm and anger).
[0044] The hidden layer implicitly extracts emotional information from the target acoustic feature signal. Emotional information is abstract information that reflects the user's emotional attributes. It should be noted that emotional information is an intermediate product between "objective features" and "identified output," and it does not have direct readability; it belongs to the "emotional representation" processed internally by the neural network model. For example, the hidden layer can extract the abstract emotional information that "the user is in an angry state" by analyzing the target feature signal of "high pitch + fast speech rate + high energy," but its form is parameters or vectors that the model can interpret, rather than a direct textual description.
[0045] The output layer is used to generate emotional state labels that clearly represent the user's emotional type.
[0046] The processing flow of a pre-trained neural network model can be as follows: The input layer of the pre-trained neural network model receives the original acoustic features. Through the feature selection logic obtained from the model's pre-training, redundant information (such as features of environmental noise interference) is extracted from the original features, and the "target acoustic feature signals" with key discriminative power for emotion recognition are selected and directed to the hidden layer. Subsequently, the hidden layer performs deep processing on the input target acoustic feature signals. Through the connection weights and activation functions between neurons, it implicitly mines the correlation between feature signals and emotions (such as "high pitch + high energy + fast speech" corresponding to anger), transforming the concrete feature signals into abstract "emotional information" that the model can interpret. Finally, the output layer receives the abstract emotional information transmitted from the hidden layer and, based on the pre-trained classification logic, transforms it into "emotional state labels" with clear semantics.
[0047] Thus, based on the pre-trained neural network model, emotional state identifiers are determined according to acoustic features. The neural network model includes an input layer, a hidden layer, and an output layer. The input layer determines the target acoustic feature signal from the acoustic features and inputs it to the hidden layer. The hidden layer implicitly captures emotional information based on the target acoustic feature signal and inputs it to the output layer. The output layer outputs the emotional state identifier based on the emotional information. In this way, the pre-trained neural network model can accurately identify different emotional state identifiers, providing a basis for subsequently generating appropriate emotionally-adaptive background dialogue.
[0048] Please see Figure 5 In some implementations, step 02 (generating emotion-adaptive background dialogue based on emotion state identifiers and voice requests) includes: 021: Perform natural language processing on voice requests to determine user intent; 022: Generate emotion-adaptive background messages based on emotion state identifiers and user intent.
[0049] In some implementations, the determination module is also used to perform natural language processing on the voice request to determine the user's intent, and to generate emotion-adaptive background dialogue based on the emotion state identifier and the user's intent.
[0050] In some implementations, the processor is also configured to perform natural language processing on the voice request to determine the user's intent, and to generate emotion-adaptive background dialogue based on the emotion state identifier and the user's intent.
[0051] Specifically, user intent refers to the core needs or goals expressed by a user through voice requests. It is the core user demand extracted from the voice content by the system after natural language processing.
[0052] After obtaining the emotion state identifier, the system performs natural language processing on the voice request to extract the core intent. Then, based on the emotion state identifier and the user's intent, it outputs targeted emotion-adaptive background dialogue.
[0053] In this way, natural language processing is performed on the voice request to determine the user's intent. Then, based on the emotion state identifier and the user's intent, emotion-adapted background dialogue is generated. By determining the user's intent and combining it with the emotion state identifier and the user's intent, the generated emotion-adapted background dialogue can be anchored to the user's actual needs, avoiding the generation of irrelevant content based solely on emotions, and ensuring that the emotion-adapted background dialogue is strongly relevant to the user's goals.
[0054] Please see Figure 6 In some implementations, step 021 (performing natural language processing on the voice request to determine the user's intent) includes: 0211: Based on the voice request, sequentially determine the first voice request segment received at the first moment and the second voice request segment received at the second moment that follows the first voice request segment; 0212: Perform speech recognition processing on the first speech request segment to determine the first recognition result; 0213: Perform natural language understanding processing on the first recognition result to determine the first understanding result; 0214: Perform speech recognition processing on the second speech request segment to determine the second recognition result; 0215: Perform natural language understanding processing on the second recognition result to determine the second understanding result; 0216: Based on the temporal relationship between the first understanding result and the second understanding result, the first understanding result and the second understanding result are logically fused to determine the user intent.
[0055] In some embodiments, the determining module is further configured to sequentially determine, based on the voice request, a first voice request segment received at a first time and a second voice request segment received at a second time that follows the first voice request segment. It also performs speech recognition processing on the first voice request segment to determine a first recognition result, and performs natural language understanding processing on the first recognition result to determine a first understanding result. The determining module is further configured to perform speech recognition processing on the second voice request segment to determine a second recognition result, and perform natural language understanding processing on the second recognition result to determine a second understanding result. Finally, based on the temporal relationship between the first understanding result and the second understanding result, it performs logical fusion processing on the first understanding result and the second understanding result to determine the user intent.
[0056] In some embodiments, the processor is further configured to, based on a voice request, sequentially determine a first voice request segment received at a first time and a second voice request segment received at a second time that follows the first voice request segment. It also performs speech recognition processing on the first voice request segment to determine a first recognition result, and performs natural language understanding processing on the first recognition result to determine a first understanding result. The processor is further configured to perform speech recognition processing on the second voice request segment to determine a second recognition result, and perform natural language understanding processing on the second recognition result to determine a second understanding result. Finally, based on the temporal relationship between the first understanding result and the second understanding result, it performs logical fusion processing on the first understanding result and the second understanding result to determine the user intent.
[0057] Specifically, the first and second voice request segments refer to consecutive voice request segments obtained after segmenting the user's voice request. The second voice request segment refers to a series of voice request segments following the first. For example, if the user's voice request is "Help me find the highest-rated seafood restaurant nearby, with parking available," then the first voice request segment might be "Help me find the highest-rated seafood restaurant nearby," and the second voice request segment might be "Parking available." Upon receiving the first voice request segment, the intelligent cockpit system begins natural language processing, performing speech recognition and natural language understanding on the first voice request segment "Help me find the highest-rated seafood restaurant nearby," determining the user's intent as "find a seafood restaurant." Subsequently, natural language processing is performed on the second voice request segment to supplement or correct the user's intent. This eliminates the need to wait for the model to complete speech recognition of the complete voice request before performing natural language understanding.
[0058] The first recognition result refers to the speech-recognized text corresponding to the first speech request segment.
[0059] The first understanding result refers to the natural language processing result corresponding to the first voice request segment.
[0060] The second recognition result refers to the speech-recognized text corresponding to the second speech request segment.
[0061] The second understanding result refers to the natural language processing result corresponding to the second speech request segment.
[0062] The temporal relationship refers to the "sequential reception and parsing order" of different voice request segments and their corresponding understanding results. That is, the second voice request segment is generated later than the first voice request segment, and the second understanding result is a supplement, continuation or correction to the first understanding result.
[0063] Logical fusion processing refers to the system determining the semantic association type of all understanding results (such as supplementary explanations, conditional constraints, and progressive intent) based on "temporal relationships" and integrating the two into a complete semantic expression, rather than viewing the parsing results of individual fragments in isolation.
[0064] The user's voice request is segmented into multiple consecutive segments using a fixed-duration window, such as a first voice request segment and a second voice request segment. After identifying the first voice request segment, speech recognition is performed on it to convert it into text, i.e., the first recognition result. Next, semantic understanding is performed on the text content of the first voice request segment (i.e., the first recognition result) to analyze its intent and meaning. Finally, the recognition and understanding results of multiple voice request segments are logically fused to determine the intent of the entire voice request, i.e., the natural language processing result.
[0065] Thus, based on the voice request, the server sequentially determines the first voice request segment received at a first moment and the second voice request segment received at a second moment, which follows the first voice request segment, with the second moment being later than the first moment. Next, speech recognition processing is performed on the first voice request segment to determine the first recognition result. Then, natural language understanding processing is performed on the first recognition result to determine the first understanding result. Subsequently, speech recognition processing is performed on the second voice request segment to determine the second recognition result. Next, natural language understanding processing is performed on the second recognition result to determine the second understanding result. Finally, based on the temporal relationship between the first and second understanding results, logical fusion processing is performed on the first and second understanding results to determine the user's intent. In this way, by segmenting the voice request into multiple voice request segments and processing them in a streaming manner, the user does not need to wait for the model to complete speech recognition of the complete voice request before natural language understanding, which shortens the model response time and improves the user experience.
[0066] Please see Figure 7 In some implementations, the emotion-adaptive pre-conversation message includes at least one of reassuring pre-conversation messages, empathetic pre-conversation messages, and lead-in pre-conversation messages. Step 022 (generating emotion-adaptive pre-conversation messages based on the emotion state identifier and the user's intent) includes: 0221: When the emotional state is identified as the primary emotional state and the user's intent is a need-based intent, generate reassuring words; or 0222: When the emotional state is identified as the second emotional identifier and the user intent is a solution-oriented user intent, generate empathetic opening remarks; or 0223: When the emotional state identifier is the first emotional identifier and the user intent is a solution-type user intent, generate a lead-in preamble.
[0067] In some embodiments, the voice interaction device further includes a generation module, which generates reassuring opening remarks when the emotional state identifier is a first emotional identifier and the user intent is a need-based user intent; or generates empathetic opening remarks when the emotional state identifier is a second emotional identifier and the user intent is a solution-based user intent; or generates leading opening remarks when the emotional state identifier is a first emotional identifier and the user intent is a solution-based user intent.
[0068] In some embodiments, the voice interaction device further includes a generation module, which generates reassuring opening remarks when the emotional state identifier is a first emotional identifier and the user intent is a need-based user intent; or generates empathetic opening remarks when the emotional state identifier is a second emotional identifier and the user intent is a solution-based user intent; or generates leading opening remarks when the emotional state identifier is a first emotional identifier and the user intent is a solution-based user intent.
[0069] Specifically, reassuring backing messages refer to emotion-adaptive backing messages generated when the emotional state is identified as the primary emotion marker and the user's intent is a need-based intent. These messages convey service progress and soothe waiting emotions. For example, if a user calmly asks, "Help me find the highest-rated seafood restaurant nearby that supports parking" (a need-based intent, primary emotion marker), the system generates the reassuring backing message, "Searching for the highest-rated seafood restaurants nearby, restaurants with parking have been specifically marked, please wait a moment."
[0070] The first emotion marker refers to a calm state where "emotional fluctuations are less than a preset emotion threshold," serving as the emotional basis for triggering soothing and introductory conversations. For example, the voice characteristics (stable tone, moderate speaking speed, and balanced energy) when a user asks questions about restaurants or traditional culture correspond to the first emotion marker.
[0071] Demand-based user intent refers to the core needs expressed by users through voice requests, specifically the need to "obtain services or specific information".
[0072] Empathic background messages refer to emotion-adaptive background messages generated when the emotional state is identified as a second emotion identifier and the user intent is a problem-solving intent. These messages can empathize with negative emotions and synchronize the progress of problem solving. For example, if a user complains angrily, "Why is the navigation stuck?" (a problem-solving intent, second emotion identifier), the system can generate an empathic background message such as, "Navigation signal fluctuations have been detected. We are working hard to restore accurate positioning. Please calm down; it will be fine soon."
[0073] The second emotion marker refers to a negative state (such as anger or anxiety) where the emotional fluctuation is greater than or equal to a preset emotion threshold. It serves as the emotional basis for triggering empathic prelude speech. For example, the voice characteristics of a user complaining about navigation lag (higher pitch, faster speech, and increased energy) correspond to the second emotion marker.
[0074] Solving-related user intents refer to the core needs expressed by users through voice requests, specifically the need to "solve problems, troubleshoot faults, or acquire knowledge." Examples include, "Why is the navigation stuck?" (solving a navigation problem), and "What are the three rings in traditional Chinese culture?"
[0075] Preliminary opening remarks refer to emotion-adaptive opening remarks generated when the emotional state identifier is the first emotion identifier and the user intent is the solution-oriented user intent. These remarks can initiate a topic and lay the groundwork for a subsequent formal response. For example, if a user asks in a calm tone, "What are the three rings in traditional Chinese culture?" (solution-oriented intent, first emotion identifier), the system generates the preliminary opening remarks, "In our traditional culture..."
[0076] Thus, when the emotional state identifier is the first emotional identifier and the user intent is a need-based user intent, a soothing interlude is generated, where the first emotional identifier indicates that the user is in a peaceful state with emotional fluctuations below a preset emotional threshold. Alternatively, when the emotional state identifier is the second emotional identifier and the user intent is a solution-based user intent, an empathetic interlude is generated, where the second emotional identifier indicates a negative state with emotional fluctuations greater than or equal to a preset emotional threshold. Or, when the emotional state identifier is the first emotional identifier and the user intent is a solution-based user intent, a leading interlude is generated. This method of generating different emotion-adaptive interludes for different scenarios makes voice interaction more natural and human-like, improving user satisfaction with the voice interaction experience.
[0077] Please see Figure 8 In some implementations, the method further includes: 03: After generating the emotion-adaptive backing message, the emotion-adaptive backing message is sent to the first target electronic device so that the emotion-adaptive backing message can be broadcast through the first target electronic device.
[0078] In some implementations, the determining module is further configured to send the emotion-adapted backing message to a first target electronic device after generating the emotion-adapted backing message, so as to broadcast the emotion-adapted backing message through the first target electronic device.
[0079] In some implementations, the processor is further configured to send the emotion-adapted backing message to a first target electronic device after generating the emotion-adapted backing message, so as to broadcast the emotion-adapted backing message through the first target electronic device.
[0080] Specifically, the first target electronic device refers to the hardware device pre-installed in the smart cockpit for receiving emotion-adaptive voice prompts and performing voice broadcasts, which is usually a car audio system or other cockpit terminal device with voice output function.
[0081] Upon receiving a user's voice request, the intelligent cockpit system analyzes features such as tone, speed, and energy to identify the user's emotional state, such as anger, frustration, or anxiety. Simultaneously, the system converts the voice request into text and performs natural language processing to analyze the user's intent and content. Next, based on the user's emotional state and the content of the voice request, the system generates appropriate emotional feedback, such as a reassuring or empathetic response. Finally, the system combines the emotional feedback and the natural language processing results to generate the final response, which is then broadcast to the user in voice format.
[0082] Thus, after generating the emotion-adaptive backing message, it is sent to the first target electronic device for broadcast. This rapid broadcast of the generated backing message through the target electronic device quickly delivers feedback to the user during speech recognition, semantic understanding, and command distribution and execution. It fills the gap between the user's voice request and the return of the formal result, preventing user anxiety due to prolonged waiting, improving the user's perception of response efficiency, and ultimately enhancing user satisfaction and trust.
[0083] Please see Figure 9 In some implementations, the method further includes: 04: Generate control commands based on the user's intent; 05: Send the control command to the second target electronic device so that the control command can be executed by the second target electronic device.
[0084] In some implementations, the determining module is further configured to generate control instructions based on the user intent, and to send the control instructions to a second target electronic device for execution of the control instructions via the second target electronic device.
[0085] In some implementations, the processor is further configured to generate control instructions based on the user intent, and to issue the control instructions to a second target electronic device for execution of the control instructions by the second target electronic device.
[0086] Specifically, control commands refer to instruction signals with clear execution logic generated through natural language processing based on the "user intent" determined while generating emotion-adaptive background dialogue. These signals can drive the second target electronic device in the smart cockpit to perform operations that directly match the user's core needs.
[0087] The second target electronic device refers to the pre-set hardware device in the smart cockpit scenario that is used to receive information from the system and perform corresponding operations.
[0088] For example, when a user's intent is "to find the highest-rated seafood restaurant nearby that supports parking" (a demand-based intent), the system might generate a control command that "calls the in-vehicle map service and performs a nearby search operation with the keywords 'seafood restaurant' and the filter conditions 'highest rating' and 'supports parking'." After this command is sent to the in-vehicle map terminal (the target electronic device), the search function is triggered.
[0089] Thus, control commands are generated based on the user's intent. These commands are then sent to a second target electronic device for execution. This method of generating control commands based on the user's intent determined through natural language processing ensures a high degree of alignment between the commands and the user's actual needs.
[0090] Please see Figure 10 In some implementations, the method further includes: 06: Receive the execution result of the control command fed back by the second target electronic device; 07: Based on the execution result, generate the corresponding execution result broadcast information and send the execution result broadcast information to the first target electronic device so that the first target electronic device can broadcast the execution result broadcast information after the emotion-adaptive backing speech has been completed.
[0091] In some embodiments, the voice interaction device further includes a receiving module and a broadcasting module. The receiving module receives the execution result of the control command fed back by the second target electronic device. The broadcasting module generates corresponding execution result broadcasting information based on the execution result and sends the execution result broadcasting information to the first target electronic device, so that the first target electronic device can broadcast the execution result broadcasting information after the emotion-adaptive background speech has been completed.
[0092] In some implementations, the processor is also used to receive the execution result of the control command fed back by the second target electronic device. That is, based on the execution result, a corresponding execution result broadcast information is generated and sent to the first target electronic device, so that the first target electronic device can broadcast the execution result broadcast information after the emotion-adaptive background speech has been broadcast.
[0093] Specifically, the execution result refers to the status information, representing the completion status or generated status, fed back to the system by the second target electronic device (i.e., the device executing the control command) after completing the operation corresponding to the control command. It should be noted that the second target electronic device is used to execute the received control command, while the first target electronic device is used to broadcast the received emotion-adaptive background dialogue and execution result broadcast information. Under the scheduling of the server, the first and second target electronic devices jointly realize a parallel link of functional execution and emotional feedback. First, the first target electronic device broadcasts the received emotion-adaptive background dialogue, filling the waiting period between the user issuing a voice request and receiving the formal result. Simultaneously, the second target electronic device (such as a navigation module and air conditioning control system) executes the received execution command and feeds back the execution result to the server in the intelligent cockpit system. After receiving the execution result from the second target electronic device, the server generates corresponding execution result broadcast information and sends it to the first target electronic device. The first target electronic device will broadcast the execution result broadcast information after the emotion-adaptive background dialogue has finished. In this way, the second target electronic device can achieve precise execution of control commands, while the first target electronic device can eliminate the potential disconnect during voice interaction.
[0094] The execution result broadcast information refers to natural language information that is generated through further processing of the "execution result" and is suitable for being conveyed to the user via voice.
[0095] For example, a user complains angrily, "Why is the navigation stuck?" (a problem-solving intent, a second emotional marker). The system generates an empathetic message: "Navigation signal fluctuations detected. We are working hard to restore accurate positioning. Please calm down, it will be fine soon," and sends it to the in-vehicle audio system (the first target electronic device). Subsequently, the smart cockpit system generates a control command: "Restart the in-vehicle navigation module and reload the route," and sends it to the navigation control module (the second target electronic device). The navigation control module then reports the execution result: "Navigation module has restarted, route data loaded, current positioning is accurate." Next, the smart cockpit system generates an execution result broadcast message: "Navigation has returned to normal. The current route has been replanned. We will continue to navigate you to your destination," and sends it to the in-vehicle audio system (the first target electronic device).
[0096] In this way, the server receives the execution result of the control command from the second target electronic device. Then, based on the execution result, it generates corresponding execution result broadcast information and sends it to the first target electronic device. This allows the first target electronic device to broadcast the execution result information after the emotion-adaptive background message has finished. In this way, receiving the execution result of the control command and generating the corresponding broadcast information accurately feeds back the final processing result of the user's needs to the user, forming a complete chain of "instant response—function execution—result confirmation" with the previous emotion-adaptive background message, thus avoiding gaps or abrupt switches between emotion-adaptive background messages and execution result feedback.
[0097] This application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it implements the steps of the voice interaction method described above.
[0098] It is understood that a computer program includes computer program code. Computer program code can be in the form of source code, object code, executable files, or some intermediate form. Computer-readable storage media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution media, etc.
[0099] In this specification, the terms "specifically," "furthermore," "particularly," "understandably," etc., refer to specific features, structures, materials, or characteristics described in connection with embodiments or examples that are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0100] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of executable request code comprising one or more steps for implementing a particular logical function or process, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order according to the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0101] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A voice interaction method, characterized in that, The method includes: Based on the acoustic features of the received voice request, an emotional state identifier is determined, wherein the acoustic features include at least one of pitch features, speech rate features, and energy features; Based on the emotional state identifier and the voice request, an emotion-adaptive backing message is generated to complete the voice interaction.
2. The method according to claim 1, characterized in that, The step of determining the emotional state identifier based on the acoustic features of the received voice request includes: Based on a pre-trained neural network model, the emotional state identifier is determined according to the acoustic features. The neural network model includes an input layer, a hidden layer, and an output layer. The input layer is used to determine the target acoustic feature signal from the acoustic features and input it to the hidden layer. The hidden layer implicitly captures emotional information based on the target acoustic feature signal and inputs it to the output layer. The output layer outputs the emotional state identifier based on the emotional information.
3. The method according to claim 1, characterized in that, The step of generating emotion-adaptive background dialogue based on the emotion state identifier and the voice request includes: Natural language processing is performed on the voice request to determine the user's intent; Based on the emotional state identifier and the user intent, generate an emotion-adaptive background message.
4. The method according to claim 3, characterized in that, The step of performing natural language processing on the voice request to determine the user's intent includes: Based on the voice request, a first voice request segment received at a first moment and a second voice request segment received at a second moment that follows the first voice request segment are determined sequentially, wherein the second moment is later than the first moment; Perform speech recognition processing on the first voice request segment to determine the first recognition result; Perform natural language understanding processing on the first recognition result to determine the first understanding result; The second voice request segment is subjected to speech recognition processing to determine the second recognition result; The second recognition result is processed by natural language understanding to determine the second understanding result; Based on the temporal relationship between the first understanding result and the second understanding result, the first understanding result and the second understanding result are logically fused to determine the user intent.
5. The method according to claim 3, characterized in that, The emotion-adaptive backing words include at least one of reassuring backing words, empathetic backing words, and lead-in backing words. The step of generating emotion-adaptive backing words based on the emotion state identifier and the user's intent includes: When the emotional state identifier is a first emotional identifier and the user intent is a need-based user intent, the reassuring words are generated, wherein the first emotional identifier is used to indicate that the user is in a peaceful state with emotional fluctuations less than a preset emotional threshold; or When the emotional state identifier is a second emotional identifier and the user intent is a solution-oriented user intent, the empathic preamble is generated, wherein the second emotional identifier is used to indicate a negative state where the emotional fluctuation is greater than or equal to the preset emotional threshold; or When the emotional state identifier is the first emotional identifier and the user intent is the solution-type user intent, the lead preamble is generated.
6. The method according to claim 5, characterized in that, The method further includes: After generating the emotion-adaptive background message, the emotion-adaptive background message is sent to the first target electronic device so that the emotion-adaptive background message can be broadcast through the first target electronic device.
7. The method according to claim 3, characterized in that, The method further includes: Generate control commands based on the user's intent; The control command is sent to the second target electronic device so that the control command is executed by the second target electronic device.
8. The method according to claim 7, characterized in that, The method further includes: Receive the execution result of the control command fed back by the second target electronic device; Based on the execution result, a corresponding execution result broadcast information is generated and sent to the first target electronic device so that the first target electronic device can broadcast the execution result broadcast information after the emotion-adaptive padding broadcast is completed.
9. A server, characterized in that, The server includes a processor and a memory, the memory storing a computer program that, when executed by the processor, implements the method according to any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method as described in any one of claims 1-8.
Citation Information
Cited By
Low-delay human-computer interaction method and system based on intention speculation and prefix streaming splicing, electronic equipment and storage medium
CN121905160A
Low-latency human-computer interaction method and system based on intention speculation and prefix streaming splicing, electronic device and storage medium
CN121905160B