Customizable speech recognition system with user feedback analysis and dynamic reorientation
The system uses facial expressions and gesture recognition with advanced neural networks and sensors to improve voice assistant accuracy and adaptability in automotive infotainment systems, addressing challenges of dialects and environmental noise.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- MERCEDES BENZ GROUP AG
- Filing Date
- 2024-12-18
- Publication Date
- 2026-05-07
AI Technical Summary
Existing voice assistant systems in automotive infotainment face challenges in accurately interpreting voice commands due to dialects, pronunciations, and environmental noise, and are subjective to user context and preferences, leading to inadequate responses.
Incorporating facial expressions, emotional feedback, and gesture recognition using deep neural networks and convolutional neural networks, along with RGB cameras and depth sensors, to detect incorrect responses and adapt user interactions.
Enhances user experience by accurately interpreting voice commands, adapting to user context, and continuously learning from feedback for improved performance over time.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND
[0001] In the field of automotive infotainment systems, voice assistants (VAs), which utilize a blend of natural language processing (NLP), speech recognition, and machine learning, have become increasingly important. These advanced systems are designed to interpret human speech, convert it into text, and process the text to execute specific tasks or commands. The goal is to provide the user with a seamless, hands-free experience, particularly while driving, where manual interaction can pose a safety risk.
[0002] However, the practical implementation of such systems can be challenging due to a variety of factors. For example, the input commands given to the voice assistant may not always elicit the desired response. This can be due to a number of reasons, such as the command not being received in the correct order or problems with dialect and pronunciation. These factors can prevent the speech recognition engine from accurately decoding the input command. This problem becomes even more apparent in a dynamic and often noisy environment like a vehicle.
[0003] Furthermore, the effectiveness of a voice assistant's response can be subjective and depend on the context, the user's expectations, and preferences. Thus, a response that seems satisfactory in one context may be perceived as inadequate or incorrect in another. This subjectivity makes the design and implementation of such systems even more complex.
[0004] On the other hand, alternative methods for improving the accuracy of voice assistants have been explored. For example, incorporating visual cues or gestures as additional input modalities may improve the system's understanding of the user's intentions. However, these methods present a number of challenges, such as the need for additional hardware or the difficulty of correctly interpreting nonverbal cues.
[0005] Therefore, it is necessary to overcome the aforementioned problems. A solution that can accurately interpret and process voice commands, handle different dialects and pronunciations, and adapt to the user's context and preferences would be extremely beneficial. Furthermore, the solution should be robust enough to function effectively in a dynamic environment such as a vehicle. In addition, it would be advantageous if the solution could learn from past interactions and adapt to improve its performance over time. SUMMARY
[0006] The primary objective of the present invention is to enable an improved user experience with a voice assistant (VA) by incorporating facial expressions, emotional feedback, and gesture recognition. This allows the VA to detect incorrect or unexpected responses based on the user's nonverbal feedback and restart the session by prompting the user to repeat their command.
[0007] Another object of the present invention is the use of deep neural network (DNN) and convolutional neural network (CNN) models for the recognition of facial expressions, emotions, and gestures. The output of these models is used to inform the VA's response and to improve the accuracy and relevance of its interactions with the user.
[0008] A further object of the present invention is to utilize hardware such as RGB cameras, depth sensors such as Microsoft Kinect, or LiDAR sensors for the precise tracking of facial features and the analysis of facial expressions. This hardware can also be used for gesture recognition by tracking the movement of hands or body parts.
[0009] A further object of the present invention is the use of machine learning and computer vision techniques for processing the data from these sensors. Libraries and frameworks such as OpenCV, TensorFlow, and PyTorch can be used to develop applications for facial expression and gesture recognition.
[0010] According to one aspect of the present invention, a computer system comprises a speech recognition module 10, a processing unit 12, a feedback detection module 14, and a control unit 18. The speech recognition module 10 is configured to receive an input command in the form of speech and convert it into text. The processing unit 12 transmits the converted text to a backend system 16 for further processing.
[0011] According to another aspect of the present invention, the detection module 14 is designed to analyze the user's feedback associated with the response to the input command. It detects indicators of an incorrect response, including facial expressions, gestures, or emotional states. When an indicator of an incorrect response is detected, the control unit restarts the session by prompting the user to repeat the input command.
[0012] In another embodiment of the present invention, the detection module 14 further comprises a deep neural network (DNN) trained to classify the user's facial expression as positive, neutral, or negative. This module can also employ a convolutional neural network (CNN) to process real-time video data of the user's facial expressions.
[0013] In another embodiment, the control module 18 can initiate a different response based on the detected emotional state, e.g., making suggestions to clarify the input command. Gesture input can be detected via a sensor interface that includes touch-based or non-touch-based input sensors.
[0014] In another embodiment, the detection module 14 is configured to detect changes in user feedback over time and dynamically update the threshold for retriggering. The control module 18 logs and stores data about the repeated commands and user feedback to improve the system in the future.
[0015] The preceding sections serve as a general introduction and are not intended to limit the scope of the following claims. The described embodiments and further advantages are best understood by referring to the following detailed description in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1: Block diagram of a computer system illustrating the workflow according to an embodiment of the invention DETAILED DESCRIPTION OF THE INVENTION
[0016] Aspects of the present invention are best understood by reference to the description contained herein. All aspects described herein will be better appreciated and understood when considered in conjunction with the following descriptions. However, it should be understood that the following descriptions, while indicating preferred aspects and numerous specific details thereof, are given for illustrative purposes only and should not be treated as limitations. Changes and modifications may be made within the scope of the present description without departing from the spirit and scope of the description, and the present invention includes all such modifications.
[0017] The present invention relates to a system as it is described in Fig.Figure 1 illustrates the system. In one embodiment, the system is located within a vehicle and comprises various components. These are explained in detail in the following sections. It includes a speech recognition module 10. The speech recognition module 10 is designed to receive an input command in the form of speech and convert the input command into text. The speech recognition module 10 is an integral part of the system and plays a crucial role as a central hub in processing the user's commands. It uses advanced speech recognition algorithms to accurately convert spoken language into written text. This enables the system to effectively understand and process the user's commands.
[0018] In one embodiment, the speech recognition module 10 can, for example, be part of an automotive / vehicle head unit / infotainment system. In some embodiments, it can include one or more processors configured to perform a transcription, speech recognition, or speech-to-text function. In some implementations, the transcription, speech recognition, or speech-to-text function can be performed on the system for processing user input (e.g., based on audio data captured by the one or more microphones).
[0019] The system also includes a processing unit 12. This unit is configured to transfer the converted text to a backend system 16 for further processing, for example, a cloud-based platform that provides the support system with additional power / capability. The processing unit 12 is a crucial component of the computer system because it is responsible for transmitting the converted text to the backend system. This step is essential to ensure that the user's command is processed correctly and the desired action is performed.
[0020] In one embodiment, the processing unit 12 may comprise one or more processors (e.g., microprocessors), one or more processing cores, a programmable logic circuit (PLC) or a programmable logic / gate array (PLA / PGA), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or any other control unit. In some embodiments, the control circuit or control system may be part of, or form, a vehicle control unit (also referred to as, the vehicle control unit) that is embedded in or otherwise arranged within the vehicle (e.g., a Mercedes-Benz® passenger car or van). The vehicle control unit may, for example, be an infotainment system control unit (e.g.,an infotainment main unit), a telematics control unit (TCU), an electronic control unit (ECU), a central powertrain control unit (CPC), a charger, a central exterior and interior control unit (CEIC), a zone control unit or any other control unit.
[0021] The system also features a feedback detection module (14). This module analyzes user feedback related to the response to an input command. It detects one or more indicators of an incorrect response, which may include a facial expression, a gesture, or an emotional state. The detection module uses advanced machine learning algorithms and data analysis techniques to accurately detect and interpret these indicators. This allows the system to determine whether the user is satisfied with the response.
[0022] A control module 18 is another important component of the system. This module is configured to retrigger the session if the feedback detection module detects an indicator of an incorrect response. It prompts the user to repeat the input command to initiate a new session. The control module plays a crucial role in ensuring that the system can effectively handle incorrect responses and gives the user the opportunity to correct their command.
[0023] Detection Module 14 also includes a deep neural network (DNN) trained to classify the user's facial expression as positive, neutral, or negative. A DNN is a type of artificial neural network with multiple layers between the input and output layers. It uses complex algorithms and large datasets to learn and make accurate predictions. In this case, the DNN is trained to classify the user's facial expressions, which can provide valuable insights into the user's emotional state and their satisfaction with the system's response.
[0024] Detection Module 14 also uses a convolutional neural network (CNN) to process real-time video data of the user's facial expressions. CNNs are a type of deep learning model that is particularly effective at processing visual data. They use a variation of multilayer perceptrons designed to require minimal preprocessing. In this case, the CNN is used to analyze real-time video data of the user's facial expressions, enabling a more accurate and comprehensive understanding of user feedback.
[0025] Based on the detected emotional state, Control Module 18 can also initiate a different response, such as suggesting ways to clarify the input command. This function of the control module allows the system to respond to user commands more individually and effectively. It can make suggestions for clarifying the input command based on the user's emotional state, thus improving the user experience and increasing the system's effectiveness.
[0026] The system can detect gestures via a sensor interface that includes touch-based or non-touch-based input sensors available through the vehicle's infotainment system. These sensors can accurately detect and interpret the user's gestures, providing the system with an additional layer of feedback. This allows the system to more comprehensively understand the user's commands and feedback, resulting in more accurate responses and an improved user experience.
[0027] The detection module is also configured to detect changes in user feedback over time and dynamically update the trigger threshold. This allows the system to adapt to changing user feedback and preferences, ensuring it remains effective and user-friendly over time.
[0028] Control Module 18 also logs and stores data on repeated commands and user feedback to improve the system in the future. This data can provide valuable insights into user preferences and system performance, which can be used for future system improvements. This ensures that the system evolves and improves over time, providing the user with a consistently high-quality experience.
[0029] In summary, the present invention represents a comprehensive and effective solution for processing user input and feedback in a speech recognition system. It incorporates advanced technologies and algorithms to accurately process and respond to user commands, taking into account the user's feedback and emotional state. The result is a system that is not only highly effective but also user-friendly and adaptable.
[0030] The block diagram illustrates a workflow for processing user input and feedback in a speech recognition system. The system consists of several interconnected components that work together to enhance the user experience with voice assistant (VA) technology.
[0031] The system begins with a speech recognition module that receives an input command from the user in the form of speech. This input command is then converted into text by the speech recognition module. Advanced speech-to-text algorithms are used in this conversion process to ensure an accurate transcription of the user's spoken command.
[0032] The text version of the input command is then transmitted to a backend processing unit (PPU). This PPU is responsible for analyzing the text command and triggering the corresponding response. The response generated by the backend processing unit is then sent back to the user as an audio response to provide the requested information or action. For example, this could be a cloud server connected to the vehicle's infotainment system.
[0033] A crucial component of this system is the feedback detection module 14. This module is designed to analyze user feedback related to the response to the input command. The detection module is equipped to identify one or more indicators of an incorrect response. These indicators can include various user reactions, such as facial expressions, gestures, or emotional states.
[0034] If detection module 14 detects an indicator of an incorrect response, a control module is triggered. This control module is configured to retrigger the session by prompting the user to repeat the input command, thus initiating a new session.
[0035] Detection Module 14 utilizes advanced machine learning models, such as Deep Neural Networks (DNNs) and Convolutional Neural Networks (CNNs), for the accurate recognition of facial expressions and gestures. These models are trained on large datasets to ensure their accuracy and reliability.
[0036] In addition to the components mentioned above, the system also includes various hardware elements for recognizing the user's facial expressions, emotions, and gestures. For example, an RGB camera is used to recognize facial expressions, while depth sensors such as Microsoft Kinect or LiDAR sensors provide depth information for more accurate tracking of facial features and expression analysis.
[0037] The system is designed to continuously learn and adapt based on user feedback. For example, if the user expresses dissatisfaction with the answer provided by the VA, the system restarts the session and requests further input. This iterative process allows the system to continuously improve and provide more accurate and personalized answers over time.
[0038] In an alternative embodiment, the system could also use other types of sensors to detect gestures, such as infrared sensors or cameras for motion tracking. Likewise, other types of machine learning models, such as recurrent neural networks (RNNs) or support vector machines (SVMs), could be used for facial expression and gesture recognition.
[0039] Furthermore, the system could also incorporate other types of feedback mechanisms, such as speech analysis or sentiment analysis, to further improve its ability to detect incorrect answers and enhance the overall user experience.
[0040] Overall, the system offers a robust and intelligent solution for improving the accuracy and user experience of voice assistant technology. By incorporating advanced machine learning models, sophisticated sensor technology, and an iterative feedback mechanism, the system is able to detect incorrect answers, retrigger sessions, and continuously learn and adapt based on user feedback.
[0041] In one embodiment of the present invention, a computer system includes a series of modules designed to enhance the functionality of a speech recognition system. The speech recognition module is designed to receive an input command in the form of speech and convert it into text. During this process, the spoken command is accurately converted into written text, enabling the system to understand and effectively process the user's commands.
[0042] From here, the converted text is forwarded to a backend system for further processing by a specially designed processing unit. This crucial component of the system ensures that the command is analyzed and an appropriate response is generated. The response, based on the entered command, is sent back to the user, often in the form of audio output, providing the requested information or action.
[0043] Given the complexity of human language and speech, it can happen that the generated response does not meet the user's expectations or needs. In such cases, the user can express their dissatisfaction through a variety of nonverbal cues, such as facial expressions, gestures, or changes in emotional state. Recognizing and responding to these nonverbal cues is an integral part of this embodiment, enhancing the user experience and minimizing potential misunderstandings.
[0044] To capture and analyze these nonverbal cues, the system features a feedback detection module. This module is designed to recognize one or more indicators of an incorrect or unexpected response. It uses advanced machine learning algorithms and data analysis techniques to accurately detect and interpret these indicators. The detection module identifies a variety of reactions, such as facial expressions, gestures, and emotional states, thus providing a comprehensive picture of the user's satisfaction with the system's response.
[0045] If an indicator of an incorrect response is detected, the system triggers a control module. This module is designed to reinitialize the session by prompting the user to repeat the input command, thus starting a new session. This iterative process ensures that the system can effectively handle erroneous or unexpected responses and gives the user the opportunity to correct or modify their command.
[0046] In further embodiments, the feedback detection module comprises a deep neural network (DNN) and a convolutional neural network (CNN). The DNN, a type of multi-layered artificial neural network, is trained to classify the user's facial expression as positive, neutral, or negative. The CNN, which is particularly effective at processing visual data, is used to analyze real-time video data of the user's facial expressions. By combining these two types of neural networks, the system is able to provide a more comprehensive and accurate understanding of user feedback.
[0047] In addition to advanced machine learning models, a range of hardware is used to recognize facial expressions, emotions, and gestures. For example, an RGB camera can be used to detect facial expressions, while depth sensors, such as a Microsoft Kinect or LiDAR sensor, provide depth information for more accurate tracking of facial features and expression analysis. Furthermore, these sensors can be used for gesture recognition, adding another layer of feedback to the system.
[0048] Another embodiment of the present invention enables the system to provide a different response based on the detected emotional state of the user. For example, the system could suggest ways to clarify the input command. This responsive approach can help improve the user experience and increase the system's effectiveness. Gesture input can be detected via a sensor interface with touch-based or non-touch-based input sensors, adding another layer of interaction to the system.
[0049] Another embodiment of the invention includes a dynamic threshold setting. The detection module is configured to detect changes in user feedback over time and dynamically update the threshold for retriggering. This adaptable function allows the system to continuously update and refine its response mechanism to accommodate changing user feedback and preferences.
[0050] In another embodiment, the control module logs and stores data on repeated commands and user feedback to improve the system in the future. Collecting this data provides valuable insights into user preferences and system performance, which can be used for future improvements to ensure a continuously evolving and improving user experience.
[0051] The present invention offers an effective solution to the existing challenges of speech recognition systems. By incorporating user feedback in the form of facial expressions, gestures, and emotional states, the system can deliver more precise and personalized responses, thereby improving the user experience. Through the use of advanced DNN and CNN models coupled with effective sensor technology, the system is able to learn from and adapt to user commands, ensuring continuous improvement over time. The dynamic retrigger threshold further enhances the system's usability and effectiveness.
[0052] In other embodiments, other types of sensors could be used for gesture recognition, such as infrared sensors or motion-tracking cameras. Similarly, other types of machine learning models, such as recurrent neural networks (RNNs) or support vector machines (SVMs), could be used for facial expression and gesture recognition. These alternatives offer additional possibilities and expand the potential application and accessibility of the invention.
[0053] This invention significantly improves the user experience compared to existing solutions, as it offers the potential for more accurate responses and an iterative feedback mechanism. By incorporating advanced machine learning models, sophisticated sensor technology, and a user-centric feedback mechanism, the system ensures improved accuracy and a highly personalized user experience. The combination of these features results in a speech recognition system that not only enhances the user experience but also offers the ability to learn and adapt from the user's commands—a significant advantage over existing solutions.
[0054] The described embodiments of the invention are illustrative and not limiting. Further embodiments and modifications that do not deviate from the spirit and scope of the invention are apparent to a person skilled in the art from the detailed description and the drawings. Therefore, the scope of the invention is not limited by the embodiments described and illustrated here.
Claims
[1] Computer system comprising the following: a speech recognition module configured to receive an input command in the form of speech and convert the input command into text; a processing unit configured to send the converted text to a backend system for further processing; a feedback detection module configured to analyze user feedback associated with the response to the input command, wherein the feedback detection module detects one or more indicators of an incorrect response, the indicators including one or more of the following: a facial expression, a gesture, or an emotional state; A control module configured to retrigger the session if the detection module detects an indicator of an incorrect response, with the control module prompting the user to repeat the input command to initiate a new session. [2] Computer system according to claim 1, wherein the feedback detection module further comprises a deep neural network (DNN) trained to classify users' facial expressions as positive, neutral or negative. [3] Computer system according to claim 1, wherein the feedback detection module uses a convolutional neural network (CNN) to process real-time video data of the user's facial expressions. [4] Computer system according to claim 1, wherein the control module also initiates another response based on the detected emotional state, such as providing suggestions to clarify the input command. [5] Computer system according to claim 1, wherein the gesture input is detected using a sensor interface comprising touch-based or non-touch-based input sensors. [6] Computer system according to claim 1, wherein the detection module is configured to detect changes in user feedback over time and dynamically update the threshold for retriggering. [7] Computer system according to claim 1, wherein the control module logs and stores data about the repeated commands and user feedback for future system improvements. [8] Computer-implemented method for processing user input and feedback in a speech recognition system, comprising the following: Receipt of an input command in the form of speech by a speech recognition module; Conversion of the input command into text by the speech recognition module and transmission of the text to a backend system; Providing a response based on the processed input command; Detecting one or more indicators of an incorrect answer using a detection module, wherein the indicators include one or more of the following: a facial expression, a gesture, or an emotional state of the user; Retriggering the session if an incorrect response is detected, including prompting the user to repeat the input command. [9] Method according to claim 8, further comprising the step of training a neural network model to classify facial expressions that are associated with the user’s satisfaction or dissatisfaction. [10] Method according to claim 8, wherein the detection of an emotional state comprises analyzing a real-time video of the user by a convolutional neural network (CNN). [11] Method according to claim 8, wherein retriggering the session includes offering the user options to clarify or modify the input command based on the detected feedback. [12] Method according to claim 8, wherein the detection of gesture-based feedback comprises the processing of inputs from a nonverbal multi-input interface. [13] Method according to claim 8, further comprising dynamically adjusting the threshold for detecting a false response based on the user's interaction history. [14] The method according to claim 8 further comprises storing data from repeated sessions in order to optimize the future processing of similar input commands.