VEHICLE USER INTERFACE SYSTEM AND METHOD FOR OPERATING SUCH A SYSTEM
Patent Information
- Application Number
- DE102024125351
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2025-09-11
- Estimated Expiration
- 2044-09-04
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
INTRODUCTION
[0001] The present invention relates generally to a vehicle user interface system and a method for operating such a system.
[0002] Some vehicles have voice command capabilities, where a driver or passenger can issue voice requests or commands to retrieve information through a vehicle user interface or control one or more vehicle functions. Separately, large language models (LLMs) are used to generate responses to user requests.
[0003] From DE 10 2019 217 751 A1, a vehicle-side user interface is known in which voice inputs are evaluated by both a target-guided dialogue analysis and a non-target-guided dialogue analysis.
[0004] From the publication CN 1 18 378 614 A a user interface is known in which speech inputs are evaluated by a language model, whereby the language model can be, for example, a statistical language model or a large language model.
[0005] Further prior art is also evident from DE 10 2023 115 462 A1, which, however, was not yet known on the priority date relevant here. SUMMARY
[0006] According to the invention, a vehicle user interface system is presented which is characterized by the features of claim 1.
[0007] The vehicle user interface system comprises: at least one vehicle speaker configured to generate audio signals within a vehicle, a vehicle user interface having a screen configured to display text, at least one vehicle microphone configured to capture the speech of a vehicle occupant, and a vehicle control module configured to: receive speech input from the vehicle occupant via the at least one vehicle microphone, classify the speech input as a deterministic speech request or a probabilistic speech request, in response to the speech input being classified as a deterministic speech request: process the speech input using a statistical language model (SLM) to generate an SLM output, in response tothat the speech input has been classified as a probabilistic speech request: processing the speech input using a large language model (LLM) to generate an LLM output, and generating an audio response output using the at least one vehicle loudspeaker and / or a text response output using the vehicle user interface screen, wherein the audio response output or the text response output is based on the SLM output generated by the statistical language model or the LLM output generated by the large language model.
[0008] In some examples, the vehicle control module is configured to automatically modify operation of at least one vehicle component in response to the voice input including a request from the occupant to operate the at least one vehicle component.
[0009] In some examples, automatically modifying the operation of the at least one vehicle component includes at least one of the following: initiating a phone call via the vehicle user interface, sending a message via the vehicle user interface, activating an entertainment feature of the vehicle user interface, or changing at least one driving setting of the vehicle.
[0010] In some examples, the vehicle control module is configured to: calculate a confidence score for the LLM output of the large language model, compare the confidence score to a specified confidence score threshold that indicates an accurate LLM output probability, and generate the audio response output or the text response output based on the LLM output in response to the confidence score being below the specified confidence score threshold.
[0011] In some examples, calculating the confidence score for the LLM output involves comparing embeddings of tokens of the LLM output with embeddings of the SLM output of the statistical language model.
[0012] In some examples, the specified confidence score threshold is a first confidence score threshold, and the vehicle control module is configured to: compare the confidence score to a second confidence score threshold, wherein the second confidence score threshold is greater than the first confidence score threshold, and generate the audio response output or the text response output based on a combination of the LLM output and the SLM output in response to the confidence score being greater than the first confidence score threshold and less than the second confidence score threshold.
[0013] In some examples, the vehicle control module is configured to: update a database of corrected output labels based on the SLM output of the statistical language model in response to the confidence score being below the specified confidence score threshold, and retrain the large language model using the database of corrected output labels.
[0014] In some examples, the vehicle control module is configured to: obtain model output guardrail data from a database having stored output data about sensitive topics, process the speech input using the large language model (LLM) to generate an intermediate output response, compare the intermediate output response to the model output guardrail data, and prevent the intermediate output response from being output if the intermediate output response includes a disallowed topic of the model output guardrail data.
[0015] In some examples, the vehicle control module is configured to determine a current geographic location of the vehicle, the sensitive topic output data stored in the database varies depending on the geographic location, and comparing the intermediate output response to the model output guardrail data includes comparing the intermediate output response only to sensitive topic output data corresponding to the current geographic location of the vehicle.
[0016] In some examples, the vehicle control module is configured to: determine an emotion score of the vehicle occupant based on the speech input received from the vehicle occupant, compare the emotion score of the vehicle occupant to a specified emotion score threshold indicative of the vehicle occupant's frustration when interacting with the output of the large language model, and generate the audio response output or the text response output based on the SLM output instead of the LLM output if the emotion score of the vehicle occupant exceeds the specified emotion score threshold.
[0017] In some examples, determining the vehicle occupant emotion score comprises: generating a first vehicle occupant emotion score based on the textual processing of the speech input, generating a second vehicle occupant emotion score based on the acoustic processing of the speech input, and combining the first vehicle occupant emotion score and the second vehicle occupant emotion score to generate an overall vehicle occupant emotion score.
[0018] In some examples, the vehicle control module is configured to convert voice input audio signals into text using automatic speech recognition (ASR).
[0019] Furthermore, according to the invention, a method for operating a vehicle user interface system is presented, which is characterized by the features of claim 6.
[0020] The method comprises: capturing speech input from a vehicle occupant using at least one vehicle microphone, classifying the speech input as a deterministic speech request or a probabilistic speech request, in response to the speech input being classified as a deterministic speech request: processing the speech input using a statistical language model (SLM) to generate an SLM output, in response to the speech input being classified as a probabilistic speech request: processing the speech input using a large language model (LLM) to generate an LLM output, and generating an audio response output using the at least one vehicle speaker and / or a text response output using the vehicle user interface screen,where the audio response output or the text response output is based on the SLM output generated by the statistical language model or the LLM output generated by the large language model.,
[0021] In some examples, the method includes automatically modifying the operation of at least one vehicle component in response to the voice input including a request from the occupant to operate the at least one vehicle component.
[0022] In some examples, automatically modifying the operation of the at least one vehicle component includes at least one of the following: initiating a phone call via the vehicle user interface, sending a message via the vehicle user interface, activating an entertainment feature of the vehicle user interface, or changing at least one driving setting of the vehicle.
[0023] In some examples, the method comprises: calculating a confidence score for the LLM output of the large language model, comparing the confidence score to a specified confidence score threshold that indicates an accurate LLM output probability, and generating the audio response output or the text response output based on the LLM output in response to the confidence score being below the specified confidence score threshold.
[0024] In some examples, calculating the confidence score for the LLM output involves comparing the embeddings of tokens from the LLM output with embeddings from the SLM output of the statistical language model.
[0025] In some examples, the specified confidence score threshold is a first confidence score threshold, and the method further comprises: comparing the confidence score to a second confidence score threshold, wherein the second confidence score threshold is greater than the first confidence score threshold, and generating the audio response output or the text response output based on a combination of the LLM output and the SLM output in response to the confidence score being greater than the first confidence score threshold and less than the second confidence score threshold.
[0026] In some examples, the method comprises: updating a database of corrected output labels based on the SLM output of the statistical language model in response to the confidence score being below the specified confidence score threshold, and retraining the large language model using the database of corrected output labels.
[0027] In some examples, the method comprises: obtaining model output guardrail data from a database having stored output data about sensitive topics, processing the speech input using the large language model (LLM) to generate an intermediate output response, comparing the intermediate output response to the model output guardrail data, and preventing the intermediate output response from being output in response to the intermediate output response containing a disallowed topic of the model output guardrail data.
[0028] Further areas of applicability of the present invention will become apparent from the detailed description, claims, and drawings. The detailed description and specific examples are provided for illustrative purposes only. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The present invention will become more fully apparent from the detailed description and the accompanying drawings, in which: Fig. 1 is a functional block diagram of an exemplary embodiment of a vehicle with a vehicle user interface and a large language model; Fig. 2 is a functional block diagram of several language models for use with the vehicle user interface of Fig. 1; Fig. 3 is a flowchart illustrating an exemplary process for automatically controlling vehicle components based on user requirements; Fig. 4 is a flowchart showing an exemplary process for processing user speech using a large language model or a statistical language model; Fig. 5 is a flowchart illustrating an example process for comparing the confidence scores of a large language model and a statistical language model; Fig. 6 is a flowchart showing an example process for using guardrail data to constrain the output of a large language model; Fig. 7 is a flowchart illustrating an exemplary process for selecting between a large language model and a statistical language model based on confidence scores; Fig. 8 is a flowchart illustrating an exemplary process for speech processing including performing acoustic emotion classification; Fig. 9A and Fig. 9B are graphical representations of examples of neural networks for generating responses to the user's speech; and Fig. Figure 10 is a flowchart showing an example process for training a machine learning model.
[0030] Reference numbers may be reused in the drawings to identify similar and / or identical elements. DETAILED DESCRIPTION
[0031] Some embodiments described herein provide vehicle user interfaces configured to prevent or inhibit generative artificial intelligence (AI) models from outputting inappropriate content, which may include hallucinations. Example systems may employ large language models (LLMs) in combination with a natural language understanding model, such as a statistical language model (SLM).
[0032] For example, a statistical language model and a semantic classifier can be used to classify context, topics, and entities from the speech of a user, e.g., a driver or passenger of a vehicle. In some examples, confidence-based heuristics are used to ensure that a large language model is reliable in its predictions or outputs, and to arbitrate the selection of outputs from a natural language processing (NLP) model. The NLP model can be used as a robust baseline to verify the results output by the LLM, and the overall confidence can be estimated for an output message (e.g., a text or audio response from the model to the user).
[0033] A user request (e.g., via voice or text input to a vehicle user interface or mobile device) can be classified as either a deterministic request or a probabilistic request. This designation allows the system to forward the user's speech or input to a large-scale language model (LLM) and / or a statistical language model for further processing. For example, deterministic requests such as "end call" or "turn on the radio" can be more easily handled by a deterministic SLM, while more complex requests such as "find fast-food restaurants near the 600XX zip code" can be more easily handled by a probabilistic LLM.
[0034] A vehicle control module can be configured to decide which type of model to use based on a semantic classification, a keyword- or topic-based rule set, etc. Requirements that go beyond a predefined set of deterministic categories (e.g., ecosystem) can be placed in a fallback context and forwarded to the large language model.
[0035] Large language models do not output confidence scores like other automatic speech recognition results. In some examples, likelihood scores for an output message generated by a large language model may be calculated by summing individual tokens and normalizing them. For example, calculating the likelihood (e.g., confidence score) for LLM output may involve summing the conditional probabilities of individual tokens in the sequence and then normalizing the conditional probabilities by the number of tokens. This may provide a new quantification method for LLMs to obtain corresponding confidence scores, e.g., using the equation ΣP(utterance / context) i = 1:N / N.
[0036] For example, a confidence threshold can be set based on a task-oriented training set that is first thoroughly validated by a statistical language model. The vehicle control module can be configured to compare the LLM output confidence scores with the specified confidence threshold to decide whether the large language model is trustworthy enough to pass the result to the user.
[0037] The utterances verified by the SLM can be used to condition the task performance of the large language model. For example, a training set can be used to optimize the large model by optimizing criteria for minimizing the training loss function (e.g., using information-theoretic metrics such as cross-entropy).
[0038] The LLM outcome can be rewarded or penalized by comparing intents and entity sets with the corresponding Natural Language Understanding (NLA) outcome, e.g., the statistical language model. Based on resonances and collisions between the LLM context and the SLM context, guardrails for system output (e.g., automatically generated audio or text responses to user requests) can be permanently adjusted, and LLM performance can improve significantly over time.
[0039] This leads to refined versions of LLMs, which can be referred to as system-memory-related LLMs (e.g., system-generated rules and refined guardrails) or user-memory-related LLMs (e.g., optimal guardrails for a logged-in user). Erroneous results correctly classified by the SLM are intended to be used to further improve self-assessment learning for the LLM.
[0040] In some examples, a vehicle control module is configured to refine the performance of a generative AI / LLM, e.g., a model for providing responses to user voice requests in a vehicle, by working with statistical language models and a semantic classifier. The performance of a generative AI model can be verified using a base of natural language understanding techniques, e.g., a statistical language model and a semantic classifier. The semantic classifier, for example, can provide robust validation of the classified intents and the corresponding / relevant entities.
[0041] Hallucinations from generative AI can be reduced or minimized, while the execution of identified user requirements is routed through an NLP model. LLMs can be corrected and improved using a starting point from natural language models, such as a statistical language.
[0042] Example vehicle control modules can be configured to use an LLM output confidence score and decision support to forward a user request to a generative AI model or a large language model, and the confidence score can suggest using both models together. A history of interaction with generative AI models and LLMs can be used to develop vendor-specific LLMs (e.g., with generated outputs corresponding to vendor-specific commands) or user-specific LLMs (e.g., with generated outputs corresponding to frequently requested requests by a specific user). A feedback mechanism can be used for LLMs with correctly labeled and verified user language (e.g., training data) to improve the training and performance of LLMs for future user requests related to the same context.
[0043] Some embodiments may provide one or more advantages, such as reduced or minimized instances of hallucinations of the generative AI model, robust performance of the generative AI model using more task-oriented training with feedback and refinement of the LLM for a better user experience, a more effective method for adapting a rule set to limit and minimize hallucinations, highly effective and efficient natural language processing (e.g., by using a methodology that includes generative AI in tandem with natural language understanding models, such as a statistical language model), an adaptive implementation for refining the LLMs for generative AI as well as NLP models using crowdsourced data, etc.
[0044] In Fig. 1 shows a vehicle 10 with front wheels 12 and rear wheels 13. A drive unit 14 delivers Fig. 1 selectively delivers torque to the front wheels 12 and / or the rear wheels 13 via drive lines 16 and 18, respectively. The vehicle 10 may include various types of drive units. The vehicle may, for example, be an electric vehicle such as a battery electric vehicle (BEV), a hybrid vehicle, a fuel cell vehicle, an internal combustion engine (ICE), or another vehicle type.
[0045] Some examples of the drive unit 14 may include any suitable electric motor, an inverter, and a motor controller configured to control power switches within the inverter to adjust motor speed and torque during propulsion and / or regeneration. The battery system supplies or receives power from the electric motor of the drive unit 14 via the inverter during propulsion or regeneration.
[0046] While the vehicle 10 in Fig. 1 includes a drive unit 14, the vehicle 10 may also have other configurations. For example, two separate drive units may drive the front wheels 12 and the rear wheels 13, one or more separate drive units may drive individual wheels, etc. It is understood that other vehicle configurations and / or drive units may also be used.
[0047] The vehicle control module 20 may be configured to control the operation of one or more vehicle components, such as the drive unit 14 (e.g., by controlling the torque settings of an electric motor of the drive unit 14). The vehicle control module 20 may receive inputs to control components of the vehicle, e.g., signals from a steering wheel, an acceleration paddle, etc. The vehicle control module 20 may monitor the vehicle's telematics data for safety purposes, e.g., vehicle speed, vehicle location, vehicle braking and acceleration, etc.
[0048] The vehicle control module 20 may receive signals from any suitable components for monitoring one or more aspects of the vehicle, including one or more vehicle sensors (e.g., cameras, microphones, pressure sensors, wheel position sensors, position sensors such as GPS antennas, etc.). Some sensors may be configured to monitor the vehicle's current motion, vehicle acceleration, steering torque, etc.
[0049] As in Fig. 1, the vehicle 10 includes a user interface 22, a vehicle microphone 24, and a vehicle speaker 26. The user interface 22 may include any suitable buttons, knobs, touchscreens, etc., for accepting input from a driver or passenger of a vehicle. The user interface 22 may include a display for displaying text or images to the driver or passenger.
[0050] One or more vehicle microphones 24 may be mounted at any suitable location within the vehicle 10 and configured to capture the speech of a driver or passenger of the vehicle 10. One or more vehicle speakers 26 may be mounted at any suitable location within the vehicle 10 to provide audio signals to the driver or passenger.
[0051] The user interface 22, the vehicle microphone 24, and the vehicle speaker 26 can be used, for example, for a voice command system, where a driver or passenger can retrieve information and control various aspects of the vehicle via voice commands. Various language models, such as a generative AI model, a large-scale language model, a statistical language model, etc., can be used to process the user's voice requests and then generate responses in the form of text and / or audio signals.
[0052] The vehicle control module 20 may communicate with another device via a wireless communication interface, which may include one or more wireless antennas for transmitting and / or receiving wireless communication signals. For example, the wireless communication interface may communicate via any suitable wireless communication protocol, including, but not limited to, vehicle-to-everything (V2V) communication, Wi-Fi communication, wireless area network (WAN) communication, cellular communication, personal area network (PAN) communication, short-range wireless communication (e.g., Bluetooth), etc. The wireless communication interface may communicate with a remote computing device via one or more wireless and / or wired networks.With regard to vehicle-to-vehicle (V2X) communication, the vehicle 10 may include one or more V2X transceivers (e.g., V2X signal transmitting and / or receiving antennas).
[0053] Fig. 2 is a functional block diagram of a system 200 having multiple speech processing models 201 for use with the vehicle control module 20 of Fig. 1. For example, the raw user speech 202 can be captured via the vehicle microphone 24 and fed to one or more speech processing models 201.
[0054] As in Fig. 2, the various language processing models 201 may include an automatic speech recognition model 204, a large language model 206, and a natural language processing (NLP) model 208, such as a statistical language model (SLM).
[0055] The automatic speech recognition model 204 may be configured to translate the driver's or passenger's acoustic speech into text for further speech processing. The large language model 206 may be a computational model configured to enable general speech generation and other natural language processing tasks, such as classification, by learning statistical relationships from large amounts of text during a computationally intensive, self-supervised or semi-supervised training process. The large language model 206 may be used for text generation, a form of generative AI, by taking an input text and repeatedly predicting the next token or word.
[0056] In some examples, the large language model 206 may be an artificial neural network using a transformer architecture, which may include a transformer-only decoder-only architecture that enables efficient processing and generation of large text data. The large language model 206 may achieve results through prompt engineering, which creates specific input prompts to guide the model's responses. The large language model 206 may acquire knowledge of syntax, semantics, and ontologies inherent in human language.
[0057] As in Fig. 2, the large language model 206 may be configured to generate an intermediate response 210, and the natural language processing model 208 may be configured to generate or use topic, entity, and rule set data 212 to generate an output response. As explained further below, an audio output selector 214 may be configured to provide a user with an audio or text response based on one or more (or a combination) of outputs from the automatic speech recognition model 204, the large language model 206, and the natural language processing model 208.
[0058] Fig. Figure 3 is a flowchart illustrating an exemplary process for automatically controlling vehicle components based on user requests. The process may, for example, be performed by the vehicle control module 20 of Fig. 1, a user's mobile device, another processing device connected to the vehicle 10 or the user, etc. At 304, the method begins by capturing the user's speech, e.g., by receiving speech from a driver or passenger of the vehicle 10 via the vehicle microphone 24 or the user interface 22.
[0059] The vehicle control module is configured to process speech at 308 using an automatic speech recognition (ASR) model. For example, one or more trained models can be configured to convert audio speech signals into text values representing words spoken by a user.
[0060] The vehicle control module is configured to create a dictation text at 312. For example, based on the processing by the automatic speech recognition model, the words spoken by the user can be recorded in a text format suitable for processing by other speech models. Any suitable ASR models or algorithms can be used to process and convert user speech to text.
[0061] The vehicle control module is configured to process the text with a statistical language model (SLM) at 316. For example, the converted or dictated text can be used as input to a suitable statistical language model to determine the intent, context, etc., of the user request.
[0062] The vehicle control module is configured to process the text with a large language model (LLM) at 320. The converted or dictated text can, for example, be used as input to a suitable large language model to determine the intent, context, etc. of the user request. In some examples, both the SLM and the LLM can be used to process the same converted text of the user speech to generate outputs that indicate the user request from each respective model (e.g., for comparison, to check accuracy, to determine which model provided a more useful output, etc.). Although Fig. 3 refers to an SLM and an LLM, in other embodiments, other suitable generative models of artificial intelligence (AI), other suitable models of natural language processing (NLP), etc. may be used.
[0063] The vehicle control module is configured to generate a confidence score based on the evaluation of the LLM output at 324. For example, embeddings of the LLM output may be compared to embeddings of the SLM output to predict how accurately or confidently the LLM output will provide a correct result or response to the user request.
[0064] The vehicle control module is configured to compare the LLM confidence score to a certain threshold (e.g., a threshold indicating that the LLM output is predicted to be correct or accurate) at 328. The vehicle control module is configured to provide a feedback prompt based on the LLM output at 340 if the LLM confidence score is above the threshold at 328. For example, the vehicle speaker 26 or user interface 22 may provide the driver or passenger with an audible or textual response based on the LLM output.
[0065] If at 328, the LLM confidence score is not greater than the specified threshold, control proceeds to 332 to update a database with corrected labels. For example, if a low confidence score indicates that the LLM model likely produced an inaccurate output, the system may fall back to the output of the SLM model (where the SLM model has a higher confidence score, indicating that its output is more likely to be correct), while storing the SLM model output for use in training the LLM model.
[0066] The vehicle control module is configured to periodically train the LLM using the database of corrected labels at 336. For example, after a specified period of time (e.g., hourly, daily, weekly, monthly, etc.) or after a specified number of user requests (e.g., 10 normal or low LLM confidence scores, 100 normal or low LLM confidence scores, etc.), the LLM model can be trained using the corrected labels from the database to make the LLM model more accurate. In this way, outputs from the SLM (which can be considered more likely to be correct if the SLM confidence score is higher than the LLM confidence score) can continue to update the LLM model to refine and make the LLM output more accurate over time.
[0067] The vehicle control module is optionally configured to automatically control one or more vehicle components according to a user request at 344. For example, if the user request is to operate a vehicle navigation system, operate a vehicle communication or entertainment interface, change a driving setting or vehicle operating setting, etc., the vehicle control module may be configured to automatically change, adjust, or modify the operation of one or more vehicle components according to the processed user request.
[0068] Fig. Figure 4 is a flowchart showing an exemplary process for processing user speech using a large language model or a statistical language model. The process may be performed, for example, by the vehicle control module 20 of Fig. 1, a user's mobile device, another processing device connected to the vehicle 10 or the user, etc.
[0069] At 404, the method begins by capturing user speech (e.g., an utterance from a driver or passenger of the vehicle), e.g., via the vehicle microphone 24 of the vehicle user interface 22. The user speech is then processed at 408 using an automatic speech recognition (ASR) model.
[0070] The vehicle control module is configured to classify the voice query as a deterministic voice query or a probabilistic voice query at 412. For example, a deterministic voice query may be a simple request or command to use a vehicle component in a specific way (e.g., "turn on the radio" or "call my spouse"). A probabilistic voice request may require more complicated processing and prediction for the user request, e.g., querying a weather forecast for a future time at a different location, querying a specific type of restaurant near a different location, etc. Any suitable classifier may be used to determine whether the user's voice query is deterministic or probabilistic.
[0071] If the user's voice query is classified as probabilistic in 416, the controller proceeds to 428 to forward the voice query to a large language model. For example, large language models such as generative AI may be better suited to handling more complex probabilistic voice queries from users.
[0072] If the user's voice query is determined to be deterministic at 416, control proceeds to 420 to pass the voice query to a statistical language model and a semantic classifier. The statistical language model may be better suited to handle more straightforward or simpler voice requests. The vehicle control module may determine a confidence score for the SLM output at 424, using any suitable confidence score calculation methods for SLM models.
[0073] The vehicle control module is configured to generate a likelihood accuracy score for the LLM output at 432. For example, embeddings of the LLM output may be compared with embeddings of the SLM output and the SLM confidence score to determine whether the LLM output should have a similar or different likelihood accuracy score than the SLM confidence score (e.g., based on whether the embeddings in the outputs of each model match, etc.). Although Fig. 4 refers to an SLM and an LLM, in other embodiments, other suitable generative models of artificial intelligence (AI), other suitable models of natural language processing (NLP), etc. may be used.
[0074] Fig. Figure 5 is a flowchart illustrating an exemplary process for comparing the confidence scores of a large language model and a statistical language model. The process may be performed, for example, by the vehicle control module 20 of Fig. 1, a user's mobile device, another processing device connected to the vehicle 10 or the user, etc.
[0075] At 504, the method begins by obtaining a sequence of tokens from a large language model. For example, all suitable tokens from the output of the large language model may be accessed to generate the likelihood score. The vehicle control module is configured to sum the token likelihood values at 508 to generate an overall likelihood score. For example, tokens, embeddings, etc., of the LLM may be compared with, for example, tokens or embeddings of the SLM or another model to determine an accuracy probability for each LLM token. These individual probability values for each token may be summed to generate an overall output probability score for the LLM output.
[0076] The vehicle control module is configured to determine an LLM confidence score based on the summed token probability values at 512. The LLM confidence score is compared to a specified threshold at 516. The specified threshold may be a confidence score indicating that the LLM output is likely to be accurate or correct.
[0077] If at 516, the LLM confidence score is greater than the specified threshold, control proceeds to 520 to output an LLM prompt to the user. For example, based on the LLM output, the system may generate an audible or textual response to the user, which is output via the vehicle speaker 26 or the user interface 22.
[0078] If at 516 the confidence score is not greater than the specified threshold, control proceeds to 524 to obtain a statistical language model confidence score at 524. The vehicle control module then compares the confidence scores from the SLM and the LLM at 528. Although Fig. 5 refers to an SLM and an LLM, in other embodiments, other suitable generative models of artificial intelligence (AI), other suitable models of natural language processing (NLP), etc. may be used.
[0079] If the LLM and SLM outputs have matching intent values and matching entities at 532, control continues to 536 to output a response from the SLM model. If the LLM and SLM outputs do not have matching intent values or entities, control continues to 540 to increment the confidence score for the LLM output. The vehicle control module is configured to output the LLM response and engage a user in an advancing N-turn dialog.
[0080] Fig. Figure 6 is a flowchart showing an example process for using guardrail data to restrict the output of a large language model. The process may be implemented, for example, by the vehicle control module 20 of Fig. 1, a user's mobile device, another processing device connected to the vehicle 10 or the user, etc.
[0081] At 604, the method begins by accessing the vehicle's location data, e.g., via a GPS antenna, a particular region assigned to the vehicle, etc. The vehicle's location data may, for example, indicate a region in which the vehicle is located, a particular city, state, county, country, etc.
[0082] The vehicle control module is configured to receive data on restricted topics at 608. For example, the discussion of certain political topics may be prohibited in certain geographic regions, or other sensitive topics that vary by location. The restricted topic data can specify information that should not be included in the output of a language model, and this information may vary depending on the vehicle's geographic location.
[0083] The vehicle control module is configured to generate a preliminary or intermediate response at 612 using a generative large AI language model. For example, the user's voice request may be submitted to the LLM to generate an intermediate response, but the intermediate response is not returned to the user until further processing and verification has occurred.
[0084] For example, the vehicle control module is configured to compare the intermediate response with the system rule set data and specified guardrail data at 616. The rule set and guardrail data may contain responses, topics, information, etc., that the LLM should not return to the user. This data may be defined, updated, etc., over time by a system administrator, a vehicle user, etc.
[0085] After comparing the intermediate response with the system rule set data and the guardrail data, if the intermediate response is a valid context at 620, control proceeds to 624 to output the intermediate response to a user. For example, after verifying the intermediate response is a valid output, the intermediate response may be output to the user via the vehicle speaker 26 or the user interface 22 as an audio signal or text.
[0086] If the intermediate response at 620 is not in a permissible context (e.g., because it contains an impermissible sensitive topic), the controller proceeds to 628 to update the prompt text using a natural language generation model. For example, a different language model may be used to generate a default response that does not contain impermissible information about sensitive topics compared to the intermediate response of the generative AI model.
[0087] The vehicle control module is configured to provide corrective inputs to the large language model at 632 for future corrections. For example, a replacement message or revised output that does not contain an illegal topic can be provided to the large language model, allowing the large language model to provide responses that do not contain illegal topics in the future.
[0088] Fig. Figure 7 is a flowchart illustrating an exemplary process for selecting between a large language model and a statistical language model based on confidence scores. The process may be performed, for example, by the vehicle control module 20 of Fig. 1, a user's mobile device, another processing device connected to the vehicle 10 or the user, etc.
[0089] In 704, the method begins by processing the user language with a large language model and a statistical language model. Although Fig. 7 refers to an SLM and an LLM, in other embodiments, other suitable generative models of artificial intelligence (AI), other suitable models of natural language processing (NLP), etc. may be used.
[0090] The controller is configured to compare the LLM output embeddings with the SLM output embeddings at 708 to generate a confidence score. If the confidence score is greater than a high threshold, the controller proceeds to 716 to respond to the user request using the LLM output.
[0091] If at 712 the confidence score is lower than the high threshold, the controller determines at 720 whether the confidence score is above a low threshold. In some examples, the high threshold may indicate a sufficient probability for an accurate LLM output such that the LLM output may be provided solely to the user.
[0092] The low threshold may indicate a moderate probability of an accurate LLM output, where the LLM output may be used in combination with the output of another model. For example, if the confidence score at 720 is above the low threshold, the controller proceeds to 724 to respond to the user request with a mixed regression based on a combination of the LLM output and the SLM output.
[0093] If at 720 the confidence score is below the low threshold, the controller proceeds to 728 to respond to the user request using the SLM output, possibly not using the LLM output at all. In the example of Fig. 7, multiple thresholds allow the use of only the LLM output, a combination of LLM output and SLM output, or the SLM output alone, based on different levels of confidence in the probability of the accuracy of the LLM output.
[0094] Fig. Figure 8 is a flowchart illustrating an exemplary process for speech processing, including performing acoustic emotion classification. The process may be performed, for example, by the vehicle control module 20 of Fig. 1, a user's mobile device, another processing device connected to the vehicle 10 or the user, etc.
[0095] At 804, the process begins by capturing the user's speech, e.g., via the vehicle microphone 24 or the user interface 22. The vehicle control module is configured to transcribe the speech using automatic speech recognition at 808. Any suitable ASR implementations may be used.
[0096] The vehicle control module is configured to detect user frustration and / or emotion levels at 812, e.g., by using an orthographic channel. For example, any suitable frustration detection or emotion detection algorithms can be used to process the user's speech and predict whether the words uttered by the user indicate that the user is frustrated, angry, or annoyed.
[0097] The vehicle control module is configured to perform acoustic emotion classification of the user speech at 816. For example, the audio signals of the user speech may be processed to determine whether the user's tone of voice indicates frustration, anger, etc. In various examples, any suitable algorithm for classifying acoustic emotions may be used.
[0098] The vehicle control module is configured to compare the user emotion score to a specified value at 820. For example, the result of textual frustration detection and the result of acoustic emotion detection can be combined and compared to individual thresholds to determine whether the user is currently frustrated or angry.
[0099] If at 824, the emotion value is greater than the specified threshold (e.g., indicating that the user still experiences pleasant or neutral emotions), control proceeds to 828 to use the LLM for further dialogue with the user. If at 824, the emotion value is less than the specified threshold (e.g., indicating that the user is experiencing frustration or anger from interacting with the LLM responses), control proceeds to 832 to switch to an N-gram-based language model or a finite-state grammar (FSG) model to continue the dialogue with the user. This allows the system to detect when a user is becoming frustrated by the responses of a generative AI model and switch to a more deterministic language response model at that time to avoid further frustration.
[0100] Fig. 9A and Fig. 9B shows an example of a neural network used to create models like those described above using machine learning techniques. Machine learning is a method for developing complex models and algorithms suitable for prediction (e.g., predictions for patient and provider matching). Models created using machine learning, like those described above, can produce reliable, repeatable decisions and outcomes and uncover hidden insights by learning from historical relationships and trends in the data.
[0101] The purpose of using a neural network-based model and training it with machine learning, as described above, can be to directly predict dependent variables without mathematically casting relationships between the variables. The neural network model includes a large number of virtual neurons operating in parallel and arranged in layers. The first layer is the input layer and receives the raw input data. Each subsequent layer modifies the outputs of a previous layer and passes them on to the next layer. Each subsequent layer optionally applies nonlinear transformation functions to the outputs of a previous layer before passing them on to the next layer. The final layer is the output layer and generates the system's output.
[0102] Fig. Figure 9A shows a fully connected neural network, where each neuron in a particular layer is connected to every neuron in the next layer. In the input layer, each input node is assigned a numerical value, which can be any real number. In each layer, each connection emanating from an input node is assigned a weight, which can also be any real number (see Fig. 9B). In the input layer, the number of neurons is equal to the number of features (columns) in a dataset. The output layer can have multiple continuous outputs.
[0103] The layers between the input and output layers are hidden layers. The number of hidden layers can be one or more (one hidden layer is likely sufficient for most applications). A neural network without hidden layers can represent linearly separable functions or decisions. A neural network with one hidden layer can perform a continuous mapping from one finite space to another. A neural network with two hidden layers can approximate any smooth mapping with arbitrary accuracy.
[0104] The number of neurons can be optimized. At the beginning of training, a network configuration is more likely to have redundant nodes. Some of the nodes can be removed from the network during training, which does not noticeably affect the network's performance. For example, nodes with weights approaching zero after training can be removed (this process is called pruning). The number of neurons can lead to underfitting (inability to adequately capture signals in the dataset) or overfitting (insufficient information to train all neurons; the network performs well on the training set but not on the test set).
[0105] Various methods and criteria can be used to measure the performance of a neural network model. The root mean squared error (RMSE), for example, measures the average distance between observed values and model predictions. The coefficient of determination (R2) measures the correlation (not the accuracy) between observed and predicted outcomes. This method may not be reliable if the data has a large variance. Other performance measures are irreducible noise, model bias, and model variance. High model bias for a model indicates that the model is unable to capture the true relationship between the predictors and the outcome. Model variance can provide an indication of whether a model is stable (a small perturbation in the data significantly changes the model fit). The neural network can adapt inputs, e.g.Vectors that can be used to create models that can be used in language processing, e.g. speech input from a driver or passenger of a vehicle.
[0106] Although Fig. 9A and Fig. 9B shows an example of neural networks, other embodiments may include other types of models or more specific types of neural networks. For example, transformers may be used for large language models; long short-term memory (LSTM) models may be used in some examples, and so on.
[0107] Fig. Figure 10 shows an example of creating a machine learning model. At 907, the controller receives data from a database 902 (e.g., a data warehouse). The data can be any data suitable for developing machine learning models.
[0108] At 911, the controller separates the data obtained from the database 902 into training data 915 and test data 919. The training data 915 is used to train the model at 923, and the test data 919 is used to test the model at 927. Typically, the amount of training data 915 is chosen to be larger than the amount of test data 919, depending on the desired parameters for model development. For example, the training data 915 can comprise approximately seventy percent of the data from the database 902, approximately eighty percent of the data, approximately ninety percent, etc. The remaining thirty percent, twenty percent, or ten percent, respectively, are then used as test data 919.
[0109] Separating a portion of the acquired data as test data 919 allows testing of the trained model against the actual output data to enable more accurate training and development of the model at 923 and 927. The model may be trained at 923 using any suitable machine learning modeling techniques, including those described herein, such as random forest, generalized linear models, decision tree, and neural networks.
[0110] At 931, the controller evaluates the results of the model test. For example, at 927, the trained model may be tested against the test data 919, and the results of the output data of the tested model may be compared with the actual outputs of the test data 919 to determine a degree of accuracy. The model results may be evaluated using any suitable machine learning model analysis, such as the exemplary techniques described below.
[0111] After evaluating the model test results at 931, the model can be deployed at 935 if the model test results are satisfactory. Deploying the model may involve using the model to generate predictions for a large input dataset with unknown outputs. If the evaluation of the model test results at 931 is not satisfactory, the model can be further developed using different parameters, different modeling techniques, different model types, etc. The machine learning model procedure of Fig. 10 can receive inputs, e.g., vectors, that can be used with language processing, such as speech input from a driver or passenger of a vehicle. In some embodiments, a machine learning model can be trained through unsupervised learning, e.g., by training generative AI models or LLMs.
Claims
[1] Vehicle user interface system comprising: at least one vehicle loudspeaker (26) configured to generate audio signals within a vehicle (10); a vehicle user interface (22) having a screen configured to display text; at least one vehicle microphone (24) configured to pick up the speech of a vehicle occupant; and a vehicle control module (20) configured to: Receiving voice inputs from the vehicle occupant via the at least one vehicle microphone (24); Classifying the speech input as a deterministic speech request or a probabilistic speech request; in response to the speech input being classified as a deterministic language request: processing the speech input using a statistical language model (SLM) to produce an SLM output; in response to the speech input being classified as a probabilistic language request: processing the speech input using a large language model (LLM) to produce an LLM output; and Generating an audio response output using the at least one vehicle loudspeaker (26) and / or a text response output using the screen of the vehicle user interface (22), wherein the audio response output or the text response output is based on the SLM output generated by the statistical language model or the LLM output generated by the large language model. [2] The vehicle user interface system of claim 1, wherein the vehicle control module (20) is configured to automatically modify the operation of at least one vehicle component (14) in response to the voice input including a request from the occupant to operate the at least one vehicle component (14). [3] The vehicle user interface system of claim 2, wherein automatically modifying the operation of the at least one vehicle component (14) comprises at least one of the following: initiating a telephone call via the vehicle user interface (22), sending a message via the vehicle user interface (22), activating an entertainment function of the vehicle user interface (22), or changing at least one driving setting of the vehicle (10). [4] The vehicle user interface system of claim 1, wherein the vehicle control module (20) is configured to: Obtaining model output guardrail data from a database (902) having stored output data about sensitive topics; Processing the speech input using the large language model (LLM) to generate an intermediate output response; Comparing the intermediate output response with the model output guardrail data; and Prevent the intermediate output response from being output if the intermediate output response contains a disallowed model output guardrail data topic. [5] A vehicle user interface system according to claim 4, wherein: the vehicle control module (20) is configured to determine a current geographical location of the vehicle (10); the output data on sensitive topics stored in the database (902) vary depending on the geographical location; and comparing the intermediate output response with the model output guardrail data comprises comparing the intermediate output response only with output data about sensitive topics corresponding to the current geographical location of the vehicle (10). [6] A method of operating a vehicle user interface system, the method comprising: Receiving voice inputs from a vehicle occupant using at least one vehicle microphone (24); Classifying the speech input as a deterministic speech request or a probabilistic speech request; in response to the speech input being classified as a deterministic language request: processing the speech input using a statistical language model (SLM) to produce an SLM output; in response to the speech input being classified as a probabilistic language request: processing the speech input using a large language model (LLM) to produce an LLM output; and Generating an audio response output using at least one vehicle loudspeaker (26) and / or a text response output using a screen of a vehicle user interface (22), wherein the audio response output or the text response output is based on the SLM output generated by the statistical language model or the LLM output generated by the large language model.
Citation Information
Patent Citations
Construction method of rewriting model, display equipment and sentence rewriting method
CN118378614A
Procedures for operating a speech dialogue system and speech dialogue system
DE102019217751A1
Method for operating a digital assistant of a vehicle, computer-readable medium, system, and vehicle
DE102023115462A1
CN000118378614A