Multimodal human-machine interface for a vehicle
The human-machine interface addresses imprecise spoken utterances by converting them to written form and using an LLM for precise responses, ensuring accurate interaction with vehicle display content and enhancing user satisfaction.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- AUDI AG
- Filing Date
- 2025-11-25
- Publication Date
- 2026-06-04
AI Technical Summary
Existing human-machine interfaces in vehicles often misunderstand spoken utterances due to imprecise or incomplete formulations, leading to inappropriate reactions and user dissatisfaction.
A human-machine interface that converts spoken utterances into written form, uses a Large Language Model (LLM) to determine a response based on the written utterance, and outputs the response audibly, incorporating a topic recognition module and a knowledge database to provide precise answers related to displayed content.
Ensures accurate responses to spoken queries about vehicle display content, enhancing user satisfaction by providing clear and relevant information, thus improving interaction precision and reducing negative perceptions.
Smart Images

Figure EP2025084171_04062026_PF_FP_ABST
Abstract
Description
[0001] AUDI AG P24541WO.O
[0002] Multimodal human-machine interface for a vehicle
[0003] DESCRIPTION:
[0004] The invention relates to a method for operating a human-machine interface of a vehicle, in which a speech converter of a human-machine interface of a vehicle converts a spoken utterance of a vehicle occupant, captured by a microphone of the human-machine interface, into a written utterance, an LLM module of the human-machine interface determines a written response to the captured spoken utterance depending on the written utterance, the speech converter converts the determined written response into a spoken response, and a loudspeaker of the human-machine interface outputs the spoken response. The invention further relates to a human-machine interface for a vehicle and a vehicle.
[0005] A human-machine interface, commonly abbreviated as HMI (Human Machine Interface), generally serves to enable a user to operate and / or monitor a machine. Accordingly, a human-machine interface comprises an input device, such as a switch, a touch-sensitive sensor surface, or a microphone, and / or an output device, such as a display screen or a speaker.
[0006] In particular, vehicles are considered machines and vehicle occupants are considered users in this sense. Accordingly, a vehicle typically comprises multiple human-machine interfaces. A human-machine interface can be operated flexibly with regard to the output medium. For example, DE 10 2022 001 872 A1 discloses a method for operating a vehicle's human-machine interface, in which the vehicle's human-machine interface determines the vehicle's output medium based on a sensor-detected driving situation and a determined characteristic of a vehicle occupant.
[0007] A human-machine interface can provide assistance to the occupant, especially the driver of the vehicle, in controlling the vehicle, i.e., in operating previously unknown or at least rarely used controls of the vehicle, which themselves can also be considered human-machine interfaces of the vehicle.
[0008] For this purpose, DE 10 2013 011 531 A1 discloses a method for operating a human-machine interface of a vehicle, in which an occupant of a vehicle indicates a control element of the vehicle with a finger, a human-machine interface of the vehicle visually detects the finger, identifies the control element indicated by the finger on the basis of the visually detected finger and outputs information concerning the identified control element.
[0009] Furthermore, a human-machine interface can save the occupant from having to tactilely operate the vehicle's controls, thereby making it easier to control the vehicle.
[0010] Such a method for operating a human-machine interface of a vehicle is described in DE 10 2017 011 498 A1. In this method, a vehicle occupant gestures to indicate an area within the vehicle and makes a spoken utterance. A human-machine interface of the vehicle visually detects the gesture and, based on the visually detected gesture, recognizes the indicated area. The human-machine interface also detects the occupant's spoken utterance and identifies a target of the detected utterance. If the gesture and the spoken utterance are detected within a predetermined time interval, the human-machine interface issues a control command to the vehicle corresponding to the detected area and target.
[0011] A human-machine interface for a vehicle, operating in the manner described above, allows a vehicle occupant to interact with the vehicle using spoken language, i.e., to communicate verbally in a bidirectional manner. This capability is provided by an artificial neural network specialized for context-sensitive verbal interaction, commonly referred to as a Large Language Model (LLM). Thanks to the human-machine interface, the occupant's hands remain free during interaction, and distraction is minimal, thus significantly enhancing vehicle safety while driving.
[0012] However, the occupant often does not formulate their request precisely and / or completely. For example, the occupant's statements may not include the expected words and / or the context of the respective request, so that the human-machine interface may misunderstand the request and react inappropriately. Based on inappropriate reactions from the human-machine interface, the occupant may then – albeit incorrectly – conclude that the human-machine interface is of insufficient quality, which should be avoided.
[0013] One object of the invention is to propose a method for operating a human-machine interface of the vehicle, enabling the human-machine interface to precisely respond to a spoken utterance of a vehicle occupant concerning a display content shown by a display unit of the vehicle. Further objects of the invention are to provide a human-machine interface for a vehicle and a vehicle.An object of the invention is a method for operating a human-machine interface of a vehicle, in which a speech converter of the human-machine interface of a vehicle converts a spoken utterance of a vehicle occupant, captured by a microphone of the human-machine interface, into a written utterance; an LLM module of the human-machine interface determines a written response to the captured spoken utterance, depending on the written utterance; and the speech converter converts the determined written response into a spoken response, and a loudspeaker of the human-machine interface outputs the spoken response. In other words, the human-machine interface receives spoken language and responds with spoken language. In short, the interaction is entirely acoustic or auditory.
[0014] Ideally, a speech recognition module first recognizes a wake-up word (WuW) and activates the human-machine interface when the recognized wake-up word is detected.
[0015] The speech converter works bidirectionally and is trained to transcribe spoken language (Speech2Text) and to pronounce written language (Text2Speech). The LLM module is trained to communicate in written language, i.e., to understand written language and to respond appropriately to the understood written language, for example, to answer a written question in writing. The speech converter can be provided separately from the LLM module or be part of the LLM module. Alternatively, the LLM module can be trained to communicate directly in spoken language, i.e., without a conversion step.
[0016] According to the invention, a topic recognition module of the human-machine interface recognizes a topic in the written utterance and, if the recognized topic relates to the displayed content, a display unit of the human-machine interface takes a screenshot of the content shown by the vehicle's display unit. Based on the screenshot, a knowledge database of the human-machine interface provides a written explanation of the displayed content. The LLM module then determines the written response based on the provided written explanation. The topic recognition module can recognize predetermined keywords in the written utterance, each corresponding to topics relevant to the vehicle, and derive the topic of the utterance from these recognized keywords.
[0017] If the detected topic relates to current display content shown by a human-machine interface display unit, the screenshot can provide a context for the spoken utterance that is explicitly missing in the spoken utterance.
[0018] For example, so-called deictic questions are spoken utterances with incomplete context. While they may include a demonstrative pronoun and / or an adjectival attribute as a context reference, the demonstrative pronoun and / or the adjectival attribute only points to a context and does not explicitly specify it.
[0019] Preferably, a warning symbol encompassed by the display content is recognized as the subject. Warning symbols are generally displayed infrequently and are therefore unfamiliar to the vehicle occupant, and in many cases not self-explanatory or suggestive. A deictic question from the occupant in this regard might be: "What does this red warning symbol mean?" According to the invention, the human-machine interface can answer this deictic question precisely as follows: "The red warning symbol is an engine alarm."
[0020] The knowledge database can be pre-trained with a large number of images and their associated texts. This prepares the database for potential spoken utterances by the inmate. For training purposes, it is advantageous to encode the majority of images using an image encoder and / or the majority of texts using a text encoder. The encoder transforms the images and texts into numbers, which can be components of mathematical vectors.
[0021] Ideally, text accompanying an image includes instructions for the occupant. Thanks to these instructions, the occupant in the example above not only learns the meaning of the warning symbol, but also receives guidance on an appropriate response, such as: "Stop the vehicle and check the engine!"
[0022] Preferably, the written explanation is provided as a text from a text-image pair in the knowledge database that exhibits maximum cosine similarity to a text / image pair formed from the written statement and the screenshot. Maximum cosine similarity corresponds to the highest probability of the written explanation being correct.
[0023] Another aspect of the invention is a human-machine interface for a vehicle, comprising a display unit, a microphone, a speech recognition module, a speech converter, an LLM module, and a loudspeaker. The aforementioned components are known per se and, in the specified combination, can form part of a human-machine interface for a vehicle.
[0024] According to the invention, the human-machine interface comprises a topic recognition module and a knowledge database, and the human-machine interface is configured to execute a method according to an embodiment of the invention. The topic recognition module, the knowledge database, and the ability of the human-machine interface to execute the method according to the invention during normal operation enable the human-machine interface to respond precisely to a spoken utterance from a vehicle occupant concerning display content shown by a vehicle display unit, even if the spoken utterance is not precisely and / or completely formulated.
[0025] The display unit can be configured as a central display screen of a vehicle's infotainment system or as a vehicle instrument cluster. The central display screen, also known as the Central Information Display (CID), can be located in the vehicle's center console. The instrument cluster can be located in the vehicle's instrument panel. The central display screen and the instrument panel are examples and not intended to be limiting. Display units according to the invention can also be located in the vehicle's headrests, side panels, headliner, or armrest.
[0026] Preferably, the knowledge database is designed as a two-dimensional multimodal database. In other words, the knowledge database comprises vectors with two different components as data records: vectors with a text component and vectors with an image component. Accordingly, a mathematical dot product can be defined for any two vectors in the database, as is standard practice. The value of this dot product is given as the product of the magnitudes of the two vectors and the cosine of an angle enclosed by the two vectors. For two parallel vectors, the enclosed angle is 0°, the cosine takes on the maximum value of 1, and consequently, the dot product of the two vectors is maximal.
[0027] Another object of the invention is a vehicle. By way of example, and without limitation, the vehicle is a road vehicle, in particular a passenger car. However, vehicles within the meaning of the invention also include trucks, lorries, and other types of road vehicles, rail vehicles such as locomotives or wagons, watercraft such as boats or ships, and aircraft such as helicopters or airplanes. Accordingly, numerous and diverse applications of the invention are possible.
[0028] According to the invention, the vehicle comprises a human-machine interface according to one embodiment of the invention. By means of the human-machine interface according to the invention, an occupant of the vehicle can communicate verbally with the vehicle with high precision.
[0029] A key advantage of the method according to the invention is that a spoken utterance by a vehicle occupant concerning content displayed by a vehicle display unit is precisely answered by a human-machine interface. This avoids negative perceptions of the human-machine interface and increases user satisfaction with the vehicle.
[0030] The invention is schematically illustrated in the drawing with reference to one embodiment and is further described with reference to the drawing. It shows:
[0031] Figure 1 shows a component-wise view of a human-machine interface according to an embodiment of the invention for a vehicle;
[0032] Figure 2 shows an exemplary view of the display content shown by the display unit of the human-machine interface shown in Figure 1.
[0033] Figure 1 shows a component-wise view of a human-machine interface 1 according to an embodiment of the invention for a vehicle 3. The human-machine interface 1 can belong to a vehicle 3 according to the invention. The human-machine interface 1 comprises a display unit 17, which can be configured as a central display screen of an infotainment system of the vehicle 3 or as an instrument cluster of the vehicle 3.
[0034] Furthermore, the human-machine interface 1 includes a microphone 10, a speech recognition module 11, a speech converter 12, a topic recognition module 13, an LLM module 14, a knowledge database 15, which can be configured as a 2-dimensional multimodal database, and a loudspeaker 16.
[0035] The human-machine interface 1 is designed to execute a method according to an embodiment of the invention as follows during operation of the human-machine interface 1.
[0036] The knowledge base 15 is preferably trained beforehand using a majority of images and their associated texts. Ideally, for training the knowledge base 15, the majority of images are encoded using an image encoder and / or the majority of texts are encoded using a text encoder. A text associated with an image can contain instructions for an occupant 2 of the vehicle.
[0037] During normal operation of the human-machine interface 1, the speech recognition module 11 of the human-machine interface 1 can recognize a wake-up word in a spoken utterance of an occupant 2 of the vehicle 3 captured by the microphone 10 of the human-machine interface 1 and activate the human-machine interface 1.
[0038] The speech converter 12 of the human-machine interface 1 of the vehicle 3 converts the captured spoken utterance into a written utterance 4.
[0039] The topic recognition module 13 of the human-machine interface 1 recognizes a topic of the written utterance 4. Figure 2 shows an exemplary view of a display content 8 shown by the display unit 17 of the human-machine interface 1 shown in Figure 1. For example, a warning symbol 80 encompassed by the display content 8 is recognized as the topic.
[0040] The display unit 17 of the human-machine interface 1 takes a screenshot 5 of the display content 8 shown by the display unit 17 if the detected topic concerns the displayed content 8.
[0041] The knowledge database 15 of the human-machine interface 1 provides a written explanation 6 of the displayed content 8 based on the screenshot 5. The written explanation 6 is preferably a text from a text / image pair in the knowledge database 15 that exhibits maximum cosine similarity with a text / image pair formed from the written utterance 4 and the screenshot 5.
[0042] The LLM module 14 of the human-machine interface 1 determines a written response 7 to the recorded spoken utterance, depending on the written utterance 4. The LLM module 14 further determines the written response 7 based on the provided written explanation 6.
[0043] The speech converter 12 converts the determined written answer 7 into a spoken answer and the loudspeaker 16 of the human-machine interface 1 outputs the spoken answer.
[0044] REFERENCE MARK LIST:
[0045] 1 Human-Machine Interface
[0046] 10 Microphone 11 Speech recognition module
[0047] 12 language converters
[0048] 13 Topic Recognition Module
[0049] 14 LLM module
[0050] 15 Knowledge database 16 Speakers
[0051] 17 Display unit
[0052] 2 occupants
[0053] 3 vehicles
[0054] 4. Written statement 5. Screenshot
[0055] 6. Written explanation
[0056] 7 written answers
[0057] 8 Warning symbol
Claims
PATENT CLAIMS:
1. Method for operating a human-machine interface (1) of a vehicle (3), wherein a speech converter (12) of a human-machine interface (1) of a vehicle (3) converts a spoken utterance of an occupant (2) of the vehicle (3), captured by a microphone (10) of the human-machine interface (1), into a written utterance (4); an LLM module (14) of the human-machine interface (1) determines a written response (7) to the captured spoken utterance, depending on the written utterance (4); the speech converter (12) converts the determined written response (7) into a spoken response; and a loudspeaker (16) of the human-machine interface (1) outputs the spoken response.a topic recognition module (13) of the human-machine interface (1) recognizes a topic of the written utterance (4) and a display unit (17) of the human-machine interface (1) takes a screenshot (5) of a display content (8) shown by the display unit (17) if the recognized topic relates to the displayed content (8); a knowledge database (15) of the human-machine interface (1) provides a written explanation (6) of the displayed content (8) based on the screenshot (5); the LLM module (14) additionally determines the written response (7) depending on the provided written explanation (6).
2. Method according to claim 1, wherein a warning symbol (8) encompassed by the display content (8) is recognized as the subject.
3. Method according to claim 1 or 2, wherein the knowledge database (15) is pre-trained with a plurality of images and texts associated with the images.
4. Method according to claim 3, wherein, for training the knowledge database (15), the plurality of images are encoded using an image encoder and / or the plurality of texts are encoded using a text encoder.
5. Method according to claim 3 or 4, wherein a text associated with an image comprises instructions for the occupant (2).
6. Method according to one of claims 1 to 5, wherein the written explanation (6) is provided as a text of a text-image pair of the knowledge database (15) which has a maximum cosine similarity with a text / image pair formed from the written utterance (4) and the screenshot (5).
7. Human-machine interface (1) for a vehicle (3) comprising a display unit (17), a microphone (10), a speech recognition module (11), a speech converter (12), a topic recognition module (13), an LLM module (14), a knowledge database (15) and a loudspeaker (16) and configured to perform a method according to any one of claims 1 to 6.
8. Human-machine interface according to claim 7, wherein the display unit (17) is designed as a central display screen of an infotainment system of the vehicle (3) or as a combination instrument of the vehicle (3).
9. Human-machine interface according to claim 7 or 8, wherein the knowledge database (15) is designed as a 2-dimensional multimodal database.
10. Vehicle (3) comprising a human-machine interface (1) according to any one of claims 7 to 9.