Controlling a call connection from a motor vehicle
By leveraging a vehicle's camera system to recognize mouth movements and process speech signals, the method enhances speech quality in vehicle calls, addressing the limitations of existing audio-focused solutions and reducing costs.
Patent Information
- Application Number
- PCT/DE2025/100378
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-08
- Filing Date
- 2025-04-14
- Publication Date
- 2026-01-15
AI Technical Summary
Existing methods for improving audio quality in vehicle voice calls, such as using high-quality microphones and active noise cancellation, are costly and often insufficient, necessitating further improvements to enhance speech clarity amidst background noise.
Utilizing a vehicle's existing camera system to recognize mouth and facial movements to process speech signals, enhancing speech quality by subtracting interference and adding missing components, without requiring additional hardware.
Achieves significant improvement in speech quality during vehicle calls by filtering out background noise and restoring incomplete speech, providing clear and natural-sounding conversations without additional hardware costs.
Smart Images

Figure DE2025100378_15012026_PF_FP_ABST
Abstract
Description
[0001] Controlling a call connection from a motor vehicle
[0002] The invention relates to a method for controlling a telephone connection from a motor vehicle, a corresponding control device for a motor vehicle, and a motor vehicle with such a control device.
[0003] It is known that conversation or voice connections from a motor vehicle are used in which a head unit, a hands-free device, etc., records the conversation share of the driver or another vehicle occupant in the form of an audio signal and transmits it to the other party of the connection.
[0004] During conversations or phone calls made from inside a vehicle, background noise is inevitably recorded. This noise can originate from the engine, the rolling of the wheels, wind noise, or rattling or creaking vehicle parts. Background noise can also come from inside the vehicle, such as from other occupants, their conversations or phone calls, or even entertainment noise. Other sources of interference include sirens, horns, engine, road, or tire noise from other nearby vehicles, and even snippets of conversation from nearby pedestrians that can be heard inside the vehicle.
[0005] Several measures are already known to improve audio quality in in-vehicle voice calls. For example, high-quality microphones can be used. Multiple microphones can be installed, and / or their mounting positions can be optimized. Other possible measures include acoustic attenuation of interference affecting voice quality, "subtraction" of background noise picked up by reference microphones, or targeted addition of "negative sound" or "anti-sound" to reduce background noise ("Active Noise Cancellation," ANC). Even when combining several of these measures, a need for improved voice quality often remains. However, further improvements using conventional methods can only be achieved through costly measures such as installing additional microphones, etc.
[0006] One of the problems underlying the present invention is the provision of improved methods for controlling telephone connections from a motor vehicle, as well as corresponding control devices and appropriately equipped motor vehicles. The invention solves this problem by means of the subject matter of the independent claims. Dependent claims describe preferred embodiments.
[0007] A first aspect of the present invention comprises a method for controlling a conversation from a motor vehicle. The method comprises capturing a speech signal stream that includes a portion of the conversation from a speaking vehicle occupant; capturing a video signal stream that includes a portion of the speaking vehicle occupant's face; recognizing sequences of sounds spoken by the vehicle occupant based on the occupant's mouth movements; and processing the speech signal stream based on the recognized sequences of sounds to improve the outgoing speech signal quality of the conversation.
[0008] In contrast to known measures that focus on improving the audio signal through more complex hardware for sound recording or acoustic signal acquisition, improved methods for filtering the audio signal, etc., the invention uses a completely different type of additional signal source to improve the speech signal quality, namely a video signal provided by a camera system of the vehicle.
[0009] Embodiments of the invention can thus achieve a noticeable improvement in outgoing speech quality. Since a camera system is already present in many vehicles, an improved speech signal can be achieved without the need for additional hardware; that is, embodiments of the invention are particularly advantageous from a cost perspective.
[0010] The voice or audio connection can be a telephone call or, more generally, a communication connection originating from within the vehicle. For example, it could be a traditional telephone call without video, or it could be a video call where an audio and video signal are transmitted in one or both directions, i.e., a combined audio / video signal (abbreviated as "AA / -signal"). When the voice connection is an AA / -connection, it is advantageous that the video signal stream used to recognize the sound sequences includes the video component of the connection. This eliminates the need for any additional video data to be acquired for improved voice quality; no separate camera needs to be activated, etc.
[0011] In embodiments of the invention, a microphone system of the vehicle can detect sound signals or acoustic signals and convert them into electrical signals that represent the detected sound signals in analog or digital form; that is, the electrical signals represent a portion of the conversation from the speaking vehicle occupant in the form of, for example, a continuous speech signal stream or data stream. When a "signal stream" is mentioned, this refers to a signal sequence that is at least partially continuous. Occasionally, the term "signal" is used for short; however, whether this refers to an instantaneous signal or a signal stream is clear to a person skilled in the art from the context. Similarly, a camera system of the vehicle can record a video signal or video signal stream, which, for example,A continuous visual representation of the speaking vehicle occupant is represented, in particular the face of the vehicle occupant. Besides the face, the visual representation can also include, for example, the shoulders and / or part of the upper body of the occupant, but also, for example, one or both hands on the steering wheel, with one or both hands making finger movements or gestures, etc.
[0012] In some implementations, the recognition of sounds, sequences of sounds, a word, several words, etc., includes the interpretation and, if necessary, plausibility check of mouth movements captured or recorded in the video. Such mouth movements are understood to mean lip movements, but additionally or alternatively also potentially visible tongue movements. Sound recognition can also include determining whether the vehicle occupant is currently speaking, and if so, what the occupant is saying, or whether the occupant is, for example, merely keeping their mouth open without speaking.
[0013] In some embodiments, facial expressions of the vehicle occupant are used to recognize sound sequences in addition to their mouth movements; for example, movements of the cheek, eyebrows, including eye movements, etc. This can make the recognition of spoken sound sequences even more reliable, but technically requires that the camera system captures the driver's face, etc., comprehensively enough and that the face is adequately illuminated. If necessary, the driver or other speaking vehicle occupant can be asked by an AI system to, for example, remove sunglasses, take a cigarette out of their mouth, remove chewing gum, candy, etc., for the conversation. In some embodiments, speech components from the other party in the conversation are also used to recognize sound sequences. In this way, plausible reactions, answers, etc., can be detected.The speech of the speaking vehicle occupant can be deduced with even greater accuracy. This requires that the corresponding control device, the AI algorithm, etc., is made aware not only of the vehicle occupant's contribution to the conversation, but also of the other party's contribution.
[0014] In some designs, processing the speech signal stream includes subtracting an interference signal. Such an interference signal can be detected, for example, using conventional methods such as reference microphones.
[0015] However, it is particularly advantageous if the recognition of sound sequences includes the provision of a plausible speech signal, such as one derived from words that, based on the captured lip movements, it is plausible that the driver or vehicle occupant is currently speaking or has spoken very recently. Such a plausible signal can additionally take into account the recorded speech signal as well as other influences, as discussed herein.
[0016] The plausible speech signal could, for example, be the most plausible of several speech signals identified by an initial AI algorithm, perhaps based solely on lip-reading. A further model (and / or a conventional algorithm) can then perform a plausibility check of the various speech signals, using additional data. In other words, the plausible speech signal could, for example, consist of a possible sequence of sounds, a sequence of spoken words, etc., for which the highest confidence values were found (compared to other potentially spoken sounds or words). It should be noted that such confidence values may not be explicitly determined if an algorithm such as an AI ("Artificial Intelligence") algorithm is used for plausibility checks (or if only a single AI algorithm is used for, for example, lip-reading and / or plausibility checks).
[0017] The interference signal can be determined as the difference between the recorded speech signal stream and the plausible speech signal. The interference signal can then be subtracted from the recorded speech signal stream. With such designs, background noise can be separated particularly well, and excellent speech quality can be achieved.
[0018] In some embodiments, processing the speech signal stream can include adding a missing speech component. For example, a sound, a sequence of sounds, part of a word, a word, several words, etc., can be added that were not recorded by the microphone system or that are not recognizable or only fragmentarily recognizable in the audio signal, e.g., due to an insufficient signal-to-noise ratio, as can occur with particularly strong interference, or if the speaker is turned away from the microphone, etc.
[0019] In some of these embodiments, processing the speech signal stream involves adding a sound not uttered by the vehicle occupant. For example, the driver might speak briefly, perhaps due to loud background noise, a traffic situation, etc., and presenting a more detailed utterance from the driver can be advantageous to facilitate understanding of the intended or spoken message for the other party, avoid confusion, etc.
[0020] In some designs, the intensity or volume of what is said can be reduced (or increased) additionally or alternatively, e.g., if the speaking vehicle occupant speaks very loudly to drown out background noise; this is particularly advantageous if the background noise is filtered out, because then the other party has no apparent reason for speaking loudly, i.e., a reduction in volume may better meet the expectations of the other party.
[0021] Another aspect of the present invention relates to a control device for controlling a conversation from a motor vehicle. The control device comprises an audio input, a video input, a computer module, and a processing unit. The audio input is configured to provide a speech signal stream received from a microphone system of the motor vehicle, wherein the speech signal stream includes a portion of the conversation from a speaking vehicle occupant. The video input is configured to provide a video signal stream received from a camera system of the vehicle, which includes a portion of the facial features of the speaking vehicle occupant. The computer module is configured to use the video signal stream to recognize spoken sequences of sounds based on the mouth movements of the vehicle occupant.The processing unit is designed to process the speech signal based on the recognized sound sequences in order to improve the outgoing speech signal quality of the conversation. Such a control device can implement particularly advantageous embodiments of the invention.
[0022] For example, in some embodiments of the control device, the AI module can include a language model. This language model can be, for instance, a Large Language Model (LLM), a Large Action Model (LAM), and / or a multimodal model, LLM (Large Multimodal Model). The high performance of these and other models, including future ones, is advantageous, as it allows, for example, real-time improvement of outgoing speech signal quality. This means delays are imperceptible or tolerable for the other party and / or the speaking vehicle occupant, such as less than or equal to 0.3 seconds, 0.1 seconds, 30 milliseconds, 10 milliseconds, 3 milliseconds, or 1 millisecond.
[0023] In some embodiments, the AI module can implement an intelligent personal assistant (IPA) for the speaking vehicle occupant. This IPA can be provided by the vehicle, a vehicle manufacturer, etc. In some embodiments, the assistant is specifically trained on the respective vehicle occupant, e.g., the driver. This can contribute to optimal recognition of spoken sounds or words, for example, because the assistant is particularly familiar with the individual face of the vehicle occupant and their mouth movements, facial expressions, etc., as well as their individually preferred vocabulary, semantics, syntax, etc.
[0024] A further aspect of the present invention relates to a motor vehicle comprising a microphone system, a camera system, and a control device described herein. In some embodiments of such a vehicle, the invention allows for very good voice quality for outgoing calls, video calls, etc. In some embodiments, it is even conceivable that the costly installation of multiple microphones can be avoided because an embodiment of the present invention is implemented instead.
[0025] The invention will now be described in more detail with reference to the attached drawings, in which:
[0026] Figure 1 shows a motor vehicle;
[0027] Figure 2 shows a control device; and
[0028] Figure 3 illustrates a flowchart of a process. Figure 1 schematically shows a motor vehicle 100 with a vehicle occupant 102, e.g. the driver (the terms "vehicle occupant" and "driver" are occasionally used synonymously here).
[0029] Vehicle 100 includes a microphone system 104, of which only a single microphone is indicated, but which can also include multiple microphones, e.g., in the interior of vehicle 100, for example, a microphone for the driver and a microphone for the passenger. The microphone system 104 can include microphones installed in vehicle 100 and / or can include additional microphones, such as a microphone from a smartphone belonging to person 102.
[0030] The vehicle 100 further includes a camera system 106, of which only a single camera is indicated, but which can include several interior cameras, e.g., a camera facing the driver and a camera facing a passenger. The camera system 106 can include cameras installed in the vehicle 100 and / or can use one or more (additional) cameras, e.g., the camera of a smartphone belonging to person 102.
[0031] Microphone system 104 and camera system 106 are connected to a control device 108, which may be implemented, for example, in a head unit (not shown) and / or an ECU ("Electronic Control Unit") of the vehicle 100, but may also be designed as a standalone module. A conversation 110 between the vehicle occupant 102 and a (not shown) counterpart outside the vehicle 100 comprises an outgoing conversation component 112 and an incoming conversation component 114. For transmitting the conversation 110, the vehicle 100 may, for example, have an antenna system 116.
[0032] Fig. 2 shows a functional block diagram illustrating the structure of the control device 108 in further detail. The control device 108 comprises an audio input 120, a video input 122, a computer module 124, and a processing unit 126.
[0033] The audio input 120 receives a speech signal stream 130 from the microphone system 104 and makes it available to the processing unit 126. The speech signal stream 130 is represented in the form of an electrical signal (or as a signal waveform that is at least segmentally continuous). Although it is described as such, the incoming or provided speech signal stream 130 represents or includes not only the acoustic contribution of the vehicle occupant 102 to the conversation 110, but also sound or acoustic signals from other sources of interference in or on the vehicle 100 and from outside.
[0034] Although Fig. 2 shows only one input 120 for, e.g., an audio track from a microphone of the microphone system 104, the input 120 can include several physical or logical inputs, e.g., for several audio tracks from several microphones of the microphone system 104.
[0035] Video input 122 receives a video signal stream 132 from camera system 106 and makes it available to the control module 124. The video signal stream 132 is represented in the form of an electrical signal (or as a signal waveform that is at least segmentally continuous). The video signal stream represents, or in particular comprises, a continuous view of the vehicle occupant 102 from the perspective of camera system 106. For example, at least one camera can capture the face of the vehicle occupant 102. Although only one input 122 is indicated, the video signal stream 132 can comprise video sequences from several cameras, for example, a driver's camera capturing a frontal view of the driver, a passenger's camera capturing the driver's face from the side, etc.
[0036] The AI module 124 receives the provided video signal stream 132 and evaluates it to interpret mouth and lip movements, and possibly facial expressions, of the vehicle occupant 102, i.e., to recognize sounds, sequences of sounds, words, multiple words, sentences, etc., spoken by this occupant 102. The AI module 124 can, for example, include a video-to-text algorithm that recognizes spoken words based on lip reading and outputs the determined sound sequence as an audio signal 134. This audio signal 134, which will be discussed in more detail later, is provided by the AI module 124 to the processing unit 126.
[0037] The processing unit 126 processes the provided speech signal stream 132 to improve the quality of the outgoing speech signal or conversation component 110. For this purpose, the processing unit 126 uses the audio signal 134 determined and provided by the AI module 124. Exemplary embodiments of the operation of AI module 124 and processing unit 126 are discussed below with reference to Fig. 3.
[0038] Fig. 3 shows in schematic form a sequence 200 of a procedure in the control device 108 for controlling the conversation connection 110, more precisely the conversation part 112 originating from the vehicle 100.
[0039] The procedure begins with an operation 202, which may, for example, involve activating the control device 108 and its components, triggered, for instance, by an incoming call or by the vehicle occupant 102 making a call to initiate the communication 110. For example, the microphone system 104 may be activated. For a video call, the video or camera system 106 may also be activated or made available.
[0040] Step 204 comprises the recording of the speech signal stream 130 by the microphone system 104 and the provision of the recorded speech signal stream 130 at the audio input 120 of the control device 108. The speech signal stream 130 contains, in particular, the conversational part 112 of the vehicle occupant 102 in the conversation 110, and also typically an interference signal that can originate from a wide variety of interference sources inside and outside the vehicle 100. In some embodiments, it is conceivable that the audio signal 130 provided at the audio input 120 has already undergone conventional preprocessing to minimize the influence of interference sources.In other embodiments, such preprocessing can be dispensed with or reduced because the methods according to the invention and described herein lead to a significant improvement in speech quality, so that separate, costly preprocessing (and / or postprocessing) can be eliminated.
[0041] Step 206 comprises the recording of the video signal stream 132 by the video or camera system 106 and the provision of the recorded video signal stream 132 at the video input 122 of the control device 108. The video signal stream 132 represents, in particular, an appearance of the vehicle occupant 102 during the conversation, i.e., the existence of the conversation connection 110. The appearance can capture not only the face (including the mouth area) of the speaking vehicle occupant 102, but also parts of the upper body, shoulders, etc., and can optionally also capture gestures that the occupant 102 performs during the conversation. For example, a camera can capture one or both hands on the steering wheel if one hand performs an involuntary or intentional gesture in the air, etc.
[0042] It should be noted that the facial expressions and gestures of the driver 102 can already be recorded for other purposes, e.g., for vehicle control. Inventive embodiments extend the scope and usefulness of corresponding camera systems if they can additionally be used according to the invention to improve voice communication from the vehicle. In step 208, the video signal 132 is used by the AI module. It is inherent in the nature of artificial intelligence, machine learning (ML), ML models, etc., that it cannot be precisely determined how the AI module 124 recognizes sound sequences, words, sentences, etc., spoken by the vehicle occupant 102 from the input video signal 132.
[0043] However, AI models are known to generate an audio signal from lip movements, representing the recognized sound sequences, words, etc. It is obvious to those skilled in the art that such recognition results can be based solely on lip movements, but can also include other aspects that contribute to or accompany speech production, such as tongue movements, mouth movements, head movements, etc. Speech utterances are often accompanied by small, possibly involuntary, head movements, as well as movements of the shoulders, upper body, etc. Utterances can also be accompanied by gestures of the fingers, hand, etc. Insofar as these circumstances are captured by the camera system, they can be used by the AI module for the most reliable sound recognition possible.
[0044] The AI module can implement a machine learning model such as a speech model, LLM, LAM, a multimodal model, etc. In one embodiment, the AI module 124 uses the audio signal 130 in addition to the video signal 132 for recognition. In yet another embodiment, the AI module additionally or alternatively uses the speech component 114 from the other party of the conversation 110 to validate the speech component of the vehicle occupant 102.
[0045] In some embodiments, the KL module 124 includes or is part of an intelligent personal assistant (IPA) for the vehicle occupant 102. The IPA can be provided, for example, by the vehicle, a vehicle manufacturer, a head unit, etc. The IPA can be provided to a driver, vehicle owner, vehicle user, such as the vehicle owner's employees, etc. In other words, the IPA can be provided to one or more drivers, co-drivers, and / or other passengers or occupants of the vehicle 100.
[0046] The IPA can, for example, be individually trained to reliably recognize sound sequences of the vehicle occupant 102. Such training can, for example, be based on feedback related to the vehicle's voice control 100. Such training can, for example, be based on recognized voice commands for the vehicle's voice control, where the recognition result 134 was correctly recognized, needs to be corrected, etc. Training can also be based on feedback within the framework of exemplary embodiments of the present invention. Such training can, for example, be based on audio data, video data, or a combination thereof.
[0047] In step 210, the processing unit 126 processes the speech signal stream 130 to improve or optimize the quality of the outgoing speech signal 110, using the recognition result or audio signal 134 of the AI module 124. For example, the processing unit 126 can subtract an interference signal from the speech signal 130, so that, in an ideal case, only a pure, i.e., undisturbed, portion of the conversation remains.
[0048] In one embodiment, such an interference signal is generated as a difference between the speech signal stream 130 at input 120 and a plausible (pure, undisturbed) speech signal that can be represented by the audio signal 134. This plausible speech signal 134 can include the sound sequences recognized by the KL module 124. Due to the design of the interference signal, these components are retained in the processed speech signal stream, at least when they are present therein. Furthermore, the control device 108 or processing unit 126 can be configured to add missing speech components to the outgoing speech signal 110 according to the plausible speech signal 134. Such missing speech components may, for example, not have been detected at all or not completely by the microphone system 102 because the speaker 102 was turned away from the microphone, or because the speech component was, for example,filtered out due to an unfavorable signal-to-noise ratio in upstream processing, etc.
[0049] In a further embodiment, the plausible speech signal 134 can also contain speech components that the KL module 124 adds to the recognized sound sequences, which were not actually spoken by the speaker 102, e.g., because the speaker 120 stumbles due to high volume of local background noise, because their attention is focused on an event in the vicinity of the vehicle, etc. These speech components cannot be contained in the recorded speech signal stream 130, regardless of whether further missing speech components were not captured by the microphone or were removed by preprocessing.
[0050] According to the nature of artificial intelligence, in some embodiments a precise distinction between noisy or uncaptured speech components on the one hand, and words not articulated or not fully articulated by the speaker on the other, is not necessary: In these embodiments, the operation of the AI module and processing unit results in syllables that were "swallowed" in some way being fully restored or completed.
[0051] In one embodiment, the processing unit 126 can remove portions of the speech signal stream 130 that represent sounds made by the vehicle occupant 102, but which do not typically belong to the conversational content, such as eating or chewing sounds, certain involuntary exclamations, etc. A simple approach would be for the AI module 124 to correctly recognize sounds originating from, for example, mouth movements of person 102, but to classify or validate these sounds as not belonging to the conversation and therefore remove or exclude them from the plausible speech signal 134. The processing unit 126 would then automatically classify these sounds as interference and filter them out of the speech signal stream 130.
[0052] In a further embodiment, the processing device 126 can adjust the volume of the resulting speech component 110, e.g. reduce it, for example because after removing background noise the speech component 110 of the speaker 102 would sound unnaturally loud to the other party.
[0053] Operation 212 terminates procedure 200. For example, disconnecting the call (110) can deactivate control device 108.
[0054] It is evident to the person skilled in the art that the process 200 comprises the fact that, during the existence of the conversation connection 110, in particular steps 208 and 210 are carried out continuously and in parallel, i.e., the processing device 126 continuously processes the speech signal stream 130 in order to improve its quality, i.e., to generate an optimized conversation component 112, and for this purpose the AI module 124 provides continuous recognition in the form of the plausible speech signal 134 based on the video signal stream 132.
[0055] In certain embodiments, Method 200 is used to improve call quality in purely audio-only calls made from within a vehicle. In some embodiments, Method 200 can also be used for video calls, video conferences, or AA / conversations. Such calls are becoming increasingly important with the growing prevalence of automated and autonomous driving. For example, the ability to conduct high-quality video calls or video conferences may become relevant from around SAEf (Society of Automotive Engineers) Level 3 onwards. Inventive methods and control devices in motor vehicles can make a significant contribution to ensuring good call and connection quality in such scenarios.
[0056] Reference sign Motor vehicle Vehicle occupant Microphone system Camera system Control device Call connection Outgoing part of the call Incoming part of the call Antenna system Audio input Video input KL module Processing unit Speech signal stream Video signal stream Detected sound sequence Procedure Activating the control device 108 Receiving and providing the speech signal stream Receiving and providing the video signal stream Recognizing sound sequences Processing the speech signal stream Deactivating the control device 108
Claims
Claims 1. Method (200) for controlling a telephone connection (110) from a motor vehicle (100), comprising the following steps: - Recording (204) a speech signal stream (130) which includes a conversation share (112) of a speaking vehicle occupant (102) in the conversation link (110); Recording (206) a video signal stream (132) that includes a facial component of the speaking vehicle occupant (102); - Recognizing (208) sequences of sounds spoken by the vehicle occupant (102) based on the mouth movements of the vehicle occupant (102); and - Processing (210) the speech signal stream (130) based on the detected sound sequences to improve the outgoing speech signal quality of the conversation connection (110).
2. Method (200) according to claim 1, wherein the processing of the speech signal stream (130) comprises subtracting an interference signal.
3. Method (200) according to claim 2, wherein the recognition of sound sequences comprises providing a plausible speech signal (134), and the interference signal is a difference between the recorded speech signal stream (130) and the plausible speech signal (134).
4. Method (200) according to one of the preceding claims, wherein the processing of the speech signal stream (130) comprises adding a missing speech component.
5. Method (200) according to claim 4, wherein the processing (210) of the speech signal stream (130) comprises adding a sound not spoken by the vehicle occupant (102).
6. Method (200) according to one of the preceding claims, wherein, in addition to the mouth movements of the vehicle occupant (102), facial expressions of the vehicle occupant (102) are used to recognize (208) the sequences of sounds.
7. Method (200) according to one of the preceding claims, wherein, in order to recognize (208) the sound sequences, additional speech components (114) of the other side of the conversation connection are used.
8. Method (200) according to one of the preceding claims, wherein the conversation link (110) is an audio / video link, and the video signal stream (132) comprises a video portion of the link (110).
9. Control device (108) for controlling a telephone connection from a motor vehicle, comprising: - one audio input (120), - a video input (122), - a KL module (124), and - a processing facility (126); wherein - the audio input (120) is configured to provide a speech signal stream (130) received from a microphone system (104) of the motor vehicle (100), which includes a speech component (112) of a speaking vehicle occupant (102) in the conversation link (110); the video input (122) is configured to provide a video signal stream (132) received from a camera system (106) of the vehicle (100), which includes a facial component of the speaking vehicle occupant (102); - the KL module (124) is designed to use the video signal stream (132) to recognize spoken sequences of sounds based on the mouth movements of the vehicle occupant (102); and - the processing unit (126) is designed to process the speech signal stream (130) based on the detected sound sequences in order to improve the outgoing speech signal quality of the conversation connection (110).
10. Control device (108) according to claim 9, wherein the AI module (124) comprises a language model.
11. Control device (108) according to claim 9 or 10, wherein the KL module (124) implements an intelligent personal assistant for the talking vehicle occupant (102).
12. Motor vehicle, including: - a microphone system (104), - a camera system (106), and - a control device (108) according to one of claims 9-11 .