Voice control of a motor vehicle

By determining speaker characteristics and integrating various data sources, the method enhances voice control systems in vehicles to interpret speech inputs accurately and adaptively, addressing language limitations and enabling personalized interaction.

DE102016217026B4Active Publication Date: 2026-03-26BAYERISCHE MOTOREN WERKE AG
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2016-09-07
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Modern voice control systems in vehicles are limited by their reliance on specific language inputs, failing to recognize and process voice commands from users speaking different languages, and lack the capability for a more natural and versatile interaction.

Method used

The method involves determining the speaker's characteristics, such as seating position, identity, role, and personal attributes, to process voice inputs intelligently and adaptively, using multiple microphones, biometric data, and non-verbal cues to enhance speech recognition and localization, and integrating data from mobile devices and vehicle systems.

Benefits of technology

Enables more natural and versatile voice control by accurately interpreting speech inputs based on the speaker's context, allowing for enhanced functionality and personalized interaction with the vehicle systems.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Procedure for voice control of a motor vehicle with the following steps - Capturing a voice input from an occupant of the vehicle, - Determining a property of the speaker, wherein the property of the speaker includes a seating position of the speaker in the motor vehicle, - Processing speech input depending on the speaker's characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to a method for voice control of a motor vehicle and to a motor vehicle with a voice control system.

[0002] Modern vehicles already offer extensive voice control capabilities. Integrated microphones allow for the capture and processing of voice commands from occupants, particularly the driver. This makes it possible to operate numerous functions via voice commands, such as entering navigation destinations, tuning radio stations, making phone calls, etc.

[0003] Modern voice control systems can be used without prior training on the user's voice. However, their use is subject to certain limitations. For example, voice control systems typically expect input in a specific language. If the user speaks a different language, the voice input usually cannot be correctly recognized and processed.

[0004] The increased use of personal assistance systems, such as smartphones, is leading users to become accustomed to an increasingly "natural" way of interacting with electronic assistance systems, one that corresponds to human interaction. This growing expectation also extends to motor vehicles, whose assistance systems, and especially their human-machine communication capabilities, are facing increasing demands. This is all the more true given the introduction of partially, highly, and fully automated vehicles.

[0005] US 2014 / 0074480 A1 describes a vehicle with a plurality of microphones, each arranged in a predefined spatial area, whereby an occupant of an area is recognized by their voice.

[0006] DE 10 2012 205 336 A1 discloses a system for the real-time detection of an emergency situation occurring in a vehicle, wherein a sensor is configured to detect an occupant-related condition.

[0007] From DE 10 2015 121 955 A1, a monitoring of a rear passenger seating area of ​​a vehicle is known.

[0008] DE 10 2009 051 508 A1 concerns a method for activating and / or guiding speech dialogues with a speaker recognition unit.

[0009] From DE 10 2013 003 059 A1 a gaze-direction-dependent voice control of a functional unit of a motor vehicle is known.

[0010] Against the background described above, the task is to improve voice control in motor vehicles and, in particular, to enable a more natural use of such voice control.

[0011] The problem is solved by a method and a motor vehicle with the features of the independent claims. Advantageous embodiments of the invention are the subject of the dependent claims.

[0012] In the inventive method for voice control of a motor vehicle, a voice input from an occupant of the motor vehicle is first recorded. The occupant can be, in particular, the driver of the motor vehicle. However, it can also be any other occupant, for example, a front-seat passenger or a passenger in a rear seat. Subsequently, at least one characteristic of the speaker is determined. The speaker is the occupant who uttered the previously recorded voice input.

[0013] According to the invention, the speech input is processed depending on the speaker's characteristics. In other words, the invention provides that the speech input is no longer processed solely as such—as is customary in the prior art—but rather that the speaker's characteristics are taken into account during processing. The core of the invention can thus be described as assigning the speech input to the speaker for the purpose of processing. This assignment enables a wide range of ways to make speech input more intelligent and versatile, thereby significantly expanding the possibilities of voice control in motor vehicles.

[0014] The invention provides that the speaker's attribute includes their seating position in the motor vehicle. In other words, the speaker's seating position is determined, and the speech input is then processed based on this position. This makes it possible to process speech inputs whose semantics can only be understood by considering the surrounding information of the seating position. For example, the speech input "Open my window" is difficult to process without knowledge of the seating position, as the meaning of the possessive pronoun "my" would be impossible to determine. However, by processing the speech input based on the speaker's seating position, it is possible to determine which window should be opened.

[0015] Determining the speaker's seating position can advantageously include localizing the sound source of the speech input. In other words, the speaker's seating position can be determined by calculating the spatial origin of the acoustic signal. For this purpose, at least two microphones must be installed in the vehicle, but preferably more than two. It is particularly preferred that one microphone be positioned in close proximity to each seat (e.g., in the vehicle roof above each seat). This enables, for example, the localization of the sound source using level measurement, triangulation, and / or other known methods.

[0016] Particularly advantageous in this context are source separation methods known from signal processing, such as independent component analysis (ICA) or principal component analysis (PCA). Besides enabling the localization of the sound source, these methods also offer the advantage of separating multiple simultaneously uttered speech inputs. This allows for parallel processing of each of the simultaneously uttered, but source-separated, speech inputs by executing the inventive method multiple times in parallel (i.e., for each speech input) after the source separation has been performed.

[0017] Alternatively or additionally, the speaker's seating position can be determined by detecting the occupant's speech activity. This is preferably done by evaluating images captured by an interior camera of the vehicle in such a way as to detect speech activity. Speech activity can be detected, in particular, by recognizing a movement of the occupant's mouth typical of speaking.

[0018] Alternatively or additionally, the speaker's seating position can be determined by the speaker operating a designated control element. It is already known to install a button on the steering wheel in motor vehicles, which the driver must press to input a voice command. It is particularly advantageous to provide several, preferably all, seats in the vehicle with such a control element. Each occupant can then operate their control element to indicate that the voice input (following the activation of the control element) originates from them. Furthermore, it can be provided that voice inputs are only considered (i.e., processed) if one of the control elements has been pressed beforehand. However, it can also be provided that every voice input is processed regardless of whether one of the control elements has been activated.In the latter case, the control element merely serves to assist the vehicle's voice control system in determining the speaker's seating position.

[0019] In a further embodiment of the method according to the invention, the speaker's characteristic includes the speaker's identity. In other words, the speaker's identity is then determined. The term "identity" is to be understood broadly in this context. It should encompass all characteristics that can be suitable for distinguishing the speaker from other persons (identity feature). For example, the identity can - the name and / or - an identification number and / or - include a biometric characteristic, in particular facial features, posture, body measurements, weight, fingerprint, iris appearance, voice characteristic of the speaker.

[0020] By determining the speaker's identity, the processing of speech input can be designed in a particularly diverse way, because then any information relating to the individual speaker can be used for processing.

[0021] The speaker's identity is determined with particular advantage by capturing at least one identifying characteristic. The speaker's identity is then established by comparing this captured characteristic with identifying characteristics stored in a database. Such a database can be a data storage device located in the vehicle, in which identifying characteristics of previously registered individuals are stored. However, such a database can also be located outside the vehicle. In this case, access can be achieved, for example, via a mobile data communication device in the vehicle.

[0022] Preferably, the acoustic signal representing the captured speech input can be analyzed with regard to identity characteristics relating to the speaker's voice. By comparing this signal with the database, it can then be determined that the speaker's voice is, for example, the voice of Mr. X or Ms. Y, i.e., that the speaker is Mr. X or the speaker is Ms. Y.

[0023] Furthermore, preferably, visual biometric data (e.g. facial features, iris, posture, etc.) of the speaker can be captured using an indoor camera and then assigned to a person (Mr. X, Mrs. Y, etc.) by comparison with the database.

[0024] Preferably, other biometric data of the speaker can be captured using a suitable data acquisition device and then matched against a database to identify a person (Mr. X, Ms. Y, etc.). For example, the speaker's weight can be captured using a weight measurement device located in the seat, or a fingerprint can be captured using a fingerprint recognition device (e.g., on the steering wheel and / or other controls).

[0025] Preferably, a speaker's identifying characteristic can be captured through self-identification. In other words, the speaker can explicitly communicate their identity to the voice control system. This can be done in various ways. For example, a fingerprint reader can be used, which each occupant uses once upon entering the vehicle, thus revealing their identity to the vehicle.

[0026] Furthermore, mobile devices carried by the occupants may be used that allow conclusions to be drawn about the identity of the respective owner. Examples of such mobile devices include smartphones, smartwatches, and personally assigned electronic vehicle keys.

[0027] Such mobile devices typically have unique identifiers and can be assigned to a user, thus enabling the determination of the user's identity. For example, if the presence of the mobile devices SMARTPHONE_ABC belonging to Mr. X and SMARTWATCH_XYZ belonging to Ms. Y is detected in the vehicle, then their users are presumably also occupants.

[0028] Particularly preferred is the determination of the position of such assigned mobile devices, and thus the seating position of the respective user of the mobile device in the motor vehicle, e.g. by means of - a charging port or docking station connected to the mobile device and / or - the signal strength of a mobile network signal (e.g., Wi-Fi, Bluetooth, NFC, RFID, LTE, etc.), whereby triangulation across several different signal types (e.g., Bluetooth and Wi-Fi) is also possible. In particular, changes in signal strength can be profitably analyzed. For example, it can be determined at the time a user enters the vehicle through which door or from which side of the vehicle the mobile device entered.

[0029] Furthermore, self-identification can be performed via voice input ("I am Mr. X"). It is particularly advantageous to include, as additional procedural steps, the determination of the speaker's identity, a corresponding request, especially an audible one, to the speaker ("Please state your name"), and subsequently the recording and evaluation of the speaker's voice input.

[0030] Furthermore, preferably, an identifying characteristic of the speaker can be determined through statistical analysis. For example, if it is known from numerous self-identifications of the driver in the past using the aforementioned fingerprint reader that the vehicle is driven by Mr. X with a very high relative frequency (e.g., over 95 percent), then even if the driver is the speaker and the speaker does not identify themselves using the fingerprint reader, the name "Mr. X" can be determined as the identifying characteristic of the speaker.

[0031] In a further embodiment of the method according to the invention, the speaker's characteristic comprises a role of the speaker relating to the operation of the motor vehicle and / or a personal characteristic of the speaker.

[0032] The role of the speaker relating to the operation of the motor vehicle is understood to be a characteristic of the speaker that he assumes in relation to the operation of the motor vehicle.

[0033] One such role is that of the driver of the motor vehicle. The driver can be, in particular, the passenger sitting in the driver's seat. However, it should be noted that this is not necessarily the case. In vehicles with fully automated driving capabilities, the responsible driver can also be seated in any other position. Their role as driver is defined by their authority to make decisions and issue commands regarding the vehicle's operation. Even in vehicles that are not fully automated, the driver may not be in the driver's seat. This is the case, for example, with driving school vehicles, where the responsible driver (namely, the driving instructor) typically sits in the passenger seat.

[0034] Another role is that of the front passenger. This person usually sits in the front of the vehicle next to the driver (i.e., in the front passenger seat). However, exceptions to this are possible. A further role is that of the passenger, which can refer in particular to occupants in the back seat of a car or passengers in the passenger seats of a bus, etc.

[0035] A personal characteristic of the speaker is understood to be a characteristic that the speaker possesses as a person, independent of the vehicle's operation. Examples of personal characteristics include "adult," "child," "man," "woman," and language skills (e.g., German-speaking, able to read Chinese characters, etc.).

[0036] The advantage of knowing the speaker's role and / or personal characteristics is that the speech input can be processed much more effectively with this additional information. For example, a speech input can have a fundamentally different meaning depending on the speaker's role and characteristics. The following example illustrates this. In a highly automated vehicle, there is an adult driver (in the driver's seat) and a child in the back seat. The child is playing a racing game. The speech input "Turn left" obviously has a completely different meaning depending on the speaker. If the driver speaks (role: driver, personal characteristics: adult), the processing of the speech input can determine that it is a control command for the vehicle.If, however, the child speaks (role: passenger, personal characteristic: child), it can be determined during the processing of the speech input that a control command for the video game was intended.

[0037] It is preferable to have the speaker's role and / or personal characteristics recognized. This can be done automatically using previously recorded biometric data. For example, voice recognition can be performed. This allows for the recording and evaluation of voice characteristics that indicate a child's or adult's voice (pitch, clarity). Other biometric parameters, which particularly allow for differentiation between adults and children, include height, weight, and many others.

[0038] Furthermore, determining the speaker's role and / or personal attribute through self-identification is conceivable. As already mentioned, a fingerprint reader, an authentication device (smartphone, smartwatch, vehicle key, etc.), or even voice input ("I am the driver," "I am of legal age") can be used for this purpose. The vehicle's voice control system can be configured to detect that the speaker's role and / or personal attribute needs to be determined and then prompt the speaker (particularly via voice output) to provide the corresponding information (via voice input or other means). The following example dialogue serves as an illustration: Driver: "Navigate to address A" - Voice control system: "Are you the driver?" or "Are you of legal age?" or "Please identify yourself as the driver using the fingerprint reader."

[0039] Furthermore, it is possible to determine the role and / or personal characteristics of the speaker by evaluating statistical information. As described above, for example, based on previous uses of the vehicle, it can be determined with a certain probability that the driver of the vehicle is sitting in the driver's seat.

[0040] It can be particularly advantageous to prioritize multiple voice inputs based on the characteristics of the speakers, especially their roles. For example, a voice input from the driver can be processed while a simultaneous voice input from a child can be ignored.

[0041] In a further embodiment of the invention, the speaker's audio signal is filtered out and speech signals from other occupants are suppressed using signal processing methods known per se. This allows for better detection of speech input. This is particularly advantageous in multi-stage speech input ("dialogue"). For example, knowledge of the speaker's seat location can be used to filter out an audio signal originating from that seat. Furthermore, a known voice profile can be used to adapt an adaptive filter accordingly.

[0042] In a further embodiment of the method according to the invention, the following steps can be carried out: - Capturing at least one nonverbal utterance of the speaker, in particular a direction of gaze and / or a gesture of the speaker, - Processing speech input depending on the speaker's nonverbal utterance.

[0043] The advantage of this embodiment is that even more background information can be used to process the speech input. This makes it possible to interpret the meaning of the speech input against the background of the speaker's captured nonverbal utterance.

[0044] It is particularly advantageous if the procedure includes the following steps: - Capturing a voice input from an occupant of the vehicle, - Determining a characteristic of the speaker, in particular the speaker's seating position, - Capturing at least one nonverbal utterance of the speaker, in particular a direction of gaze and / or a gesture of the speaker, depending on the characteristic of the speaker, in particular the speaker's seating position, and - Processing speech input depending on the speaker's nonverbal utterance.

[0045] In other words, for example, knowledge of the seating position can be used to analyze an interior camera pointed at that position and thus capture the speaker's nonverbal utterance. Simultaneous nonverbal utterances from occupants other than the speaker could be disregarded due to the knowledge of the speaker's seating position.

[0046] Capturing at least one nonverbal utterance of the speaker enables advantageous use of the invention in the exemplary situations briefly outlined below: - An occupant points to a vehicle control with a finger and says "Activate function" or "Operate control"; - An occupant points a finger at a specific interior light and says "Turn on this light" or points a finger at a specific window and says "Roll down this window"; - An occupant points to an object outside the vehicle (e.g., a restaurant visible through the window in the distance) and asks, "How long is this restaurant open today?" - An occupant, who is the driver, points to a vehicle ahead and says: "Please follow this vehicle".

[0047] The last two examples refer to objects outside the vehicle. It should be noted that the disclosure is expressly intended to include the case where such objects are virtual, i.e., projected onto the vehicle windshield by means of a contact-analog head-up display.

[0048] In a further embodiment of the method according to the invention, the following steps can be carried out: - Retrieving data assigned to the speaker, - Processing speech input based on data assigned to the speaker.

[0049] Data associated with the speaker can be any data that is (even if not exclusively) associated with them. In particular, this includes data stored on a mobile device (e.g., smartphone, smartwatch) associated with the speaker, to which the vehicle can establish a data connection, for example, via Bluetooth or Wi-Fi. Furthermore, this includes data on an external server, such as an internet server, which the vehicle can access, for example, via a Wi-Fi or mobile network connection. This also includes data stored within the vehicle.

[0050] This advanced technology offers the distinct advantage of enabling the speaker to reference data within their voice input. For example, if the voice input is "Take me home," the speaker's home address can be retrieved from their smartphone, allowing the voice input to be processed successfully. Similarly, the voice input "Call my wife" can be processed in the same way. Furthermore, it is conceivable that several occupants have the contact details of different people with the same name (for example, "Mr. Müller") stored on their respective mobile devices. The invention makes it possible to successfully process the voice input "Call Mr. Müller" by dialing the phone number of the Mr. Müller whose contact details are stored on the speaker's mobile device.

[0051] In a further embodiment of the method according to the invention, the following steps can be carried out: - Determining the speaker's access authorization, in particular access authorization to a driving function of the motor vehicle, - Processing speech input depending on the speaker's access rights.

[0052] In other words, voice input is processed taking access permissions into account, which means that a voice command is either executed fully, with restrictions, or not at all. Determining access permissions can be particularly advantageous if it is based on the speaker's identity, role, and / or personal characteristics.

[0053] For example, a driver can be granted more and / or different access rights than a passenger. Similarly, an adult can be granted more and / or different access rights than a child. For instance, access rights can be configured so that a child is not allowed to control vehicle functions via voice commands, but is permitted to control certain sub-functions of a rear-seat entertainment system.

[0054] Simplified handling, particularly automated assignment, of access rights is conceivable. For example, the occupant in the driver's seat could be automatically granted access to vehicle controls, while passengers in the back seat would only receive voice-activated access to specific entertainment functions. Such automated rights assignment based on the speaker's seating position could be combined with identity-based rights assignment. For instance, it could be stipulated that Mr. X may always execute all voice commands, regardless of his seating position. Conversely, unknown occupants (i.e., those whose identity cannot be established or is unknown) would only be allowed to control entertainment functionalities, regardless of their seating position, even if it is the driver's seat.

[0055] It is particularly advantageous to determine access rights based on the speaker's previously defined attribute and / or the chosen method for determining that attribute. In other words, the speaker can be granted broader access rights if their identity is known than if only their seating position and / or role has been determined. For example, Mr. X (identity) receives broader access rights than the driver (seating position / role). Alternatively or additionally, the more robust the authentication method used to determine their attribute (e.g., identity), the broader the access rights granted to the speaker. For example, the speaker is granted broader access rights if their identity is verified via a fingerprint reader than if it is verified via voice input ("I am Mr. X").

[0056] In a further embodiment of the method according to the invention, the following steps can be carried out: - Determining a property of another occupant, - Processing speech input depending on the characteristics of the other occupant.

[0057] In other words, the voice input is then processed not only based on the previously determined characteristics of the speaker, but also based on the previously determined characteristics of the other occupant. This allows, with particular advantage, the ability to refer to other people during voice input. By determining, for example, the identities, roles, and / or personal characteristics of all vehicle occupants, the following example voice commands can be processed. The characteristics of the other occupant required for processing and therefore determined are indicated in parentheses in each case. - "Close Mrs. Y's window" (identity and seat of Mrs. Y); - "Start children's film on the screens of A and B" (identities and seats of A and B); - "End the children's video games" (personal characteristics (children) and children's seats in the vehicle); - "Set Mrs. Y's seat heating to level 2" (identity and seat location of Mrs. Y); - "Drive Z home" (identity of Z); Z's address can be obtained, for example, from their mobile device or from data associated with them on a server; - "I want to see the same movie on my screen that A has on his screen" (identity and seat of A).

[0058] A particularly advantageous further development of the invention is achieved by generating and outputting a message, especially a voice message, depending on the characteristics of the speaker and / or a characteristic of another occupant. In other words, the voice control method is further developed into a voice dialogue method. The vehicle is enabled to "respond" using the information acquired according to the invention.

[0059] In this way, specific occupants can be addressed based on their characteristics, particularly their role. For example, when important decisions need to be made regarding the vehicle's operation (e.g., whether to bypass a traffic jam or how to react to a detected technical problem, such as stopping or proceeding slowly), an occupant with a suitable role and / or personal characteristic (e.g., "driver," "owner," "proprietor," adult) could be specifically addressed. Including the occupant's name in the address is particularly advantageous. Furthermore, only voice input from the addressed person is accepted as feedback, i.e., as a result of the voice message. An extension is conceivable such that the addressed user does not need to be physically present in the vehicle but can be contacted via telephone or computer device.

[0060] The aspect described above will be illustrated by two examples: - The fuel supply of the (fully automated) vehicle is running low; the occupant in the role of "driver" is specifically asked whether the next gas station should be visited; - The (fully automated) vehicle picks up children from kindergarten; there is a long traffic jam on the way; the vehicle asks the parents (role: parents or vehicle owner) by telephone whether it should bypass the traffic jam.

[0061] It may be intended that the voice message is output via an audio playback device if the occupant being addressed is currently using that device. For example, if an occupant has headphones connected to their mobile device, it is advisable to output the voice message via the mobile device and thus via the headphones, as the addressed occupant might miss an output via the vehicle's speakers.

[0062] It can also be advantageous to focus the voice message output on the seat of the addressed occupant. For example, the output can be made via multiple vehicle speakers in such a way that the output at the addressed occupant's seat has maximum volume. This design is particularly advantageous when several voice messages are output at least partially simultaneously, because the output messages then overlap as little as possible, thus ensuring that each occupant can hear the voice message addressed to them in the best possible way.

[0063] A further advantage is that the message can be or include visual feedback. A particular advantage is that the message can include both a voice message and visual feedback.

[0064] In voice control systems of existing motor vehicles, a so-called "accompanying visual feedback," for example on a vehicle display, is already often provided alongside acoustic feedback to the voice command. For instance, the phone book is displayed on a vehicle display when the command "phonebook" is given; when the weather is requested, an overview of the weather for the next few days is shown. The aforementioned embodiment of the invention improves upon this prior art as follows: By knowing which occupant made the voice command, the appropriate display (in the case of multiple displays in the vehicle) can be selected for the visual notification to that occupant. For example, if a person in the back seat made the voice command, the weather report can only be displayed on the screen assigned to that speaker.This “suitable display” can be a display of the vehicle, but also a display of a mobile device belonging to the occupant to be addressed (e.g. a display of a smartphone that the occupant is holding).

[0065] In a further advantageous embodiment of the invention, the following steps are carried out: - Retrieving at least one previous voice input from the speaker and / or one of the other occupants, - Processing speech input depending on previous speech input.

[0066] In other words, the voice input is processed depending on the speaker's characteristics and on one or more previous voice inputs. This further expands the possibilities of voice control.

[0067] With regard to immediately preceding voice inputs, it is thus possible to conduct a kind of dialogue and to speak in a shortened form in the context of preceding voice inputs (e.g.: "Call Mr. Müller" - "You have three people with the last name Müller in your telephone book" - "Max").

[0068] It is particularly advantageous to stipulate that recently received voice inputs are only considered within a predetermined maximum time period. If more than this predetermined time period has elapsed since a voice input, that input will no longer be considered when processing subsequent voice inputs.

[0069] This method offers a particular advantage: it allows the voice inputs of different occupants to be correlated. To achieve this, previous voice inputs from other occupants are taken into account when processing the current (i.e., most recently uttered) voice input. This enables occupants to refer to and react to voice inputs. The following situation serves as an example: - Inmate 1: "Who did I call yesterday?"; Response from the voice control system: "Max Müller"; - Inmate 2: "Navigate to the address of this Max Müller"; - Inmate 3: "Is there a fast-food restaurant on the way to Max Müller's?"

[0070] In another example, the following sequence of events occurs: - Child A: "Roll down my window"; - Passenger Y (parent): "Open the window again and then lock the window opening function."

[0071] Storing, analyzing, and considering past voice inputs opens up a range of advantageous possibilities. For example, it's conceivable to adapt the speech recognition to different users to improve its quality (e.g., recognition rate), taking into account factors such as voice, word choice, typically used voice commands, and typical acoustic characteristics (seating position, height, distance to the microphone, etc.). Personal preferences discovered in this way can also be used to improve speech recognition (e.g., a child always wants to watch a specific movie on the rear-seat entertainment system; a driver mostly uses voice control for navigation but almost never for phone calls).

[0072] Furthermore, each user's level of knowledge can be evaluated and taken into account. In other words, the voice control system can behave differently depending on whether it is interacting with a more experienced or less experienced user. For example, a child could be addressed differently than an adult. Similarly, a new vehicle user who is not yet familiar with the vehicle's functions and interaction possibilities could be addressed differently than an experienced user. For example, older users could be addressed differently (e.g., more slowly, clearly, and loudly) than younger users. For users with visual impairments, additional visual output could be provided.

[0073] Furthermore, the voice control system could use different voices when communicating with different users, for example, a funny, child-friendly voice for communicating with a child in the back seat and a more "serious" voice for communicating with the driver. This also makes it easier to distinguish between potentially overlapping voice messages. Additionally, the voice control system can adapt the language (e.g., German, English, etc.) to the individual user. For this purpose, it can use, for example, the language in which the speaker made their voice input or a known preferred language of the speaker. This preferred language can be determined by analyzing past voice inputs or by other means, such as through the settings on the user's mobile device.

[0074] By using past voice inputs, it is also possible to differentiate between parallel "conversations" of different occupants. For example, while the driver in the front seat is using voice input to specify the navigation destination, a child in the back seat is simultaneously using voice input to select a film on the rear-seat entertainment system.

[0075] It goes without saying that saving the voice input is possible. It is particularly advantageous to save every voice input (or at least some information derived from it), as this allows the processing of subsequent voice inputs to draw on the largest possible amount of background information (generated by previous voice inputs).

[0076] It may be possible to store voice input not only separately for each speaker, but also with additional attributes. For example, the time of the voice input, its language, and much more can be saved. Furthermore, content-related attributes can also be stored. In the example above, the driver's spoken navigation destination could be assigned the attribute "Navigation," whereas the film title mentioned by the child would be assigned the attribute "Film selection in the infotainment system."

[0077] When issuing multiple messages, especially voice messages, it can be advantageous to stagger their timing, i.e., to issue them with as little overlap as possible or even completely separately. Additionally, the voice control system can address different users by name or other identifying characteristic, so it is clear who is being addressed.

[0078] Previous voice commands can be stored in the vehicle's internal data storage or in external data storage devices. Data storage devices assigned to individual users can also be used to save their respective voice commands, such as mobile devices or personally assigned vehicle keys. This allows for the use of previous voice commands in different vehicles, which is particularly advantageous for car sharing.

[0079] By saving voice input and entire “conversation histories” for each user, detailed information can be collected (e.g., which user executed which voice commands and received which result, when sitting in which position, and which destination was targeted at what time?).

[0080] Using voice inputs from further back in time allows for a more natural and more complex communication with the vehicle (e.g.: "Turn my seat massage function back on as it was yesterday when I drove to Nuremberg (with another car-sharing vehicle)"; "Order two more cinema tickets next to the two seats from yesterday").

Claims

[1] Method for voice control of a motor vehicle comprising the steps - Capturing a voice input from an occupant of the vehicle, - Determining a property of the speaker, wherein the property of the speaker includes a seating position of the speaker in the motor vehicle, - Processing speech input depending on the speaker's characteristics. [2] Method according to the preceding claim, wherein the speaker property comprises an identity of the speaker. [3] Method according to any of the preceding claims, wherein the speaker's characteristic comprises a role of the speaker relating to the operation of the motor vehicle and / or a personal characteristic of the speaker. [4] Method according to any of the preceding claims comprising the steps - Capturing at least one nonverbal utterance of the speaker, in particular a direction of gaze and / or a gesture of the speaker, - Processing speech input depending on the speaker's nonverbal utterance. [5] Method according to any of the preceding claims comprising the steps - Retrieving data assigned to the speaker, - Processing speech input based on data assigned to the speaker. [6] Method according to any of the preceding claims comprising the steps - Determining the speaker's access authorization, in particular access authorization to a driving function of the motor vehicle, - Processing speech input depending on the speaker's access rights. [7] Method according to any of the preceding claims comprising the steps - Determining a property of another occupant, - Processing speech input depending on the characteristics of the other occupant. [8] Method according to any of the preceding claims comprising the steps - Generating a message, in particular a speech message, depending on the characteristic of the speaker and / or a characteristic of another occupant, - Issuing the message. [9] Method according to any of the preceding claims comprising the steps - Retrieving at least one previous voice input from the speaker and / or one of the other occupants, - Processing speech input depending on previous speech input. [10] Motor vehicle with a voice control system configured to carry out the method according to any of the preceding claims.

Citation Information

Patent Citations

  • Device, system and method for activating and guiding speech dialogue

    DE102009051508A1

  • System and procedure for the real-time detection of an emergency situation occurring in a vehicle

    DE102012205336A1

  • Method for controlling functional unit of motor vehicle, involves automatically detecting whether view of vehicle occupant to given area is directed by view detection unit

    DE102013003059A1

  • Method and device for monitoring a rear passenger seating area of ​​a vehicle

    DE102015121955A1

  • Voice stamp-driven in-vehicle functions

    US20140074480A1