Method for operating a vehicle-integrated voice assistant, voice assistant and vehicle

The method enhances vehicle voice assistants by logging user interactions, assigning timeliness factors, and using machine learning to improve natural interaction and reduce response time.

DE102023003428B4Active Publication Date: 2025-06-18MERCEDES BENZ GROUP AG
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
DE102023003428
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-08-21
Publication Date
2025-06-18
Estimated Expiration
2043-08-21

AI Technical Summary

Technical Problem

Existing voice assistants in vehicles often misinterpret voice commands or require excessive time to execute them, leading to unnatural and inefficient user interactions.

Method used

A method for operating a vehicle-integrated voice assistant that logs user interactions, assigns a timeliness factor to contexts, and selects the most relevant context for executing vehicle functions, using machine learning to enhance natural and intuitive interaction.

Benefits of technology

Enables more natural and efficient voice command execution by accurately interpreting user intentions, reducing response time and enhancing user convenience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000008_0000
    Figure 00000008_0000
  • Figure 00000008_0001
    Figure 00000008_0001
  • Figure 00000008_0002
    Figure 00000008_0002
Patent Text Reader

Abstract

Method for operating a vehicle-integrated voice assistant, wherein the voice assistant detects and processes a voice input made by a user and determines from the voice input a context-dependent operating intention for operating a vehicle functionality, comprising the following method steps carried out by a vehicle-internal computing unit: - (101): Logging usage behavior of the vehicle functionalities, whereby each user interaction with a vehicle functionality is assigned a context; - (102): Matching the context-dependent operating intention with the contexts logged during user interaction with the vehicle functionalities; and - if an agreement between a logged context and the operating intention is determined: (103): operating the vehicle functionality to be controlled by the voice input, taking the context into account; characterized in that the computing unit, when logging the usage behavior of the vehicle functionalities, assigns a timeliness factor to a respective context, wherein the timeliness factor assumes its maximum value upon assignment and then steadily decreases over time, and wherein, in the case of several contexts that match an operating intention, the computing unit selects the context for operating the vehicle functionality whose timeliness factor has the greatest value at that moment.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for operating a vehicle-integrated voice assistant according to the type defined in more detail in the preamble of claim 1, a corresponding voice assistant and a vehicle with such a voice assistant.

[0002] Voice assistants allow functions to be controlled by voice. A user can issue a corresponding voice command that causes a corresponding device to perform actions. The device could be a mobile device such as a smartphone, tablet computer, or laptop. For example, an internet search can be initiated, the weather forecast can be checked, an alarm can be set, a calendar entry can be generated, a text message can be written, a call can be initiated, and so on.

[0003] The use of voice assistants is also well known in vehicles. In this context, voice assistants are particularly advantageous because they allow user interaction with vehicle functions without requiring the user to take their eyes off the road and / or manually operate buttons or a touch-sensitive display to enter commands. This increases user convenience and improves road safety.

[0004] While voice assistants are usually capable of correctly capturing voice input and reliably deriving the intent contained in a voice command, voice assistants are not yet fully developed, meaning that voice commands may not be understood or misinterpreted, or their execution may take an unacceptably long time. In particular, interacting with a voice assistant often feels unnatural because the user must speak to the voice assistant in a specific way, particularly using predefined commands and / or sentence structures.

[0005] Therefore, there is a need to provide methods and means that allow for more natural interaction with voice assistants, preferably while shortening the time required to execute a corresponding voice command. The naturalness of voice interaction is further ensured by the fact that in human communication, the interpretation of data always takes place using the knowledge of the recipient. This unconscious, yet permanent feature of human communication is also to be achieved by a voice assistant, according to the invention.

[0006] A speech recognition system and a corresponding method for its operation are known, for example, from DE 10 2015 211 101 A1. The speech recognition system comprises a mobile unit that is in wireless communication with a server. The mobile unit is able to download expressions required for speech analysis or to generate a speech output from the server. Expressions stored on the mobile unit can be supplemented or replaced. Which expressions are to be integrated into the mobile unit is determined depending on events, such as a concert or a sporting event, and their respective start and end times. In addition, context data can be taken into account to determine the expressions to be implemented in the mobile unit. Expressions that are rarely used or not used at all can be deleted from the mobile unit.While bringing other expressions into the mobile device increases the reliability that the voice assistant running on the mobile device actually understands a user, it does not improve the natural interaction with the voice assistant.

[0007] Furthermore, DE 10 2015 213 722 A1 discloses a method for operating a speech recognition system in a vehicle and a corresponding speech recognition system. The speech recognition system is capable of operating a vehicle function while taking into account the context of the respective situation.

[0008] Furthermore, DE 10 2022 000 387 A1 discloses a method for processing voice inputs and an operating device for controlling vehicle functions. Several vehicle functions that are suitable for execution are displayed on a display device. The execution of a respective vehicle function can be initiated by a voice command, whereby a vehicle function highlighted on the display device can be initiated by a shortcut command.

[0009] Furthermore, DE 10 2016 218 270 B4 discloses a method for operating a motor vehicle operating device with a speech recognizer and a corresponding operating device. Depending on a given driving situation, a command word is predicted for a user based on a behavior model. A check is then performed to determine whether the command word is present in a list of recognition results. One of the recognition results contained in the list can be selected for the speech recognizer, or the recognition results can be re-sorted according to a confidence order.

[0010] Furthermore, WO 2014 / 060 054 A1 discloses speech recognition in a motor vehicle. At least a portion of a speech input recognized in the vehicle is fed to both an in-vehicle speech recognizer and an external speech recognizer. The speech text underlying a speech input is determined by a processing device in agreement with a recognition result provided by the in-vehicle and external speech recognizers.

[0011] The present invention is based on the object of providing an improved method for operating a vehicle-integrated voice assistant, with the aid of which a natural interaction with the voice assistant is enabled, in particular by shortening the time required for the voice assistant to react.

[0012] According to the invention, this object is achieved by a method for operating a vehicle-integrated voice assistant having the features of claim 1. Advantageous embodiments and further developments as well as a correspondingly operable voice assistant and a vehicle with such a voice assistant emerge from the dependent claims.

[0013] A generic method for operating a vehicle-integrated voice assistant, wherein the voice assistant captures and processes a voice input made by a user and determines from the voice input a context-dependent operating intention for operating a vehicle functionality, comprises the following method steps executed by a vehicle-internal processing unit: - Logging the usage behavior of the vehicle functionalities, whereby each user interaction with a vehicle functionality is assigned a context; - Matching the context-dependent operating intention with the contexts logged during user interaction with the vehicle functionalities; and - if an agreement is found between a recorded context and the operating intention: operating the vehicle functionality to be controlled by the voice input, taking the context into account.

[0014] According to the invention, it is provided that the computing unit assigns a timeliness factor to a respective context when logging the usage behavior of the vehicle functionalities, wherein the timeliness factor assumes its maximum value during assignment and then steadily decreases over time, and wherein the computing unit selects the context for operating the vehicle functionality whose timeliness factor has the greatest value at that moment when there are several contexts that match an operating intention.

[0015] The method according to the invention enables particularly intuitive user interaction with the voice assistant. The voice assistant is able to determine which operating input should be made, taking the context into account, so that the user no longer has to fully formulate corresponding voice commands as usual. Instead, individual semantic content is determined by the voice assistant itself, taking the context into account. For example, the user no longer has to say: "Hey Mercedes, call the contact Anna Schmidt from the phone book," but rather: "Call Anna." This is possible if, for example, the user has viewed an address book within the last 10 minutes, to which the computing unit has at least read access. Thus, in this example, the user interaction with the vehicle functionality corresponds to opening and viewing the address book.The computing unit logs this usage behavior and determines the respective context. In this example, the context could be determined to be that a name from the address book is to be provided as an input for a respective vehicle function. The operating intention underlying the voice input is seen as starting a telephone call. Thus, a telephony function is to be used as a vehicle functionality. Telephony functions are generally associated with the fact that contact names can be provided as an input, in addition to, for example, the direct entry of a telephone number. Reading a name from the address book is available as the context. As additional information, the computing unit determines from the voice input that “Anna” should be selected as the name. The computing unit is therefore able to correctly interpret the shortened voice command to initiate the call.This shortened voice command is formulated in natural language, so that the voice interaction between the user and the voice assistant can be designed to be particularly natural and therefore intuitive.

[0016] The vehicle-integrated voice assistant can otherwise be implemented as usual. The necessary hardware components, such as acoustic detection devices such as microphones and corresponding processing components such as the aforementioned processing unit, are present. A signal generated by the microphone can be read by a speech recognition module and converted into a character string, such as a text string. This character string can be read and processed using a natural language recognition module. This enables the recognition of corresponding operating intentions and other information, such as, referring back to the previous example, the name to be called, in the character string.

[0017] The vehicle functionalities that can be recorded can preferably include vehicle functionalities that go beyond the voice assistant, such as operating the vehicle's infotainment system, adjusting the air conditioning, selecting a sports driving mode, selecting a typical speed of travel, i.e. driving behavior in the broader sense, and the like.

[0018] For example, context could also be determined: a location, a time of day, a parameter of a vehicle functionality such as a tuned radio station, a target temperature, a driving mode, etc.

[0019] According to the invention, as already described, the method provides that the computing unit assigns a timeliness factor to a respective context when logging the usage behavior of the vehicle functionalities. The timeliness factor assumes its maximum value upon assignment and then steadily decreases over time. If there are several contexts that match an operating intention, the computing unit selects the context for operating the vehicle functionality whose timeliness factor has the greatest value at that moment. With the help of the timeliness factor, suitable contexts can thus be selected even more accurately, increasing the probability that, even when considering natural language or naturally formulated voice inputs, the respective correct operating intention will be recognized and the respective vehicle functionality will be operated accordingly.

[0020] This approach is based on the idea that the voice inputs made by the user most likely relate to the last vehicle function activated or operated. A further differentiation is made based on the respective type of vehicle functionality. For example, if the user opens the address book and then sets the target temperature of the air conditioning system to a different value, the processing unit is able to relate the respective context to the address book rather than the climate control functionality when a corresponding voice command is given to initiate a call. However, several similar uses of the same vehicle functionality could be made, whereby the most recent user interaction is appropriately taken into account, taking into account the timeliness factor.

[0021] The timeliness factor can take on any value when generated, for example 1. Over time, the timeliness factor decays, for example to a value of 0. The respective expiration time can be chosen differently depending on various boundary conditions, which will be discussed in more detail below.

[0022] According to an advantageous embodiment of the method according to the invention, the computing unit uses a suitably trained machine learning model to compare the context-dependent operating intention with the logged contexts. With the help of artificial intelligence, in particular using machine learning models, a sufficiently trained model can reliably reference corresponding operating intentions with corresponding contexts, thereby increasing the probability that the operating intention underlying a particular voice input will be correctly implemented. The machine learning model can, for example, have been initially trained by the vehicle manufacturer during the development of the voice assistant.

[0023] A further advantageous embodiment of the method according to the invention further provides that the speed at which a respective topicality factor expires is set by a respective topicality factor-specific expiration factor. The expiration factor can, for example, define a ratio by how many points the topicality factor is to be reduced per unit of time, for example by 0.1 points per 10 minutes or the like. Individual expiration factors can then be defined for different vehicle functionalities or different contexts. Operating actions and corresponding vehicle functionalities that are better retained in the user's long-term memory can then be linked to a low expiration factor, so that the respective update factor decreases more slowly and corresponding operating actions orVehicle functions that the user forgets more quickly can be associated with a comparatively higher expiration factor. This increases the reliability that the correct operating intention is recognized and the respective vehicle function is operated appropriately.

[0024] Preferably, the computing unit sets the amount of at least one expiration factor depending on: - the type of vehicle functionality underlying the respective context; - a manual user setting; or - the detection of an interruption of the operation of the vehicle functionality carried out in accordance with the voice input.

[0025] An expiration factor dependent on the respective vehicle functionality can, for example, be hard-coded into a database, e.g., a table, by the vehicle manufacturer. It would also be possible for the user to manually specify a corresponding expiration factor. For this purpose, the vehicle can be equipped with appropriate human-machine interfaces that allow the reception of corresponding user inputs or user specifications. For example, the expiration factor for a specific context of a specific vehicle functionality can be set to the value 6 on a scale ranging from 0 to 10, with the factor 6 corresponding, for example, to a deterioration of 0.6 points of the timeliness factor within 10 minutes. The user can then manually change this value to, for example, 2. This allows the personal preferences of different users to be taken into account.

[0026] In a particularly preferred embodiment, the computing unit is capable of automatically adapting the expiration factor of a respective context to a respective vehicle functionality. To this end, the computing unit monitors user interaction with the voice assistant. If the user aborts the execution of a corresponding voice input shortly after its output, this indicates that the voice assistant misinterpreted the operating intention and / or the context. The cause could therefore be the consideration of a context that is no longer current. By automatically adapting the expiration factor, it is possible to ensure that the corresponding context expires more quickly, so that the most appropriate context can be selected instead.

[0027] A further advantageous embodiment of the method according to the invention further provides that the computing unit deletes a logged context as soon as the timeliness factor assigned to the respective context has expired. Contexts whose timeliness factor has expired are likely no longer relevant for operating vehicle functions. Accordingly, the contexts logged in this way can be deleted, so that less storage space needs to be reserved on a corresponding computer-readable storage medium of the computing unit. The freed-up storage space can then be used for other purposes.

[0028] According to a further advantageous embodiment of the method according to the invention, the computing unit predicts at least one potentially next expected speech input from the currently logged contexts, in particular several potentially next expected speech inputs, before the next speech input occurs, wherein the computing unit assigns to each potentially next expected speech input an input probability determined depending on the respective context. If several potentially next expected speech inputs are determined, the input probability can be used to determine an order which of these potentially expected speech inputs will actually follow. If the computing unit only determines one expected speech input, the input probability does not have to be determined.To determine the input probability, the computing unit can particularly preferentially consider the recency factor. Thus, the input probability increases with a higher value of the respective recency factor, so that voice inputs that relate to recent user interactions with the respective vehicle functionalities are particularly expected.

[0029] Predicting expected speech inputs makes it possible to provide entirely new functionalities. A corresponding advantageous development of the method according to the invention provides that, at least for the most likely next expected speech input, the computing unit, before it is entered, - refers to background information required to formulate a voice response; and / or - carries out at least one intermediate step required to operate the respective vehicle functionality.

[0030] Particularly advantageously, these steps can also be performed early for several voice inputs expected next. This means that, should the processing unit misjudge which voice input will actually be performed, the background information required for the actual voice input is still available or the aforementioned intermediate vehicle functionality steps have already been performed. By obtaining the background information early or performing the respective intermediate steps, the execution time of the respective operation of the vehicle functionality can be shortened. This increases user convenience, as the user has to wait less time. This improves the efficiency of the voice assistant.

[0031] For example, the user can activate route guidance to a destination. The recorded vehicle functionality is then navigation. The destination, for example, can be determined as context. As a voice input, the user can then say, for example, "What will the weather be like?" The processing unit then determines that the user wants to know the weather at the destination at the time of arrival.

[0032] Even before the user makes this voice input, the computing unit can then preferably request the relevant weather information at the destination at the arrival time from a weather service. This is only possible after the user has activated route guidance to the destination. In this case, the timeliness factor or the expiration factor can be selected so that the timeliness factor for the context of destination and destination arrival time will expire at the earliest upon reaching the destination. This enables the correct interpretation of corresponding naturally formulated voice inputs for requesting the weather at any time during the journey.

[0033] Since the information required for this in the form of weather report data was already requested when programming the navigation route, it can be output directly after the voice input has actually been received.

[0034] In a voice assistant comprising acoustic detection means, a computing unit, and a connection for operating a vehicle function, the detection means, the computing unit, and the connection for operating the vehicle function are configured according to the invention to execute a method described above. With the aid of the computing unit, signals generated by the acoustic detection means can be processed, thereby correctly interpreting the respective voice inputs. The computing unit is then able to control the respective vehicle functions for operation via a corresponding interface.

[0035] A vehicle according to the invention has such a voice assistant. This enables particularly convenient and intuitive use of the voice assistant while driving. It can be any road vehicle such as a car, truck, van, bus, or the like. Generally, it could also be a rail vehicle, watercraft, or aircraft.

[0036] Further advantageous embodiments of the method according to the invention for operating a vehicle-integrated voice assistant also emerge from the exemplary embodiments which are described in more detail below with reference to the figures.

[0037] Showing: Fig. 1 shows a flowchart of a method according to the invention for operating a vehicle-integrated voice assistant; Fig. 2 is a flowchart of the intermediate steps performed by a voice assistant according to a first embodiment to respond to a voice input; and Fig. 3 a flowchart of the intermediate steps performed by a voice assistant according to another embodiment to respond to a voice input.

[0038] Interaction with a voice assistant can be uncomfortable for a user if commands must be delivered in a specific format, especially when certain keywords are taken into account. It is therefore desirable to design the interaction with a voice assistant as natural and thus as intuitive as possible for the user. This is where a method according to the invention for operating vehicle-integrated voice assistants comes into play.

[0039] The procedure involves carrying out the Fig. 1. In method step 101, an in-vehicle computing unit logs the usage behavior of vehicle functionalities, whereby each user interaction with a vehicle functionality is assigned a context. In the subsequent method step 102, the computing unit compares the context-dependent operating intention, which was determined from a corresponding voice input, with the contexts logged during the user interaction with the vehicle functionalities. If a match is found between a logged context and the operating intention, in method step 103 the vehicle functionality to be controlled by the voice input is operated taking the context into account.

[0040] Taking the logged context into account, the computing unit is able to determine semantic content for interpreting voice inputs from previous user interactions with vehicle functionalities, so that corresponding voice inputs or voice commands can be formulated more briefly and in simpler language.

[0041] Based on the Fig. 2 and Fig. Figure 3 illustrates a preferred embodiment of the method according to the invention. Thus, the vehicle's processing unit can predict expected future voice inputs, which can be used to obtain background information required to respond to a voice input before the respective voice input is actually input, or to execute intermediate steps required to respond to the voice input in advance. This can shorten the user's waiting time after making the voice input.

[0042] Fig. 2 shows the procedure for responding to a voice command according to a first embodiment of the voice assistant. In method step 201, the vehicle user is proactively asked whether a charging station should be searched for. In method step 202, the user confirms this, for example by saying "yes." This is where the actual waiting time for the user begins. In method step 203, the computing unit transmits a request to a central computing device, such as the cloud server operated by a vehicle manufacturer, to retrieve information about nearby charging stations. In method step 204, the computing unit receives the corresponding results and prepares them. For this purpose, suitable charging stations can be suggested according to known optimization algorithms, and the results can be formatted for easy interpretation. In method step 205, the user then selects a suitable charging station.In method step 206, the computing unit queries the current status of a route guidance. In method step 207, the computing unit makes a decision regarding an action to be performed. In method step 208, the computing unit executes the corresponding action, such as setting the specified charging station as an intermediate destination for the route guidance. This is where the user's waiting time ends.

[0043] In Fig. 3 illustrates the sequence of method steps set out according to a preferred embodiment of the voice assistant. In method step 301, the computing unit already sends the request to the central computing device to obtain information about charging stations. In method step 302, the computing unit prepares the received results. In method step 303, the computing unit queries the current status of the route guidance. In method step 304, the computing unit makes a decision about the actions to be performed. Only in method step 305 is the proactive dialog initiated: "Should I search for a charging station?". In method step 306, the user responds, for example, by saying "Yes." Thus, the waiting time for the user only begins at this point.In method step 307, the computing unit can then directly select a result already prepared from the corresponding background information and perform the corresponding action in method step 308, for example, setting the charging station as an intermediate destination. This ends the user's waiting time.

[0044] According to the Fig. 2, the waiting time for the user lasts from step 202 to step 208, whereas the waiting time according to the method steps shown in Fig. 3 only lasts from step 306 to step 308. This clearly shows the reduction in waiting time for the user.

[0045] The method according to the invention makes it possible to interpret the user's actual intentions more reliably, especially when voice inputs are given through indirect speech or when disambiguating voice commands. Differentiating between information types also provides additional value, since the relevance of contexts also differs for different types of information. Furthermore, the user's waiting time for the voice assistant's response can be shortened.

Claims

[1] Method for operating a vehicle-integrated voice assistant, wherein the voice assistant detects and processes a voice input made by a user and determines from the voice input a context-dependent operating intention for operating a vehicle functionality, comprising the following method steps carried out by a vehicle-internal computing unit: - (101): Logging usage behavior of the vehicle functionalities, whereby each user interaction with a vehicle functionality is assigned a context; - (102): Matching the context-dependent operating intention with the contexts logged during user interaction with the vehicle functionalities; and - if an agreement is found between a recorded context and the operating intention: (103): operating the vehicle functionality to be controlled by the voice input, taking the context into account; characterized bythat the computing unit assigns a timeliness factor to a respective context when logging the usage behavior of the vehicle functionalities, wherein the timeliness factor assumes its maximum value during assignment and then steadily decreases over time, and wherein, in the case of several contexts that match an operating intention, the computing unit selects the context for operating the vehicle functionality whose timeliness factor has the greatest value at that moment. [2] Method according to claim 1, characterized by that the computing unit uses an appropriately trained machine learning model to compare the context-dependent operating intention with the logged contexts. [3] Method according to claim 1 or 2, characterized by that the speed with which a respective timeliness factor expires is set by a respective timeliness factor specific decay factor. [4] Method according to claim 3, characterized by that the calculation unit sets the level of at least one expiration factor depending on: - the type of vehicle functionality underlying the respective context; - a manual user setting; or - the detection of an interruption of the operation of the vehicle functionality carried out in accordance with the voice input. [5] Method according to one of claims 1 to 4, characterized by that the processing unit deletes a logged context as soon as the timeliness factor assigned to the respective context has expired. [6] Method according to one of claims 1 to 5, characterized bythat the computing unit predicts at least one potentially next expected speech input from the currently logged contexts, in particular several potentially next expected speech inputs, before a next speech input occurs, wherein the computing unit assigns to each potentially next expected speech input an input probability determined depending on the respective contexts. [7] Method according to claim 6, characterized by that the processing unit at least for the most likely next expected speech input: - refers to background information required to formulate a voice response; and / or - carries out at least one intermediate step required to operate the respective vehicle functionality. [8] Voice assistant, comprising acoustic detection means, a computing unit and a connection for operating a vehicle functionality, characterized bythat the detection means, the computing unit and the connection for operating the vehicle functionality are set up to carry out a method according to one of claims 1 to 7. [9] Vehicle, characterized by a voice assistant according to claim 8.

Citation Information

Patent Citations

  • Speech recognition system and method for operating a speech recognition system with a mobile unit and an external server

    DE102015211101A1

  • Method for operating a speech recognition system in a vehicle and speech recognition system

    DE102015213722A1

  • Method for operating a motor vehicle operating device with a speech recognizer, operating device and motor vehicle

    DE102016218270B4

  • Method for processing speech input and control device for controlling vehicle functions

    DE102022000387A1

  • Speech recognition in a motor vehicle

    WO2014060054A1