Method for operating a vehicle-integrated voice assistant, voice assistant and vehicle

EP4713916A1Pending Publication Date: 2026-03-25MERCEDES BENZ GROUP AG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2026-03-25

Smart Images

  • Figure EP2024071999_27022025_PF_FP_ABST
    Figure EP2024071999_27022025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for operating a vehicle-integrated voice assistant, wherein the voice assistant captures and processes a voice input made by a user and, from the voice input, determines a context-dependent operating intention for operating a vehicle functionality. The method according to the invention is characterised by the following method steps executed by an in-vehicle computing unit: - (101): logging the usage behaviour relating to the vehicle functionalities, wherein a context is assigned to each user interaction with a vehicle functionality; - (102): comparing the context-dependent operating intention with the contexts logged in relation to the user interaction with the vehicle functionalities; and - if agreement between a logged context and the operating intention is identified: (103): operating the vehicle functionality to be controlled by the voice input taking the context into account.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method for operating a vehicle-integrated voice assistant, voice assistant and vehicle

[0002] The invention relates to a method for operating a vehicle-integrated voice assistant according to the type defined in more detail in the preamble of claim 1, a corresponding voice assistant and a vehicle with such a voice assistant.

[0003] Voice assistants allow functions to be controlled by voice. A user can issue a corresponding voice command that causes a corresponding device to perform actions. The device could be a mobile device such as a smartphone, tablet computer, or laptop. For example, an internet search can be initiated, the weather forecast can be checked, an alarm can be set, a calendar entry can be generated, a text message can be written, a call can be initiated, and so on.

[0004] The use of voice assistants is also well known in vehicles. In this context, voice assistants can be used particularly advantageously, as they allow user interaction with vehicle functions without requiring the user to take their eyes off the road and / or manually operate buttons or a touch-sensitive display to enter commands. This increases user convenience and improves road safety.

[0005] Voice assistants are usually able to correctly capture voice input and reliably derive the intention contained in a voice command.

[0006] However, voice assistants are not yet fully developed, so voice commands may not be understood or misinterpreted, or their execution may take an unacceptably long time. In particular, interacting with a voice assistant often feels unnatural, as the user is required to speak to the voice assistant in a specific way, especially using predefined commands and / or sentence structures.

[0007] Therefore, there is a need to provide methods and means that allow for more natural interaction with voice assistants, preferably while shortening the time required to execute a corresponding voice command. The naturalness of voice interaction is further ensured by the fact that in human communication, the interpretation of data always takes place using the knowledge of the recipient. This unconscious, yet permanent feature of human communication is also to be achieved by a voice assistant, according to the invention.

[0008] A speech recognition system and a corresponding method for its operation are known, for example, from DE 102015211 101 A1. The speech recognition system comprises a mobile unit that is in wireless communication with a server. The mobile unit is able to load expressions required for speech analysis or to generate speech output from the server. Expressions stored on the mobile unit can be supplemented or replaced. Which expressions are to be integrated into the mobile unit is determined depending on events, such as a concert or a sporting event, and their respective start and end times. In addition, context data can be taken into account to determine the expressions to be implemented in the mobile unit. Expressions that are rarely used or not used at all can be deleted from the mobile unit.While introducing other expressions into the mobile device increases the reliability that the voice assistant running on the mobile device actually understands a user, it does not improve the natural interaction with the voice assistant.

[0009] Furthermore, DE 102015213722 A1 discloses a method for operating a speech recognition system in a vehicle and a corresponding speech recognition system. The speech recognition system is capable of operating a vehicle function taking into account the context prevailing in the respective situation. Furthermore, DE 102022 000 387 A1 discloses a method for processing speech inputs and an operating device for controlling vehicle functions. Several vehicle functions suitable for execution are displayed on a display device. The execution of a respective vehicle function can be initiated by a voice command, whereby a vehicle function highlighted on the display device can be started by a shortcut command.

[0010] Furthermore, DE 102016218270 B4 discloses a method for operating a motor vehicle operating device with a speech recognizer and a corresponding operating device. Depending on the current driving situation, a command word is predicted for a user based on a behavior model. Subsequently, a check is performed to determine whether the command word is present in a list of recognition results. One of the recognition results contained in the list can be selected for the speech recognizer, or the recognition results can be reordered according to a confidence order.

[0011] Furthermore, WO 2014 / 060054 A1 discloses speech recognition in a motor vehicle. At least a portion of a speech input recognized in the vehicle is fed to both an in-vehicle speech recognizer and an external speech recognizer. The speech text underlying a speech input is determined by a processing device in agreement with a recognition result provided by the in-vehicle and external speech recognizers.

[0012] The present invention is based on the object of providing an improved method for operating a vehicle-integrated voice assistant, with the aid of which a natural interaction with the voice assistant is enabled, in particular by shortening the time required for the voice assistant to react.

[0013] According to the invention, this object is achieved by a method for operating a vehicle-integrated voice assistant having the features of claim 1. Advantageous embodiments and further developments, as well as a correspondingly operable voice assistant and a vehicle with such a voice assistant, are disclosed in the dependent claims. A generic method for operating a vehicle-integrated voice assistant, wherein the voice assistant captures and processes a voice input made by a user and determines a context-dependent operating intention for operating a vehicle function from the voice input, comprises the following method steps executed by an in-vehicle computing unit:

[0014] - Logging the usage behavior of the vehicle functionalities, whereby each user interaction with a vehicle functionality is assigned a context;

[0015] - Matching the context-dependent operating intention with the contexts logged during user interaction with the vehicle functionalities; and

[0016] - if an agreement is found between a recorded context and the operating intention: operating the vehicle functionality to be controlled by the voice input, taking the context into account.

[0017] According to the invention, it is provided that the computing unit assigns a timeliness factor to a respective context when logging the usage behavior of the vehicle functionalities, wherein the timeliness factor assumes its maximum value during assignment and then steadily decreases over time, and wherein, in the case of several contexts that match an operating intention, the computing unit selects the context for operating the vehicle functionality whose timeliness factor has the greatest value at that moment.

[0018] The method according to the invention enables particularly intuitive user interaction with the voice assistant. The voice assistant is able to determine which operating input is to be made, taking the context into account, so that the user no longer has to formulate corresponding voice commands in full, as is usual. Instead, individual semantic content is determined by the voice assistant itself, taking the context into account. For example, the user no longer has to say: "Hey Mercedes, call the contact Anna Schmidt from the phone book," but rather: "Call Anna." This is possible if, for example, the user has viewed an address book within the last 10 minutes, to which the computing unit has at least read access. Thus, in this example, the user interaction with the vehicle functionality corresponds to opening and viewing the address book.The computing unit logs this usage behavior and determines the respective context. In this example, the context could be that a name from the address book is to be provided as an input for a respective vehicle function. The operating intention underlying the voice input is seen as starting a telephone call. Thus, a telephony function is to be used as a vehicle functionality. Telephony functions are generally associated with the fact that contact names can be provided as an input, in addition to, for example, the direct entry of a telephone number. Reading a name from the address book is available as the context. As additional information, the computing unit determines from the voice input that “Anna” should be selected as the name. The computing unit is therefore able to correctly interpret the shortened voice command to initiate the call.This shortened voice command is formulated in natural language, so that the voice interaction between the user and the voice assistant can be designed to be particularly natural and therefore intuitive.

[0019] The vehicle-integrated voice assistant can otherwise be implemented as usual. The necessary hardware components, such as acoustic detection devices like microphones and corresponding processing components like the aforementioned processing unit, are present. A signal generated by the microphone can be read by a speech recognition module and converted into a character string, such as a text string. This character string can be read and processed using a natural language recognition module. This enables the recognition of corresponding operating intentions and other information, such as, referring back to the previous example, the name to be called, in the character string.

[0020] The vehicle functionalities that can be recorded can preferably include vehicle functionalities that go beyond the voice assistant, such as operating the vehicle's infotainment system, adjusting the air conditioning, selecting a sports driving mode, selecting a typical speed of travel, i.e. driving behavior in the broader sense, and the like.

[0021] The following could, for example, also be determined as context: a location, a time of day, a parameter of a vehicle functionality such as a tuned radio station, a target temperature, a driving mode, etc. According to the invention, as already described, the method provides that the computing unit assigns a timeliness factor to a respective context when logging the usage behavior of the vehicle functionalities, wherein the timeliness factor assumes its maximum value upon assignment and then steadily decreases over time, and wherein, if there are several contexts that match an operating intention, the computing unit selects the context for operating the vehicle functionality whose timeliness factor has the greatest value at that moment. With the help of the timeliness factor, suitable contexts can thus be selected even more accurately, increasing the probability that, even taking natural language ornaturally formulated voice inputs, the correct operating intention is recognized and the respective vehicle functionality is operated accordingly.

[0022] This approach is based on the idea that the voice inputs made by the user most likely relate to the last vehicle function activated or operated. A further differentiation is made based on the respective type of vehicle functionality. For example, if the user opens the address book and then sets the target temperature of the air conditioning system to a different value, the processing unit is able to relate the respective context to the address book rather than the climate control functionality when a corresponding voice command is given to initiate a call. However, several similar uses of the same vehicle functionality could be made, whereby the most recent user interaction is appropriately considered, taking the timeliness factor into account.

[0023] The timeliness factor can take on any value when generated, for example 1. Over time, the timeliness factor decays, for example to a value of 0. The respective expiration time can be chosen differently depending on various boundary conditions, which will be discussed in more detail below.

[0024] According to an advantageous embodiment of the method according to the invention, the computing unit uses a suitably trained machine learning model to compare the context-dependent operating intention with the logged contexts. With the help of artificial intelligence, in particular using machine learning models, a sufficiently trained model can reliably reference corresponding operating intentions with corresponding contexts, thereby increasing the probability that the operating intention underlying a given voice input will be correctly implemented. The machine learning model can, for example, have been initially trained by the vehicle manufacturer during the development of the voice assistant.

[0025] A further advantageous embodiment of the method according to the invention further provides that the speed at which a respective topicality factor expires is set by a respective topicality factor-specific decay factor. The decay factor can, for example, define a ratio by how many points the topicality factor is to be reduced per unit of time, for example by 0.1 points per 10 minutes or the like. Individual decay factors can then be defined for different vehicle functionalities or different contexts. Operating actions and corresponding vehicle functionalities that are better retained in the user's long-term memory can then be linked to a low decay factor, so that the respective update factor decreases more slowly and corresponding operating actions orVehicle functions that the user forgets more quickly can be associated with a comparatively higher decay factor. This increases the reliability that the correct operating intention is recognized and the respective vehicle function is operated appropriately.

[0026] Preferably, the computing unit sets the level of at least one expiration factor depending on:

[0027] - the type of vehicle functionality underlying the respective context;

[0028] - a manual user setting; or

[0029] - the detection of an interruption of the operation of the vehicle functionality carried out in accordance with the voice input.

[0030] An expiration factor dependent on the respective vehicle functionality can, for example, be hard-coded into a database, e.g., a table, by the vehicle manufacturer. It would also be possible for the user to manually specify a corresponding expiration factor. For this purpose, the vehicle can be equipped with appropriate human-machine interfaces that allow the receipt of corresponding user inputs or user specifications. For example, the expiration factor for a specific context of a specific vehicle functionality can be set to the value 6 on a scale ranging from 0 to 10, where the factor 6 corresponds, for example, to a deterioration of 0.6 points of the timeliness factor within 10 minutes. The user can then manually change this value to, for example, 2. This allows the personal preferences of different users to be taken into account.

[0031] In a particularly preferred embodiment, the computing unit is capable of automatically adjusting the expiration factor of a respective context of a respective vehicle functionality. To this end, the computing unit monitors user interaction with the voice assistant. If the user aborts the execution of a corresponding voice input shortly after its output, this indicates that the voice assistant misinterpreted the operating intention and / or the context. The cause could therefore be the consideration of a context that is no longer current. By automatically adjusting the expiration factor, the corresponding context can then be expired more quickly, so that the appropriate context can be selected instead.

[0032] A further advantageous embodiment of the method according to the invention further provides that the computing unit deletes a logged context as soon as the relevance factor assigned to the respective context has expired. Contexts whose relevance factor has expired are likely no longer relevant for operating vehicle functions. Accordingly, the contexts logged in this way can be deleted, so that less storage space needs to be reserved on a corresponding computer-readable storage medium of the computing unit. The freed-up storage space can then be used for other purposes.

[0033] According to a further advantageous embodiment of the method according to the invention, the computing unit predicts at least one potentially next expected speech input from the currently logged contexts, in particular several potentially next expected speech inputs, before the next speech input occurs, wherein the computing unit assigns to each potentially next expected speech input an input probability determined depending on the respective context. If several potentially next expected speech inputs are determined, the input probability can be used to determine an order which of these potentially expected speech inputs will actually follow. If the computing unit only determines one expected speech input, the input probability does not need to be determined.To determine the input probability, the computing unit can particularly preferentially consider the recency factor. Thus, the input probability increases with a higher value of the respective recency factor, so that voice inputs are particularly expected that relate to recent user interactions with the respective vehicle functionalities.

[0034] Predicting expected speech inputs makes it possible to provide entirely new functionalities. A corresponding advantageous development of the method according to the invention provides that, at least for the most likely next expected speech input, the computing unit, before it is entered,

[0035] - refers to background information required to formulate a voice response; and / or

[0036] - carries out at least one intermediate step required to operate the respective vehicle functionality.

[0037] Particularly advantageously, these steps can also be performed for several voice inputs expected next, so that if the computing unit misjudges which voice input will actually be made, the background information required for the actual voice input is still available or the aforementioned intermediate vehicle functionality steps have already been performed. By obtaining the background information early, or

[0038] By performing the respective intermediate steps, the execution time of each operation of the vehicle's functionality can be shortened. This increases user comfort, as the user has to wait less time. This improves the efficiency of the voice assistant.

[0039] For example, the user can activate route guidance to a destination. The logged vehicle functionality is then navigation. The destination, for example, can be determined as the context. The user can then say, for example, “What will the weather be like?” as a voice input. The computing unit then determines from this that the user wants to know the weather at the destination at the time of arrival. Even before the user makes this voice input, the computing unit can then preferably already query the relevant weather information at the destination at the time of arrival from a weather service. This is only possible after the user has activated route guidance to the destination. In this case, the timeliness factor or the expiration factor can be selected such that the timeliness factor for the destination and destination arrival time context will expire at the earliest when the destination is reached.This makes it possible to correctly interpret naturally formulated voice inputs for requesting the weather at any time during the journey.

[0040] Since the information required for this in the form of weather report data was already requested when programming the navigation route, this information can be output directly after the voice input has actually been received.

[0041] In a voice assistant comprising acoustic detection means, a computing unit, and a connection for operating a vehicle function, the detection means, the computing unit, and the connection for operating the vehicle function are configured according to the invention to execute a method described above. With the aid of the computing unit, signals generated by the acoustic detection means can be processed, thereby correctly interpreting the respective voice inputs. The computing unit is then capable of controlling the respective vehicle functions for operation via a corresponding interface.

[0042] A vehicle according to the invention has such a voice assistant. This enables particularly convenient and intuitive use of the voice assistant while driving. It can be any road vehicle such as a car, truck, van, bus, or the like. Generally, it could also be a rail vehicle, watercraft, or aircraft.

[0043] Further advantageous embodiments of the method according to the invention for operating a vehicle-integrated voice assistant also emerge from the exemplary embodiments which are described in more detail below with reference to the figures.

[0044] 1 shows a flowchart of a method according to the invention for operating a vehicle-integrated voice assistant;

[0045] Fig. 2 is a flowchart of the intermediate steps performed by a voice assistant according to a first embodiment to respond to a voice input; and

[0046] Fig. 3 is a flowchart of the intermediate steps performed by a voice assistant according to another embodiment to respond to a voice input.

[0047] Interacting with a voice assistant can be uncomfortable for a user if commands must be delivered in a specific format, especially when certain keywords are used. It is therefore desirable to make interaction with a voice assistant as natural and intuitive as possible for the user. This is where an inventive method for operating vehicle-integrated voice assistants comes into play.

[0048] The method involves carrying out the method steps shown in Figure 1. In method step 101, an in-vehicle computing unit logs the usage behavior of vehicle functionalities, with each user interaction with a vehicle functionality being assigned a context. In the subsequent method step 102, the computing unit compares the context-dependent operating intention, which was determined from a corresponding voice input, with the contexts logged during the user interaction with the vehicle functionalities. If a match is found between a logged context and the operating intention, in method step 103 the vehicle functionality to be controlled by the voice input is operated taking the context into account.

[0049] Taking the logged context into account, the computing unit is able to determine semantic content for interpreting voice inputs from previous user interactions with vehicle functionalities, so that corresponding voice inputs or voice commands can be formulated more briefly and in simpler language.

[0050] A preferred embodiment of the method according to the invention is illustrated in Figures 2 and 3. Thus, the vehicle's processing unit can predict expected future voice inputs, which can be used to obtain background information required to respond to a voice input before the respective voice input is actually input, or to execute intermediate steps required to respond to the voice input in advance. This can shorten the user's waiting time after making the voice input.

[0051] Figure 2 shows the procedure for responding to a voice command according to a first embodiment of the voice assistant. In method step 201, the vehicle user is proactively asked whether a charging station should be searched for. In method step 202, the user confirms this, for example by saying "yes." This is where the actual waiting time for the user begins. In method step 203, the computing unit transmits a request to a central computing device, such as the cloud server operated by a vehicle manufacturer, to retrieve information about nearby charging stations. In method step 204, the computing unit receives the corresponding results and prepares them. For this purpose, suitable charging stations can be suggested according to known optimization algorithms, and the results can be formatted for easy interpretation. In method step 205, the user then selects a suitable charging station.In process step 206, the computing unit queries the current status of a route guidance. In process step 207, the computing unit makes a decision regarding an action to be performed. In process step 208, the computing unit executes the corresponding action, such as setting the specified charging station as an intermediate destination for the route guidance. This ends the user's waiting time.

[0052] Figure 3 illustrates the sequence of method steps according to a preferred embodiment of the voice assistant. In method step 301, the computing unit sends a request to the central computing device to obtain information about charging stations. In method step 302, the computing unit processes the received results. In method step 303, the computing unit queries the current status of the route guidance. In method step 304, the computing unit makes a decision about the actions to be performed. Only in method step 305 is the proactive dialog initiated: "Should I search for a charging station?". In method step 306, the user responds, for example, by saying "Yes." Thus, the waiting time for the user only begins at this point.In process step 307, the computing unit can then directly select a result already prepared from the corresponding background information and perform the corresponding action in process step 308, for example, setting the charging station as an intermediate destination. This ends the user's waiting time.

[0053] According to the method steps illustrated in Figure 2, the waiting time for the user lasts from method step 202 to method step 208, whereas the waiting time according to the method steps shown in Figure 3 only lasts from method step 306 to method step 308. This clearly demonstrates the shortening of the waiting time for the user.

[0054] The method according to the invention makes it possible to interpret the user's actual intentions more reliably, especially when voice input is given through indirect speech or when disambiguating voice commands. Distinguishing between information types also provides additional value, since the relevance of contexts also differs for different types of information. Furthermore, the user's waiting time for the voice assistant's response can be shortened.

Claims

Patent claims 1. A method for operating a vehicle-integrated voice assistant, wherein the voice assistant detects and processes a voice input made by a user and determines from the voice input a context-dependent operating intention for operating a vehicle functionality, comprising the following method steps carried out by a vehicle-internal processing unit: - (101): Logging the usage behavior of the vehicle functionalities, whereby each user interaction with a vehicle functionality is assigned a context; - (102): comparing the context-dependent operating intention with the contexts logged during the user interaction with the vehicle functionalities; and - if an agreement is established between a logged context and the operating intention: (103): operating the vehicle functionality to be controlled by the voice input, taking the context into account; characterized in that the computing unit, when logging the usage behavior of the vehicle functionalities, assigns a timeliness factor to a respective context, wherein the timeliness factor assumes its maximum value upon assignment and then steadily decreases over time, and wherein, in the case of several contexts that match an operating intention, the computing unit selects the context for operating the vehicle functionality whose timeliness factor has the greatest value at that moment.

2. Method according to claim 1, characterized in that the computing unit uses a correspondingly trained machine learning model to compare the context-dependent operating intention with the logged contexts.

3. Method according to claim 1 or 2, characterized in that the speed with which a respective timeliness factor expires is set by a respective timeliness factor-specific decay factor.

4. Method according to claim 3, characterized in that the computing unit sets the level of at least one decay factor depending on: - the type of vehicle functionality underlying the respective context; - a manual user setting; or - the detection of an interruption of the operation of the vehicle functionality carried out in accordance with the voice input.

5. Method according to one of claims 1 to 4, characterized in that the computing unit deletes a logged context as soon as the timeliness factor assigned to the respective context has expired.

6. Method according to one of claims 1 to 5, characterized in that the computing unit, before the next speech input occurs, predicts at least one potentially next expected speech input from the currently logged contexts, in particular a plurality of potentially next expected speech inputs, wherein the computing unit assigns to each potentially next expected speech input an input probability determined as a function of the respective contexts.

7. Method according to claim 6, characterized in that the computing unit at least for the most likely next expected speech input: - refers to background information required to formulate a voice response; and / or - carries out at least one intermediate step required to operate the respective vehicle functionality.

8. Voice assistant, comprising acoustic detection means, a computing unit and a connection for operating a vehicle functionality, characterized in that the detection means, the computing unit and the connection for operating the vehicle functionality for carrying out a method according to one of the claims 1 to 7 are set up.

9. Vehicle characterized by a voice assistant according to claim 8.