Method for transforming a voice input into a voice control instruction, multimodal control system and method for controlling a multimodal control system

The method improves multimodal user interface usability by integrating graphical and contextual data to transform natural language speech inputs into voice control instructions, enhancing efficiency and reliability.

WO2026012780A1PCT designated stage Publication Date: 2026-01-15MERCEDES BENZ GROUP AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/068388
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-08
Filing Date
2025-06-27
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing multimodal user interfaces in vehicles do not adequately consider the underlying graphical representation on the GUI when processing natural language inputs, leading to reduced usability and efficiency.

Method used

A method that transforms natural language speech input into voice control instructions by utilizing both graphical user interface data and contextual data, including interaction history and sensor data, to accurately determine user intentions and reduce the search space for matching voice control commands.

Benefits of technology

Enhances the efficiency and reliability of multimodal user interface operations by accurately determining user intents and reducing the need for additional interaction steps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025068388_15012026_PF_FP_ABST
    Figure EP2025068388_15012026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a computer-implemented method for transforming a natural-language voice input (SE) into a voice control instruction (SA) for a multimodal user interface (multimodal UI) (1) comprising a graphical user interface (graphical UI, GUI) (2) and a voice user interface (voice UI) (3), a voice control instruction (SA) comprising at least one voice control command (SB) and at least one voice control parameter (SP), each of which is associated with a parameter slot (SL). GUI text data (2.T) currently displayed by the GUI (2) are acquired. Based on an interaction profile of the multimodal UI (1), additional current context data (K) that go beyond the displayed GUI text data (2.T) and are the basis for a current display by the GUI (2) are continuously determined. The voice input (SE) is broken down into voice input tokens. At least one voice control command (SB) and / or at least one voice control parameter (SP) are / is determined from in each case at least one voice input token on the basis of the current context data (K). The search space for associating a voice control command (SB) and / or a voice control parameter (SP) with at least one voice input token is limited in accordance with the current context data (K) and the current GUI text data (2.T). The invention also relates to a multimodal control system (100) and a method for controlling a multimodal control system (100).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method for transforming a speech input into a speech control instruction, multimodal operating system and method for controlling a multimodal operating system

[0002] The invention relates to a method for transforming natural language speech input into a voice control instruction. Furthermore, the invention relates to a multimodal operating system comprising a graphical user interface (GUI) and a voice user interface (Voice-UL), thus enabling multiple modes of user interaction. The invention also relates to a method for controlling such a multimodal operating system.

[0003] Natural language speech input refers to spoken language that is not subject to any particular formal grammar, especially not to restrictions on certain keywords and / or rules regarding the arrangement of speech elements. A "voice control instruction" refers to a well-formed expression, according to a predefined formal grammar, for interacting with a user interface (UI). Specifically, a voice control instruction comprises a voice control command, which may be configured by one or more voice control parameters.

[0004] Modern vehicles feature user interfaces (UIs) that can be operated via graphical input and output. For example, infotainment systems and navigation systems can be operated via such UIs. Even mobile devices that are only temporarily assigned to a vehicle, such as mobile phones, can be operated via a vehicle-provided UI. With a graphical input and output (hereinafter referred to as a Graphical User Interface, or GUI), output is typically displayed on a screen. Input can also be via the screen, which might be a pressure-sensitive touchscreen, for example. Alternatively or additionally, input can be via physical controls such as buttons, switches, or rotary push-buttons that are integrated into the screen.

[0005] Voice-controlled ULs (hereinafter referred to as voice ULs), which can be operated via voice input and output, are also known. Spoken language input and output is handled via at least one loudspeaker and at least one microphone, respectively. Methods and devices are known for transforming a user's naturally spoken language into instructions that are transmitted by the UL to at least one associated device or system.

[0006] Furthermore, multimodal user interfaces are known that interact with a vehicle occupant via a GUI as well as via spoken language input / output and optionally use additional sensors to determine an operating intention.

[0007] For example, document DE 102015210430 A1 describes a method for recognizing a language context for voice control in a vehicle. The method comprises a step of reading gaze direction information about the current gaze direction of a vehicle occupant, a step of assigning the gaze direction information to a viewing zone in the vehicle's interior to obtain viewing zone information about a viewing zone currently being viewed by the occupant, and a step of determining language context information about a predetermined language context assigned to the currently viewed viewing zone using the viewing zone information.

[0008] Document EP 3 955244 A1 discloses a voice control method for a vehicle. The voice control method comprises: receiving voice input information; transmitting the voice input information and current graphical user interface (GUI) information of the vehicle to a server; receiving a voice control operating instruction generated by the server based on the voice input information, the current GUI information, and the voice interaction information corresponding to the current GUI information; and parsing the voice control operating instruction and executing an action as if it were in response to a touch operation in accordance with the voice control operating instruction.In the voice control method for the vehicle according to one embodiment of the present disclosure, a semantic understanding of the voice input information is performed in combination with the current GUI information of the vehicle during the voice control process. This improves the semantic understanding capability of a voice assistant and allows elements in the GUI to be operated by voice, thus providing the user with a more convenient mode of interaction, a higher level of intelligence, and a better user experience. The document also discloses an information processing method, a vehicle, a server, and a storage medium.

[0009] Document DE 10360 655 A1 describes an operating system for a vehicle with a screen display with multiple display areas for showing entries of a menu structure with multiple menu levels, a manual operating means for selecting and / or activating at least one entry in a current menu level from the menu structure, and voice control means for a redundant selection and / or activation of at least one entry from the menu structure, which simultaneously forms a keyword for the voice control means.The menu structure entries are divided into different groups, with a first group comprising entries that can only be selected and / or activated using manual controls, a second group comprising entries that can be selected and / or activated using manual controls and / or voice control, and the second group being divided into at least two groups of terms that can be defined by simple rules and that determine which keywords can currently be entered for menu navigation.

[0010] Document US 2022 / 0284904 A1 describes a method for displaying text messages containing n-grams via a user interface of a user-generated operating system.

[0011] Document US 2007 / 0256027 A1 describes a control system for a vehicle comprising a screen with multiple display areas for showing entries of a menu structure, a manual operating device for selecting and / or activating menu entries, and a voice input control for redundant selection and / or activation of a menu entry of the menu structure, which also forms a keyword for the voice input control.

[0012] However, known speech analysis methods do not consider, or only partially consider, the content underlying a current graphical representation on a GUI when operating such multimodal ULs. In particular, natural language input (i.e., user-speak) is limited to text currently displayed on the GUI. For example, a GUI can often only display a portion of a list of data (addresses, phone numbers, audio titles, or similar data) retrieved based on a search criterion, due to space constraints. In such cases, known speech analysis methods only use the portion of the list currently visible on the GUI to compare the user's natural language input with the GUI. This makes the usability of a multimodal UL more difficult.

[0013] According to a first aspect, the invention is based on the objective of providing an improved method for transforming a natural language speech input into a voice control instruction for a multimodal UL comprising a GUI and a voice UL. This objective is achieved according to the invention by a method with the features of claim 1.

[0014] According to a second aspect, the invention aims to provide an improved multimodal Ul. This objective is achieved according to the invention by a method with the features of claim 6.

[0015] According to a third aspect, the invention aims to provide a method for controlling a multimodal Ul according to the second aspect of the invention. This objective is achieved according to the invention by a method with the features of claim 9.

[0016] Advantageous embodiments of the invention are the subject of the dependent claims. In a computer-implemented method for transforming natural language speech input into a voice control instruction for a multimodal user interface, a voice control instruction comprises at least one voice control command and at least one voice control parameter, each of which is assigned to a parameter slot of the voice control instruction. The multimodal user interface comprises a GUI and a voice UL. The method can be implemented on a control unit, which can be configured, for example, as a head unit or as a vehicle control unit.

[0017] The GUI is designed to display both graphical data (such as icons, backgrounds, maps, and similar elements) and textual data (such as contact lists, points of interest, radio stations, and similar information). Textual data currently displayed on the GUI is recorded as GUI text data.

[0018] A user interacts with the multimodal UL through a sequence of steps, such as spoken language and / or the use of buttons or rotary pushbuttons, which are recorded as an interaction history. Based on this interaction history, current contextual data is continuously determined, describing the current state of the multimodal UL in a way that goes beyond the currently displayed text.

[0019] For example, in an interaction history designed to display contact information based on a specific search criterion entered during the interaction, the contextual data can capture all contacts matching that criterion. The textual display on the GUI, however, can be limited to just a few contacts. Furthermore, the contextual data can capture all information recorded for each contact (e.g., last name, first name, nickname, mobile phone number, work landline number, home landline number, date of birth, and similar information), while the GUI text data is limited to the last name and first name.

[0020] Voice input is broken down into voice input tokens. A voice input token comprises a spoken word or a plurality of words spoken within a meaningful context, for example, "call," "on the mobile phone," "make an appointment," "this week," and similar phrases. A voice control command and / or a voice control parameter are determined from at least one voice input token based on the current context data. The search space for matching a voice control command and / or a voice control parameter to at least one voice input token is restricted according to the current context data and the current GUI text data.

[0021] For example, if the current context data indicates a search for a contact, only voice commands relevant to a contact or a list of selected contacts will be considered. Furthermore, voice input tokens that are ambiguous in relation to the context data will be limited by the currently displayed GUI text data.

[0022] This allows user intents to be formulated as voice input that cannot be expressed using only the GUI (for example, making a call via a connection option not displayed for a contact in the GUI, or selecting a point of interest based on criteria not represented textually in the GUI). Furthermore, user intents that are ambiguous with respect to the entirety of the context data can be disambiguated using the GUI text data (for example, if the entire search results list includes multiple contacts with the same name, but only one is represented textually).

[0023] Thus, the proposed method can improve the efficiency and reliability of operating a multimodal UL.

[0024] In one embodiment of the method, if only the GUI text data changes as a result of user interaction, but not the context data, GUI index data is captured. GUI index data indicates the portion of the context data that is currently displayed as text in the GUI.

[0025] For example, a list of contacts filtered according to a search criterion from the past interaction history can be displayed only partially in the GUI. By turning a rotary push-button, the portion of the list displayed in the GUI is changed, while the search result (the list of filtered contacts) remains unchanged. The GUI index data, for example via field indexes of a list implemented as a field, indicates the currently displayed portion of the list. This implementation allows for particularly efficient execution. The communication overhead between the GUI and the voice UL is reduced by exchanging only the GUI index data (for example, a field of a few integer values) instead of the GUI text data and / or context data.

[0026] In one embodiment, the current context data comprises one, preferably several, operating options that are displayed textually on the GUI and thus captured in the GUI text data. At least one of these operating options is assigned an index element and visualized on the GUI. An "index element" here is understood as a graphically or textually visualizable display element of the GUI that can be easily and unambiguously distinguished from other index elements by an indexing phrase, for example, a single digit that can be spoken as a number word.

[0027] The search space for assigning a language input token is limited to the recognition of natural language indexing words that correspond to the visualized index elements (for example, the recognition of number words).

[0028] When transforming speech input into a speech control instruction, a speech control command or a speech control parameter is assigned according to the operating option, to which the index element is assigned, which was identified based on the natural language indexing wording (for example, one of the number words "one", "two", "three" corresponding to the index elements "1", "2", "3" of a first, second, and third operating option). An operating option can also be assigned both a speech control command and a speech control parameter, or multiple speech control parameters.

[0029] This embodiment enables a particularly easy and also very reliable assignment of a voice control instruction.

[0030] In a further development of this implementation, a subset of textually displayed operating options is selected for visual indexing via the voice UL. For example, if the GUI displays several lists of operating options textually, exactly one of these lists can be visually indexed (i.e., marked with numbers or similar index elements). This enables particularly easy and also very reliable selection of operating options, even with complex GUI displays.

[0031] In one embodiment of the method, the current context data includes at least one complete list of operating options, each determined during the preceding interaction history. The GUI text data comprises only an (incomplete) part of such a list. In other words, the at least one list is only partially displayed as text on the GUI.

[0032] The search space for mapping a voice input token to a voice control command and / or a voice control parameter is adapted to the complete list of operating options. "Search space" refers to a (not necessarily countable) set of possible results for mapping voice input tokens to a voice control command or a voice control parameter.

[0033] In other words, the system also searches for voice control commands or parameters that match a control option not displayed textually on the GUI. This eliminates interaction steps that would otherwise be required by operating the GUI (for example, scrolling through a list with a rotary push-button) or by interacting with the voice UL (for example, querying further filter criteria).

[0034] According to a second aspect of the invention, a multimodal operating system comprises an operating control unit, a graphical user interface (GUI) configured to implement a GUI, and a voice control unit configured to implement a voice UI. The graphical user interface is configured to display and transmit GUI text data to the operating control unit. The operating control unit is configured to continuously acquire current context data based on an interaction history and to transmit the current context data and the GUI text data to the voice control unit. The voice control unit is configured to implement a method according to the first aspect of the invention.

[0035] One advantage of such a multimodal operating system is that, to determine operating intentions from a voice command, not only textual representations of the GUI but also contextual data can be used. This enables easier and more reliable operation. Further advantages correspond to the advantages of the method according to the first aspect of the invention.

[0036] In one embodiment of a multimodal operating system, the operating control unit is designed as a control unit of a vehicle and is configured to acquire context data by means of sensors of the vehicle, which is suitable for limiting the search space for the assignment of a voice control command and / or a voice control parameter to at least one voice input token.

[0037] Such in-vehicle sensors could, for example, be a camera pointed at a user, capturing their gestures, facial expressions, or emotions. This data can then be used to determine user intent. For instance, a swiping motion of the user's hand could be identified as an intention to scroll through a list of operating options displayed on the GUI.

[0038] Such in-vehicle sensors can also record vehicle operating parameters, such as vehicle position, planned route, remaining fuel level, or battery capacity. These parameters can also be used to narrow down the search space for assigning voice input tokens to voice control parameters. For example, in a voice input...

[0039] "Hello Ul, looking for the nearest gas station" means that the nearest gas station is determined as the gas station that can be reached from the planned route with the smallest possible detour and with the remaining amount of fuel.

[0040] This design allows for particularly easy operation.

[0041] According to a third aspect of the invention, in a method for controlling a multimodal operating system according to the second aspect of the invention, the at least one graphical user interface and / or the at least one voice user interface are configured via a voice UL. For example, one of several available microphones can be selected using the voice UL. Additionally or alternatively, one of several displays on which a significant GUI is to be displayed can be selected, so that GUI text data and GUI index data only refer to the significant GUI and displays on other screens are ignored.

[0042] This simplifies the operation of the multimodal control system. Further advantages correspond to the advantages of the method according to the first aspect of the invention and the advantages of a multimodal control system according to the second aspect of the invention.

[0043] In addition to or as an alternative to configuration via a voice UL, in a method for controlling a multimodal operating system according to the second aspect of the invention, the search space for assigning a voice input token to a voice control command and / or a voice control parameter in a multi-step dialog is restricted. A multi-step dialog is understood to be a sequence of user interaction steps that take place via at least one voice UL.

[0044] This allows even more complex operating sequences to be implemented easily and without impairing the user's visual attention.

[0045] Exemplary embodiments of the invention are explained in more detail below with reference to drawings.

[0046] This shows:

[0047] Fig. 1 schematically shows a multimodal user interface,

[0048] Fig. 2 schematically shows a graphical user interface (GUI),

[0049] Fig. 3 schematically shows a context module for capturing context data,

[0050] Fig. 4 schematically shows a context module for capturing context data.

[0051] Calendar data and data retrieved from a backend

[0052] Fig. 5 schematically shows the process of a procedure for transforming a speech input into a speech control instruction,

[0053] Fig. 6 schematically shows a graphical user interface (GUI) with index elements as well as

[0054] Fig. 7 schematically shows a multimodal operating system. Figure 1 schematically shows a multimodal user interface 1 of an infotainment system of a vehicle (not shown in detail), hereinafter referred to as multimodal III 1. The multimodal III 1 comprises a graphical user interface 2, hereinafter referred to as graphical user interface GUI 2, and a voice user interface 3 coupled to it, hereinafter referred to as Voice-UL 3.

[0055] The Voice-Ul 3 includes at least one loudspeaker 3.1, one microphone 3.2 and one speech processor 3.3. The speech processor 3.3 is configured to analyze speech input SE captured by the microphone 3.2.

[0056] GUI 2 is displayed on a screen and, in this case, comprises graphical display elements as well as four text fields T1 to T4, each displaying a text. The entirety of the text displayed by GUI 2 will be referred to here and in the following as GUI text data 2.T. As a purely illustrative example, GUI 2 could display part of a list L obtained as the result of a search query for all contacts containing the surname "Maier".

[0057] The first (topmost) text field T1 displays the search query, in this case the string with the search criterion "Maier". The second to fourth text fields T2 to T4 display the strings "Andreas Maier", "Frank Maier" and "Michael Maier" respectively.

[0058] Conventional methods evaluate the graphical representation of the GUI 2 and, in particular, capture the GUI text data 2.T using a method of optical character recognition (OCR) in order to determine an intended operation of the GUI 2 from speech inputs SE.

[0059] In contrast, the invention proposes to determine context data K from an interaction history on the UL 1 and to transfer it to the voice UL 3. In this case, the interaction history includes the search query for the surname "Maier". The context data K comprises the complete list L of contacts found according to the search query. Due to the space limitations of the GUI 2, this complete list L includes, in addition to the contacts displayed in the text fields T2 to T3, further contacts, for example, "Peter Maier", "Siegfried Maier", "Thomas Maier", and others. Furthermore, the complete list L transmitted with the context data K includes additional information not displayed in the GUI 2, for example, a nickname "Andi" assigned to the contact "Andreas Maier", a telephone number and / or address assigned to the respective contact, and optionally, other data.

[0060] The totality of this context data K enables a more accurate and comprehensive comparison of the speech input SE captured by the Voice-Ul 3 and enables a more accurate, convenient and reliable determination of an operating intention with respect to the multimodal Ul 1 than is possible using only the extracted GUI text data 2.T.

[0061] Furthermore, the context data K includes the current state of the multimodal Ul 1 as a result of the preceding interaction sequence. For example, a voice input SE can take the form

[0062] The voice command "Ul, please call ANDREAS" is disambiguated by Voice-Ul 3 to indicate that ANDREAS refers to the contact "Andreas Maier" displayed in the second text field T2, even though the contact database may contain other contacts with the string "Andreas" (for example, "Andreas Müller", "Andreas Zimmermann", or "Friederike Andreas"). This reduces the user effort compared to disambiguation through complete user input (either via GUI 2 or Voice-Ul 3).

[0063] In another embodiment, also illustrated with reference to Figure 1, a search query for restaurants located within a certain radius of the vehicle's current position is initiated during the interaction. In this embodiment, the first text field T1 displays the string "Restaurants". The second to fourth text fields T2 to T4 display a portion of a list L of restaurants, sorted, for example, by their respective distances. For instance, the second text field T2 displays the string "Restaurant Pizzeria MamaMia", the third text field T3 displays the string "Restaurant El Greco", and the fourth text field T4 displays the string "Restaurant Rasputin". In addition to the visually visible GUI text data 2.T (“Restaurants”, “Restaurant Pizzeria MamaMia”, “Restaurant El Greco”, “Restaurant Rasputin”) of the list L include the associated contextual data K, further information, for example, about the regional characteristics of the cuisine, opening hours, user ratings, distance from the current vehicle position, or similar data. This allows voice input SE in the form

[0064] The command "Ul, call the Italian restaurant" can be clearly resolved without further interaction, as only one of the three displayed restaurants has the attribute "Italian cuisine." Furthermore, other restaurants detected by the search query but not listed in GUI 2 are not considered when evaluating the voice input SE. For this purpose, the state of GUI 2, also transmitted in the context data K, is evaluated by the Voice-Ul 3. From the context data K, the Voice-Ul 3 determines which restaurants are currently displayed and automatically limits the execution of the request to make a call to the context visually accessible to the user.

[0065] By evaluating the context data K, the multimodal Ul 1 can thus react particularly quickly to especially likely operating intentions, especially without further user interaction, as is explained in more detail in Figure 2.

[0066] Figure 2 schematically shows a GUI 2 displaying the list L of results from a search for contacts with the surname "Maier", as already explained using the first embodiment shown in Figure 1. The text fields T1 to T4 are populated with strings as follows:

[0067] From a voice input SE containing the spoken text "Hello Ul, please call ANDREAS on his mobile phone," a voice control instruction SA is created, comprising a voice control command SB and one or more voice control parameters SP. Each voice control parameter SP is assigned to a parameter slot SL. Depending on the voice control command SB, a voice control instruction SA can contain a fixed or variable number of parameter slots SL.

[0068] For this purpose, the speech input SE is first decomposed into speech input tokens using the speech processor 3.3. Possible speech input tokens could be: "CALL", "ANDREAS", "ON THE PHONE". Methods for decomposing speech input SE into speech input tokens are known from the prior art, for example from the publication by Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016. Google's neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv: 1609.08144.

[0069] The voice input token "CALL" is used to create a voice command SB "Call". This takes into account the context data K, which captures the partial display of a list L of callable contact data on the GUI 2.

[0070] From further voice input tokens, voice command parameters SP are derived, each of which is assigned to a parameter slot SL. The first voice command parameter SP assigned to the voice input token "ANDREAS" is the contact "Andreas Maier" specified by the second text field T2. This assignment is unambiguous because the visually displayed part of the list L contains only a single contact with the string "Andreas". This assignment is therefore possible without further interaction even if the list L of callable contact data includes other contacts ("Andreas Müller", "Andreas Zimmermann", or "Friederike Andreas") that are not displayed on GUI 2.

[0071] As a second voice command parameter, SP, the voice input token "ON THE MOBILE PHONE" is assigned the mobile phone number of the contact "Andreas Maier" stored in the context data K. Here, the speech processor 3.3 uses the synonym relationship "Handy" - "Mobile Phone". Such synonym relationships can be captured in a language model that is queried by the speech processor 3.3.

[0072] Optionally, the voice control instruction SA can include further voice command parameters SP determined from the voice input SE and the context data K, which are assigned to the subsequent parameter slots SL.

[0073] One advantage of the described method is that by evaluating the context data K (recognizing that a list L of callable contact details was requested in the previous interaction) and the GUI text data 2.T (recognizing that only a single contact containing the name segment "Andreas" is displayed on GUI 2), the search space for assigning voice input tokens to voice control commands SB or voice control parameters SP is narrowed. This allows for a more precise and reliable assignment that requires fewer interaction steps than conventional methods. Thus, in the present example, the prompt asking which contact should be assigned to the voice input token "ANDREAS" is eliminated.

[0074] Figure 3 schematically illustrates a possible implementation of the procedure. A context module KM continuously captures context data K, which is provided by a GUI 2 and a voice UL 3. The context module KM can be run on the

[0075] Speech processor 3.3, but can also be implemented on a processing unit outside the Voice-UL 3.

[0076] The GUI 2 continuously provides context data K. This allows a voice control instruction SA to be determined with minimal delay, efficiency, and reliability at the time of a voice input SE, as already explained in Figure 2. The evaluation of a voice input SE (that is, the assignment of voice control commands SB and voice control parameters SP) can be restricted to the search space resulting from the currently available context data K.

[0077] Furthermore, GUI 2 continuously transmits the currently displayed GUI text data 2.T or graphic data, from which the context module KM extracts the GUI text data 2.T using OCR. "Continuous" transmission (of context data K, of GUI text data 2.T) here refers to transmission that occurs at discrete intervals. These intervals are chosen to be so frequent that the context module KM is always notified of any change to context data K and / or GUI text data 2.T resulting from a user interaction. This notification occurs regardless of whether such user interaction takes place via GUI 2 or via the Voice UL 3.

[0078] Voice-Ul 3 provides a voice input SE to the context module KM when Voice-Ul 3 is activated.

[0079] The context module KM determines a voice control command SB and one or more voice control parameters SP for a voice control instruction SA based on the context data K and the GUI text data 2.T.

[0080] Conventional methods require the parameter slots SL of a voice control instruction SA to be populated through interaction with the operator (i.e., assigned a voice control parameter SP to each slot). This typically requires several interaction sequences, for example, by displaying multiple selection options on the GUI 2, of which the operator repeats or otherwise linguistically identifies only one. The proposed method eliminates at least some of these interaction sequences by deriving the voice control parameters SP through evaluation of the context data K.

[0081] By evaluating the context data K, a user's intent can be determined more accurately, flexibly, and comprehensively than by evaluating voice input SE and / or GUI text data 2.T alone. In addition to the GUI 2, context data K can also be provided by other vehicle sensors, such as a camera that captures gestures or (based on facial images) facial expressions or emotions of a user, or a control unit that monitors the vehicle's status.

[0082] Figure 4 shows, in an exemplary and purely schematic way, a context module KM, which is equipped with a

[0083] Voice-UL 3 is coupled, which is coupled to a user's calendar KL and to a backend B. Analogous to the situation shown in Figure 2, GUI 2 displays GUI text data 2.T to indicate a due maintenance, which contains the string

[0084] The GUI text data 2.T can be the result of an interaction history in which the display of a list L of open error messages and pending maintenance tasks is requested via various selection menus. This interaction history is recorded in the context data K. The list L can include further pending maintenance tasks that are not currently displayed on the GUI 2.

[0085] A voice input SE with the spoken text

[0086] "Hello Ul, can THIS be completed BY THEN?" can be transformed into a speech control instruction SA by evaluating both the context data K and the GUI text data 2.T with

[0087] - the voice control command SB "BOOK WORKSHOP APPOINTMENT", a first voice control parameter SP "WITHIN 10 DAYS", a second voice control parameter SP "TAKING MY CALENDAR INTO ACCOUNT".

[0088] The language processor 3.3 extracts from the GUI text data 2.T which of the outstanding maintenance work recorded in list L (currently displayed on the GUI 2) the language input SE refers to (in this case: “Service A”).

[0089] The Voice Processor 3.3 extracts the type of maintenance to be performed from the context data K. This context data K can include additional or more precise information than what is displayed on the GUI 2, such as the estimated duration of a workshop visit. Such context data K can also be obtained from a backend B, through which the Voice-Ul 3 accesses an unspecified cloud service. Furthermore, the Voice-Ul 3 provides context data K that includes data from a calendar KL (or authentication data that enables access to such a calendar KL). For example, a calendar KL on a smartphone (not specified) can be paired with the Voice-Ul 3 via a wireless connection, such as Bluetooth.

[0090] Using conventional methods that are limited to the analysis of GUI text data 2.T, such a speech control instruction SA cannot be obtained.

[0091] Figure 5 shows a schematic flowchart of the proposed procedure. In a first step S1, context data K is provided to the context module KM. In a subsequent second step S2, the context data K is classified. For example, from a plurality of possible operating domains (purely by way of example: establishing a telephone connection, operating a vehicle function, playing back infotainment content, operating a navigation system, querying a vehicle status), an operating domain is selected based on the context data K to which a subsequent interaction refers.

[0092] Optionally, additional information that may be assigned to a control domain is determined (purely by way of example: list L of selected contacts that can be called by telephone, list L of selectable parameters of a vehicle function, list L of playable audio content, list L of navigation destinations, list L of queryable parameters of a vehicle status).

[0093] In a subsequent third step S3, a context is provided that includes, for example, an operating domain, a list L of choices for this operating domain, a calendar KL assigned to the user, or other data to be assigned to a user and / or the vehicle.

[0094] In a subsequent fourth step, S4, a speech input SE is captured. Based on the context data K, processed as described, and the GUI text data 2.T currently displayed by GUI 2, a speech control instruction SA is determined in a subsequent fifth step, S5. Taking into account the information obtained from the context data K and the GUI text data 2.T, a speech control command SB and one or more speech control parameters SP are determined from the speech input SE. In a subsequent sixth step, S6, the speech control instruction SA is executed.

[0095] It is possible that in the fifth step S5, a voice control command SB cannot be uniquely determined and / or not all required voice control parameters SP can be uniquely determined and parameter slots SL assigned. As a result, a complete voice control instruction SA cannot be determined.

[0096] In one embodiment of the method, missing information (for example: identification of a single voice control command SB in a set of multiple voice control commands SB that are compatible with the current voice input SE, the current context data K and the current GUI text data 2.T; parameter slots SL not yet filled with voice control parameters SP) is added.

[0097] For this purpose, the context data K is updated according to the information already recognized from the speech input SE (already recognized voice control command SB; already recognized voice control parameters SP). Simultaneously, information is provided to the user via the GUI 2 and / or the Voice-UL 3, enabling supplementary selections. Using such iterations, referred to as Multi-Step Dialogs (MSD), the voice instruction SA is incrementally supplemented until complete.

[0098] Figure 6 schematically shows an embodiment of the method in which index elements 11 to I3 are displayed on the GUI 2. By way of example, in a representation of list elements of a list L on the GUI 2, as already explained with reference to Figure 3 (List L of contact details) and Figure 4 (List L of restaurants), each list element can be assigned an index element 11 to I3 in addition to a text field T2 to T4.

[0099] Index elements 11 to I3 can, for example, represent digits, letters, or similar text elements that are particularly easy to pronounce and readily identifiable as speech input tokens. The type of displayed list L can be determined using the context data K. Furthermore, the context data K, along with the GUI text data 2.T, can be used to determine which index element 11 to I3 is assigned to which text field T2 to T4. A speech input SE corresponding to the representation of an index element 11 to I3 can thus be directly assigned a speech control instruction SA.

[0100] For example, the speech input SE can be used with the number word assigned to the first index element 11.

[0101] “One” the voice control instruction SA

[0102] The command "Hello III, call «Participant 1> on their default phone contact» is assigned, where the string «Participant 1>» is replaced by the full name of the contact displayed in the second text field T2. For this assignment, it is not necessary for the full contact name to be displayed in the second text field T2, as it is provided by the context data K. Displaying the default phone contact (e.g., a landline phone number), also provided by the context data K, in text field T2 is also unnecessary. Furthermore, other contacts with the same name in list L, which are not displayed on GUI 2, do not prevent this unique assignment.

[0103] This makes interaction with the Voice-Ul 3 particularly easy.

[0104] Figure 7 schematically shows a multimodal operating system 100 comprising an operating control unit 10, a graphic operating unit 20 and a voice operating unit 30.

[0105] The graphical user interface 20 is configured to control a GUI 2. By way of example, the graphical user interface 20 includes a display (not shown in detail), preferably a touch display. Additionally, the graphical user interface 20 can include input devices (not shown in detail), such as buttons, switches, rotary push-buttons, or a keypad, which are linked to the display.

[0106] The graphical control unit 20 is also configured for exchanging context data K and GUI text data 2.T with the control unit 10 and with the GUI 2. The voice control unit 30 is configured for controlling a Voice-Ul 3. Furthermore, the voice control unit 30 is configured for exchanging context data K with the control unit 10 and with the Voice-Ul 3.

[0107] The control unit 10 is configured to implement a previously described procedure for transforming a speech input SE into a speech control instruction SA. In particular, the control unit 10 is configured to implement a context module KM.

[0108] The control unit 10 can, but does not necessarily have to, provide its own computing resources, such as its own processor. To implement the transformation of a speech input SE into a speech control instruction SA, the control unit 10 can also access the unspecified speech processor 3.3 of the Voice-Ul 3 or another computing resource of the GUI 2, the graphical control unit 20, or the speech control unit 30.

[0109] In one embodiment, the operating control unit 10 is designed as the main control unit (head unit) of a vehicle not shown in detail.

Claims

Patent claims 1. Computer-implemented method for transforming a natural language speech input (SE) into a speech control instruction (SA) for a multimodal user interface (multimodal III) (1) comprising a graphical user interface (graphical III, GUI) (2) and a speech user interface (Voice-UL) (3), wherein a speech control instruction (SA) comprises at least one speech control command (SB) and at least one speech control parameter (SP), each of which is assigned to a parameter slot (SL), characterized in that - the GUI text data (2.T) currently displayed by the GUI (2) is captured, - based on an interaction history of the multimodal Ul (1), additional current context data (K) extending beyond the displayed GUI text data (2.T) are continuously determined, which underlie a current display of the GUI (2), - the speech input (SE) is broken down into speech input tokens, - at least one voice control command (SB) and / or at least one voice control parameter (SP) is determined from at least one voice input token based on the current context data (K), wherein - the search space for an assignment of a voice control command (SB) and / or a voice control parameter (SP) to at least one voice input token according to the current context data (K) and the current GUI text data (2.T) is restricted.

2. Method according to claim 1, characterized in that - when the GUI text data (2.T) is changed without changing the context data (K), GUI index data is recorded, which designates the part of the current context data (K) displayed in the GUI (2).

3. Method according to one of the preceding claims, characterized in that - the current context data (K) must include at least one operating option (T2 to T4) determined in the interaction history and displayed in the GUI text data (2.T), - at least one of the operating options (T2 to T4) shown in the GUI text data (2.T) is assigned an index element (11 to I3), whereby - the assigned index element (11 to I3) is visualized on the GUI (2), - the search space for assigning a language input token is limited to the recognition of natural language indexing words, each corresponding to an index element (11 to I3) of one of the indexed visualized operating options (T2 to T4) and - a voice control command (SB) and / or a voice control parameter (SP) is / are assigned according to the operating option (T2 to T4) selected based on the natural language indexing wording.

4. Method according to claim 3, characterized in that a subset of operating options (T2 to T4) represented in the GUI text data (2.T) is selected for visual indexing via the Voice-Ul (3).

5. Method according to one of the preceding claims, characterized in that - the current context data (K) must include at least a complete list (L) of operating options (T2 to T4) identified in the interaction history, - where the GUI text data (2.T) represents only a part of the complete list (L) of operating options (T2 to T4) and where - the search space for assigning a voice input token to a voice control command (SB) and / or to a voice control parameter (SP) is adapted to the complete list (L) of operating options (T2 to T4).

6. Multimodal operating system (100) comprising an operating control unit (10), at least one graphic operating unit (20) configured for implementing a GUI (2), and at least one voice operating unit (30) configured for implementing a voice UL (3), characterized in that - the graphical user interface (20) is set up for displaying and transmitting GUI text data (2.T) to the user control unit (10), - the operating control unit (10) is set up for the continuous acquisition of current context data (K) based on an interaction history and for the transmission of the current context data (K), the GUI text data (2.T) to the voice control unit (30), wherein - the voice control unit (30) is set up to implement a method according to one of the preceding claims.

7. Multimodal operating system according to claim 6, characterized in that the operating control unit (10) is designed as a control unit of a vehicle and is configured to acquire context data (K) by means of sensors of the vehicle, which are suitable for limiting the search space for the assignment of a voice control command (SB) and / or a voice control parameter (SP) to at least one voice input token.

8. Method for controlling a multimodal operating system (100) according to claim 7, characterized in that the at least one graphic operating unit (20) and / or the at least one voice operating unit (30) are configured via a voice UL (3), wherein at least one natural language voice input (SE) is transformed into at least one voice control instruction (SA) for configuring the at least one graphic operating unit (20) and / or the at least one voice operating unit (30) using a method according to one of claims 1 to 5. Method for controlling a multimodal operating system (100) according to claim 7, characterized in that the search space for assigning a speech input token to a speech control command (SB) and / or a speech control parameter (SP) in a multi-step dialog (MSD) is restricted to a plurality of interaction steps implemented via the Voice UL (3), wherein at least one natural language speech input (SE) is transformed into at least one speech control instruction (SA) for configuring the at least one graphical operating unit (20) and / or the at least one speech operating unit (30) using a method according to one of claims 1 to 5.