Method for transforming naturally spoken language into a control command for operating a vehicle function.

A method using a catalog of implausible word sequences to filter out unintended voice commands in vehicle systems improves user experience and safety by accurately transforming spoken language into control commands.

DE102026112941A1Pending Publication Date: 2026-05-13MERCEDES BENZ GROUP AG
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
MERCEDES BENZ GROUP AG
Filing Date
2026-03-27
Publication Date
2026-05-13

AI Technical Summary

Technical Problem

Existing speech-based vehicle control systems misinterpret naturally spoken language, particularly idiomatic expressions and phrases without operational intent, leading to reduced user-friendliness and safety risks.

Method used

A method that includes a catalog of implausible word sequences to identify and eliminate unintended voice commands by comparing recorded speech against a catalog of implausible phrases, and optionally confirming intended commands through user interaction.

Benefits of technology

Reduces misinterpretation of unintended speech inputs, enhancing user-friendliness and safety by ensuring accurate vehicle function control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for transforming naturally spoken language into at least one control command (30) designed to operate a vehicle function. A catalog (20) of implausible word sequences for controlling vehicle functions is created, wherein each implausible word sequence comprises at least one word that is permissible as part of a valid voice command for operating a vehicle function. In the naturally spoken language, at least one current word sequence (11) is captured and compared with the catalog (20). If it matches at least one of the implausible word sequences in the catalog (20), it is eliminated or checked.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The invention relates to a method for transforming naturally spoken language into at least one control command designed to operate a vehicle function of a vehicle.

[0002] Speech-based input methods for vehicles capture the naturally spoken language of a vehicle occupant and generate control commands from it, which can be used to operate and parameterize vehicle functions. To do this, a speech instruction is first extracted from the naturally spoken language, which must adhere to certain predetermined syntactic and semantic rules.

[0003] A valid (that is, syntactically and semantically correct) speech instruction could, for example, be: "Go to Jena - Paradies train station!"

[0004] Such a voice instruction is transformed in the conventional manner into a control command, which prompts the vehicle's navigation system to calculate a route from the current location to the geoposition assigned to Jena - Paradies station and to determine and output corresponding navigation and driving instructions.

[0005] Modern speech-based input methods allow for a significant variation in the speech instructions recognized as valid, for example in the form "Go to Paradise Station" or "Take me to paradise!" where ambiguities can be resolved context-dependently (for example, by using statistical language models - large language models, LLMs - and taking into account the current geoposition of the vehicle).

[0006] Such methods save the user the tedious adherence to a strict syntax, such as the use of precisely prescribed keywords for certain vehicle functions. However, they also increase the likelihood that words, phrases, or sentences uttered without any intention of operation will be misinterpreted as voice commands. Such misinterpretation is particularly probable for imperative idiomatic expressions (for example, "Go to hell!").

[0007] Naturally spoken language, misinterpreted as a voice command and without any intention to control the vehicle, leads to unclear situations when operating vehicle functions. This can reduce user-friendliness and, in extreme cases, compromise operational and traffic safety. Therefore, there is a need for an improved method for transforming naturally spoken language into control commands, one that verifies the plausibility of the intended use.

[0008] Document US 7,437,297 B2 describes a system and a method for processing and executing commands in automated systems, for example, for determining, evaluating, or piecemeal predicting the results of executing incorrectly recorded or misinterpreted user instructions in such automated systems.

[0009] The invention is based on the objective of providing an improved method for transforming naturally spoken language into at least one control command designed to operate a vehicle function, in particular a method by which the probability and / or the degree of impact of a misinterpretation of naturally spoken language as a voice instruction is reduced.

[0010] The problem is solved according to the invention by a method having the features of claim 1.

[0011] Advantageous embodiments of the invention are the subject of the dependent claims.

[0012] In a method for transforming naturally spoken language into at least one control command designed to operate a vehicle function, a catalog of implausible word sequences is created.

[0013] An implausible word sequence includes at least one word that is permissible as part of a valid voice instruction for operating a vehicle function. For example, in the phrase “Go to the devil!”

[0014] The sentence element "Fahr' zum" is a syntactically and semantically isolated partial word sequence of a speech instruction for formulating a service request to a navigation system.

[0015] However, the entire sequence of words as such a speech instruction is implausible, since on the one hand a destination "executioner" is typically not available and on the other hand this sequence of words represents a common idiomatic expression whose overall meaning cannot be derived from the individual elements (that is, by literal interpretation).

[0016] Similarly, the catalog of implausible word sequences can include curses, curse formulas, phrases or expressives that express an emotional state without expressing an intention to act.

[0017] Furthermore, the system continuously records naturally spoken language, for example using one or more microphones inside the vehicle. At least one current word sequence is recorded and compared to a catalog of implausible word sequences.

[0018] Preferred methods are to capture current word sequences that are candidates for valid speech instructions (that is, for which there is at least a probability of an intention to use the command). Natural language processing (NLP) methods are known and available for capturing such current word sequences; these methods are also used in conventional speech-based input methods.

[0019] A captured current word sequence (potentially assignable to a valid voice command) is compared with the implausible word sequences recorded in the catalog. If there is a match with at least one of these implausible word sequences, the captured current word sequence is eliminated (that is, it is not transformed into a voice command and therefore not into a control command).

[0020] Alternatively or additionally, the captured current word sequence can be checked. For example, if the current word sequence is ambiguous and only matches an implausible word sequence in one of its interpretations, a confirmation question can be asked. For example, several options for a possible control command can be described, and the user can be given the opportunity to choose.

[0021] The correspondence between a currently recorded word sequence and an implausible word sequence stored in the catalog does not necessarily have to be determined by a word-by-word comparison, but can be determined as a metric that describes the content-related (semantic) coverage and which is, for example, invariant with respect to elisions.

[0022] The method enables flexible speech input, especially without rigidly defined keywords and syntax rules for formulating speech instructions, and at the same time reduces the likelihood that words, phrases or sentences spoken without intention of operation will be misinterpreted as speech instructions.

[0023] This improves ease of use, accuracy, and reliability when controlling vehicle functions via natural speech input. This contributes to increased operational and road safety. Furthermore, by recognizing implausible word sequences in a timely manner, their further processing can be avoided. This reduces the need for computing power, energy, and other resources.

[0024] In one embodiment, at least one operating state parameter and / or one environmental parameter of the vehicle is included in the comparison of the at least one currently detected word sequence with the catalog of implausible word sequences.

[0025] The catalog may include word sequences that are plausible in certain situations but implausible in others. By evaluating the current operating state and the situation in the vehicle's environment, it is determined whether a situation exists in which a word sequence recorded in the catalog is implausible or not. For example, the phrase "Open all vehicle windows!"

[0026] Items are considered implausible and will be included in the catalog only if the windshield wipers are activated and / or heavy rain is detected by a rain sensor.

[0027] Similarly, phrases such as "Open the vehicle's convertible top!" or "Set audio playback volume to maximum level!" can be included in the catalog as implausible under certain conditions (moving vehicle and / or precipitation or more than one vehicle occupant).

[0028] This design improves the operational safety of vehicle functions that can be triggered and / or controlled via natural language. It contributes to the protection of vehicle occupants and people in the vehicle's vicinity, as well as to the vehicle's operational safety.

[0029] Exemplary embodiments of the invention are explained in more detail below with reference to a drawing.

[0030] This shows: Fig. 1. Schematically, a flowchart for a plausibility check of naturally spoken language before its transformation into a control command for controlling a vehicle function.

[0031] Fig. Figure 1 shows a schematic flowchart in the form of an activity diagram. Starting from a starting point S, an acoustic signal 10 is detected in a first activity A1, which was recorded, for example, with a microphone (not shown) in the vehicle interior. It is checked (for example, using signal processing and / or an artificial intelligence method) whether the acoustic signal 10 contains a potential speech input.

[0032] If no potential voice input is detected, the first decision step E1 proceeds along a "no" branch N to endpoint E. Otherwise, the process proceeds along a "yes" branch J to a second activity A2.

[0033] In the second activity A2, a current word sequence 11 is extracted from the potential speech input and passed to a subsequent third activity A3.

[0034] In the third activity, A3, the current word sequence 11 is compared with a catalog 20 of implausible word sequences. If there is sufficient literal and / or semantic match between the current word sequence 11 and one or more of the implausible word sequences recorded in catalog 20, a subsequent second decision step, E2, proceeds along the "yes" branch J to the endpoint E. Optionally, a verification activity (not shown here) can be executed to check the current word sequence 11.

[0035] If no sufficient match with at least one implausible word sequence is found, the current word sequence 11 is analyzed in a fourth activity A4 following the No branch N of the second decision step E2 according to a conventional procedure and transformed into a control command 30.

[0036] For example, the currently detected acoustic signal 10 could simply consist of murmuring or other non-verbal articulations. In that case, the procedure ends after the execution of the first activity A1 and the first decision step E1.

[0037] If the currently detected acoustic signal 10 contains an idiomatic phrase, for example the word sequence “zum Kruzifix” in the current word sequence 11 “Fahr’, zum Kruzifix” or the word sequence “zum Teufel” in the current word sequence 11 “Fahr’, zum Teufel”, then this current word sequence 11 is first identified in the second activity A2.

[0038] In the third activity, A3, a comparison with catalog 20 will reveal a sufficient match with the implausible word sequences "zum Kruzifix" and "zum Teufel" contained therein. In the subsequent second decision step, E2, the current word sequence 11 will be rejected, thus ending the process. Alternatively (not shown here), a check of the current word sequence 11 can be triggered.

[0039] If the currently detected acoustic signal 10 contains a current word sequence 11 which, although it has an overlap with an implausible word sequence from catalog 20, does not show a sufficient match, the procedure is carried out up to and including the fourth activity A4 and ends with the provision of a control command 30.

[0040] In an embodiment not shown here, a current word sequence 11 can be checked after an implausible word sequence that is at least situationally appropriate has been identified in the catalog 20 in the third activity A3.

[0041] For example, the current word sequence 11 “Please set the audio playback volume to maximum level” may be recorded as implausible in catalog 20 if other vehicle occupants besides the driver are detected.

[0042] In this case, a follow-up question might be phrased, for example, "Are you sure you want to set the audio playback to maximum level? This may be too loud for the other passengers."

[0043] Upon user confirmation, the procedure can trigger a secure execution of control command 30, which has been assigned to the current word sequence 11, for example, a gradual increase in audio playback volume. The procedure can also refuse to execute a control command 30, for example, depending on the time of day and / or an area description taken from a high-resolution (HD) map of the vehicle's surroundings. QUOTES INCLUDED IN THE DESCRIPTION

[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature

[0000] US 7,437,297 B2

[0008]

Claims

[1] Method for transforming naturally spoken language into at least one control command designed to operate a vehicle function of a vehicle (30), characterized by , that - a catalogue (20) of implausible word sequences for controlling vehicle functions is created, wherein each implausible word sequence includes at least one word that is permissible as part of a valid voice instruction for operating a vehicle function, - in naturally spoken language at least one current word sequence (11) is recorded and each is compared with the catalogue (20) and - is eliminated or checked if it matches at least one of the implausible word sequences in the catalogue (20). [2] Method according to claim 1, characterized by, that when checking a recorded current word sequence (11) which matches at least one implausible word sequence stored in the catalogue (20), at least one query and / or a corrected word sequence is generated. [3] Method according to any one of the preceding claims, characterized by , that in the comparison of the at least one recorded current word sequence (11) with the catalogue (20) at least one operating condition parameter and / or environmental parameter of the vehicle is included. [4] Method according to any one of the preceding claims, characterized by , that at least one vehicle occupant is acoustically identified and an individually assigned catalogue (20) of implausible word sequences is selected and / or continuously adapted.