Device and method for controlling vehicle functions

WO2026201348A1PCT designated stage Publication Date: 2026-10-01BAYERISCHE MOTOREN WERKE AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/052249
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2026-01-29
Publication Date
2026-10-01

Smart Images

  • Figure EP2026052249_01102026_PF_FP_ABST
    Figure EP2026052249_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a device (100) of a vehicle (102) for controlling vehicle functions, comprising a storage element (122), on which a language model (116) and a first classification layer (117) are stored, a receiving module (104), which is designed to receive spoken user input from a vehicle occupant (110) which corresponds to a vehicle function desired by the vehicle occupant (110) and comprises at least one parameter from a parameter class associated with the desired vehicle function, a processing module (106) which is designed to load and execute the language model (116) and the first classification layer (117) from the storage element (122), and a control module (108). The control module is designed to determine a vehicle function on the basis of the user input, and it is designed to determine, on the basis of the determined vehicle function and the first output, a word of the user input that is most likely to be included in the parameter class, as a parameter of the vehicle function, and to control same while taking into account at least the determined parameter of the vehicle function.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Device and method for controlling vehicle functions

[0002] The invention relates to a device for a vehicle for controlling vehicle functions. The invention further relates to a method for controlling vehicle functions.

[0003] Many vehicle functions in modern vehicles can be activated with spoken user input. This allows, for example, a driver to operate vehicle functions without taking their hands off the steering wheel or their eyes off the road. However, so-called voice assistants are limited by their language understanding and can often only correctly interpret specific, predefined commands. Variations in voice input that deviate from these predefined commands pose a challenge. Since modern vehicles have a multitude of functions, for example, 1000, the challenge of selecting and correctly activating the right function is particularly significant. Users can express their commands in very different ways, which increases the complexity of correctly classifying and assigning voice inputs to the corresponding vehicle functions.

[0004] Document WO 2020 / 226665 A1 discloses a method for the selective activation of speech recognition in a vehicle. In this method, activation occurs only when, based on recognized text, it is determined that processing is required. This can be achieved using a machine learning model.

[0005] Document WO 2021 / 076304 Al discloses a method for improving speech recognition accuracy for assistance systems. The method comprises receiving an initial audio input and processing it into multiple transcripts using several automatic speech recognition systems.

[0006] 25-0655The object of the invention is to provide a device for a vehicle and a method for controlling vehicle functions that can implement spoken user inputs better than known devices and methods.

[0007] This problem is solved by a device having the features of claim 1 and by the subject matter of the dependent claim. Further developments are specified in the dependent claims.

[0008] The proposed device for controlling vehicle functions comprises a receiver module configured to receive spoken user input from a vehicle occupant. This spoken user input corresponds to a vehicle function requested by the occupant and includes at least one parameter from a parameter class associated with that desired vehicle function. The device further includes a memory element containing a language model and a first classification layer. The language model is trained to generate an embedding for each word of the user input, based on the semantic meaning of the word within the context of the user input.The first classification layer is trained to generate an initial output for each word of the user input, based on the embeddings. This output indicates, for each parameter class from a predefined list of parameter classes, the probability that the user input word is encompassed by that parameter class. The device also includes a processing module configured to load and execute the language model and the first classification layer from the memory element. Furthermore, the device includes a control module configured to determine, based on the user input, a vehicle function that most likely corresponds to the desired vehicle function. The control module is further configured, based on the determined vehicle function and the initial output, to assign at least one word of the user input to the determined function.

[0009] The parameter class most likely to be included in the vehicle function (25-0655) is to be determined as a parameter of the vehicle function. The control module is further configured to control the vehicle function by incorporating the determined parameter of the vehicle function.

[0010] The proposed device implements a voice assistant for the vehicle that controls a vehicle function based on spoken user input, taking into account a parameter associated with that function. The user input is first processed by the language model. The output of the language model is the embedding, a mathematical representation of the semantic content of each individual word in the user input within the context of the user input's semantic content. For example, the output of the language model is a vector for each word of the user input in a high-dimensional vector space. Embeddings of words with similar semantic content will be located closer together in this high-dimensional vector space, the embedding space, than embeddings of words with different semantic content.

[0011] This allows the appropriately trained first classification layer to assign different words related to the same parameter class to that same parameter class. For example, let's say user input is used to start navigation to a destination. The user input could be: "Navigate to Grail Street in Munich," where "Grail Street" is a parameter from the parameter class<Straßenname> An alternative user input could also be: "How do I get to Unter den Linden?", where "Unter den Linden" is the parameter from the parameter class.<Straßenname> The language model now assigns an embedding to each word in the user input. The first classification layer then assigns a probability vector to each word using its respective embedding. This vector specifies the probability, for each parameter class included in the predefined list, that a word will be associated with the embedding.

[0012] The parameter class 25-0655 is included. Thus, using the first classification layer, the word "Gralstraße" or the words "Unter den Linden" would each be assigned a high probability of belonging to the parameter class.<Straßenname> to correspond, since the embeddings are located, for example, in a limited neighborhood of the embedding space, which corresponds to the parameter class<Straßenname> is assigned. The word "Munich" would be assigned a high probability according to the same principle, belonging to the parameter class <stadtname>to correspond. The remaining words would, for example, each be assigned a high probability of belonging to a parameter class. <unbekannt>to correspond if the respective words only correspond to one of the other parameter classes with a very low probability.

[0013] The control module then determines, based on the user input, for example using the embeddings generated by the language model or using other methods, a vehicle function that corresponds to the vehicle function desired by the vehicle occupant. In an exemplary embodiment, based on the user input "Navigate to Gralstraße in Munich", the control module determines the vehicle function "FN_NAVIGIERE_ZU(<Straßenname> ; <stadtname>), a function with two parameters, defined by the parameter class<Straßenname> or <stadtname>In such an embodiment, the control module, based on the first output, determines at least one word that is most likely encompassed by the parameter class belonging to the determined vehicle function, as the parameter of the determined vehicle function. In the aforementioned example, the parameter thus determined is the word "Gralstraße," which is highly likely to belong to the parameter class<Straßenname> is included. Since, in the aforementioned example, the further parameter class for the determined vehicle function is also included. <stadtname>Based on the first output, the control module additionally determines at least one further word as a parameter of the vehicle function, which is most likely to be included in the further parameter class. In the example, this is the word "Munich", which is with a high probability (25-0655) included in the parameter class. <stadtname>is included. Then the control module executes the determined vehicle function, taking into account at least the determined parameter of the vehicle function. In the example, the vehicle function FN_NAVIGIERE_ZU(<Straßenname> ; <stadtname>) with the parameters<Straßenname> = “Grail Road” and <stadtname>= “Munich” executed.

[0014] Classifying user input words into different parameter classes enables the voice assistant generated by the device to reliably determine the parameters required to control vehicle functions, even based on complex and variably worded user input, and to extract them for further processing. Furthermore, word classification can be performed in parallel with other user input processing operations. This allows for a more efficient implementation of spoken user input, particularly on limited hardware such as that typically found in vehicles, compared to known devices and methods.

[0015] In one embodiment, the first classification layer is trained using a training dataset containing sample user inputs, each word of which is labeled with one of the parameter classes from the predefined list. The language model is trained by feeding these sample user inputs into the language model. The first classification layer is trained using a supervised learning approach. During the training of the first classification layer, the language model remains unchanged. For example, one element of the training dataset consists of the user input "Open the window halfway" and the corresponding parameter class for the vehicle function FN_FENSTER_OEFFNEN_BIS( <fensterposition>) associated parameter class <fensterposition>The language model generates an embedding for each individual word in the example user input, which is then fed into the first 25-0655 classification layer as training input. The word "half" in the aforementioned example user input is, for instance, represented by the parameter class <fensterposition>The word is labeled because it specifies the target window position. The remaining words in the sample user input are labeled with the parameter class. <unbekannt>These words are labeled because they are not included in any parameter class belonging to a vehicle function. The embeddings of words with similar semantic content generated by the language model are closer together in the embedding space than the embeddings of words with different semantic content. This allows the trained classification layer to robustly assign words in user input to the corresponding parameter class, even if these words were not part of the training dataset. For example, the following sample user inputs were assigned to the parameter class <fensterposition>the vehicle function FN_WINDOWS_OPEN_UNTIL( <fensterposition>The following were used as part of the training dataset: "Open the window halfway", "Open the window completely", "Open the window halfway", and "Open the window halfway". For example, the words "half", "completely", and "halfway" were used as words of the parameter class. <fensterposition>labeled. After training, variations of these user input words, such as "window halfway up," "open window," or "please open the window halfway up," are also recognized accordingly. For example, the words "halfway up," "open," and "halfway up" are each considered highly likely to be used by the parameter class <fensterposition>The data is determined. Alternatively or additionally, a training dataset can be used which includes the corresponding output of the language model instead of the example user inputs.

[0016] In one embodiment, a second classification layer is stored on the memory element. This second classification layer is trained to generate a second output for each word of the user input, based on the embeddings. This second output indicates which of the vehicle functions controllable by the control module most likely corresponds to the desired vehicle function. The processing module is configured to load and execute the second classification layer from the memory element, and the control module is configured to determine, based on the second output of the second classification layer, the vehicle function that most likely corresponds to the desired vehicle function. In such an embodiment, the control module is further configured, for example, to control the vehicle function based on the first and second outputs, taking into account the parameter associated with the vehicle function.The close proximity of user input embeddings with similar semantic content within the embedding space allows the appropriately trained second classification layer to assign different user inputs, but those relating to the same vehicle function, to the same vehicle function. For example, the user input might be to open a vehicle window. The user input could be: "Open the window." However, the vehicle occupant could also use alternative phrases, such as: "Open the window" or "Lower the window." The classification layer then assigns the vehicle function FN_FENSTER_ÖFFNEN_BIS(. to all embeddings in this neighborhood. <fensterposition>The control module then executes the vehicle function determined by the classification layer, taking into account at least the determined parameter of the vehicle function. This enables the voice assistant created by the device to control vehicle functions based on complex and variably formulated user inputs, also taking into account the associated parameters of the vehicle functions. Furthermore, the voice assistant is able to perform the determination of the desired vehicle function and the generation of the initial output for the parameter associated with the vehicle function in parallel, thus reducing the latency between receiving the user input and controlling the vehicle function.Furthermore, the computational load of the processing module can be reduced, since the embeddings are only created once with the computationally intensive language model and can then be used in non-computationally intensive steps to determine the desired 25-0655 vehicle function and to generate the first output.

[0017] In one embodiment, the second classification layer is trained using a training dataset containing sample user inputs, each labeled with a corresponding vehicle function, and the language model itself. This training is achieved by feeding the sample user inputs to the language model. The training of the second classification layer can also be performed using a supervised learning approach. During the training of the second classification layer, the language model remains unchanged. For example, one element of the training dataset consists of the user input "Open the window" and the vehicle function FN_WINDOW_OPEN_UNTIL as its label. The language model generates an embedding for each individual word of the sample user inputs, which are then fed back to the second classification layer as input.The training enables the classification layer to robustly assign variations in user input to the corresponding vehicle function, even if these variations were not part of the training dataset. For example, the following sample user inputs were used as part of the training dataset for the vehicle function FN_FENSTER_ÖFFNEN_BIS: "Open the window" and "Open the window." After training, variations of these user inputs, such as "Window up," "Open window," and "Please open the window," will also have a high probability of being correctly assigned to the vehicle function FN_FENSTER_ÖFFNEN_BIS.

[0018] In one embodiment, the second classification layer comprises at least one neural network. Neural networks are capable of reliably recognizing complex patterns even in high-dimensional datasets, such as the embedding space. This enables the neural network to correctly assign embeddings to the same vehicle function, even if they relate to the same vehicle function but have different underlying characteristics.

[0019] 25-0655, however, the semantic content is so different that they do not lie within a geometrically easy-to-define neighborhood within the embedding space, for example, "Open the window" and "the air is very bad." Alternatively or additionally, the second classification layer can include further elements, for example, elements of a transformer architecture such as one or more attentionheads.

[0020] In one embodiment, the language model is trained using a generic text corpus. This generic text corpus comprises texts that are not limited to a specific subject area. For example, the Toronto Book Corpus or a filtered version of Wikipedia can be used as the generic text corpus. Training with the generic text corpus gives the language model a general understanding of language and a broad knowledge base. This general knowledge is also referred to as world knowledge. This world knowledge enables the language model to assign different user inputs with the same semantic content to an embedding in the same neighborhood of the embedding space, even if it was not specifically trained on these user inputs.

[0021] In one embodiment, the language model is a Large Language Model or a Small Language Model. Large Language Models (LLM) and Small Language Models (SLM) are classes of language models that differ primarily in the size of the training dataset used to train them. SLMs, in particular, can be optimized for a specific task. LLMs are typically trained with a text corpus that can be several hundred gigabytes in size. SLMs are typically trained with a text corpus that is only a few gigabytes in size. LLMs and SLMs also differ in the number of variables and thus the size of the model itself. An LLM can have one hundred billion variables; for example, GPT-3 has 175 billion variables, while an SLM typically has no more than one billion variables; for example, BERT has 340 million variables.By appropriately selecting the model size, low latency can be ensured, and the speech model can also be run on the vehicle's limited hardware. For example, BERT can be operated with an inference time of 20 ms, which is imperceptible to humans. The speech model can be either BERT or one of its many successors and enhancements, such as DistilBERT, ALBERT, roBERTa, ELECTRA, and T5.

[0022] In one embodiment, the receiving module is designed to convert the user input into a text format that can be processed by the language model.

[0023] For example, the receiving module can be trained to generate a list of words (tokens) in text form based on the spoken user input, corresponding to the user's input. The language model can be kept particularly simple if the input to the language model is in text form.

[0024] In one embodiment, the first output for each word of user input comprises an ordered list containing a numerical value for each parameter class in the predefined list, indicating the probability that the word is encompassed by that parameter class. In such an embodiment, the first classification layer resolves the embeddings by reducing each high-dimensional embedding to a vector whose dimensionality corresponds to the number of parameter classes in the predefined list. Each entry in this vector represents the probability that the word is encompassed by the respective parameter class. The control module can then, for example, identify a word that is most likely to be encompassed by a parameter class of a determined vehicle function and define it as a parameter of that function.If the parameter cannot be clearly determined, the control module can, for example, activate an output unit of the vehicle to prompt the vehicle occupant to repeat the user input.

[0025] 25-0655 In one embodiment, the second output comprises an ordered list containing a numerical value for each of the vehicle functions controllable by the control module, indicating the probability that the vehicle function corresponds to the desired vehicle function. In this embodiment as well, the second classification layer resolves the embedding(s) generated by the language model based on user input. The second classification layer reduces each high-dimensional embedding to a vector whose dimensionality corresponds to the number of controllable vehicle functions. Each entry in this vector corresponds to the probability that one of the vehicle functions is the desired vehicle function. The control module can then, for example, determine the vehicle function with the highest probability and control it.If the desired vehicle function cannot be clearly determined, the control module can, for example, activate an output unit of the vehicle to prompt the vehicle occupant to repeat the user input.

[0026] In one embodiment, the processing module is part of the vehicle.

[0027] For example, the processing module is part of a vehicle processing unit, such as a central vehicle computer. Alternatively, the processing module can be implemented, at least partially, by a processing unit located remote from the vehicle, such as a server, or in a cloud computing environment. For example, the language model is executed on a server remote from the vehicle or in a cloud computing environment. In such an embodiment, the language model is preferably stored on a memory element of the processing unit remote from the vehicle. This allows the use of a language model that might not be able to run on the limited hardware of the vehicle, or only with very high latency.The processing module 25-0655 can also be part of a mobile device that can be installed in the vehicle, for example, a smartphone or a tablet computer belonging to a vehicle occupant. In such an embodiment, the language model is preferably stored on a memory element of the mobile device.

[0028] According to another aspect, the invention relates to a method for controlling vehicle functions of a vehicle.The procedure involves at least the following steps: a) Spoken user input is received from a vehicle occupant, corresponding to a vehicle function requested by the occupant and including at least one parameter from a parameter class associated with the requested vehicle function; b) Using a language model and based on the user input, an embedding is created for each word of the user input, corresponding to the semantic meaning of the word within the context of the semantic meaning of the user input; c) Using a first classification layer and based on the embeddings, a first output is generated for each word of the user input, indicating, for each parameter class from a predefined list of parameter classes, the probability that the word of the user input is included in that parameter class.d) At least one vehicle function is determined based on the user input, which most likely corresponds to the desired vehicle function. e) Based on the determined vehicle function and the first output, at least one word from the user input, which is most likely to be included in the parameter class belonging to the determined vehicle function, is determined as a parameter of the vehicle function. f) The vehicle function is controlled by including at least the parameter of the vehicle function.

[0029] The method has the same advantages as the claimed device.

[0030] In particular, the method can be further developed with features described in this document in connection with the device. Furthermore, the claimed device can be further developed with features described in this document in connection with the method.

[0031] In one embodiment, a second classification layer is used, and based on the embeddings for each word of the user input, a second output is generated that indicates which of the vehicle functions controllable by the control module most likely corresponds to the desired vehicle function. In such an embodiment, for example, the vehicle function that most likely corresponds to the desired vehicle function is determined based on the second output. The second classification layer assigns different user inputs, but related to the same vehicle function, to the same vehicle function. This enables a voice assistant operating according to this embodiment to control vehicle functions even based on complex and variably formulated user inputs.

[0032] In one embodiment, the first classification layer and / or the second classification layer are retrained when a vehicle function changes, when a previously available vehicle function is no longer available, and / or when a new vehicle function becomes available. In this embodiment, specifically, only the training of the respective classification layer is repeated when, for example, new vehicle functions become available. This saves the considerable computational effort required for training the language model.

[0033] Exemplary embodiments of the invention are explained in more detail below with reference to the figures. These show:

[0034] Figure 1 is a schematic representation of a device of a vehicle for controlling vehicle functions according to an embodiment; and Figure 2 is a flow chart of a method for controlling vehicle functions according to an embodiment.

[0035] Figure 1 shows a schematic representation of a device 100 of a vehicle 102 for controlling vehicle functions according to one embodiment. Vehicle functions controllable by the device 100 include, for example, opening a window of the vehicle 102, starting route guidance, activating seat heating, activating ventilation, initiating a call, controlling lighting, controlling an entertainment system, controlling a vehicle mode, providing an operating aid, querying the status, and providing a support function for the vehicle 102. The device 100 comprises a receiver module 104, a processing module 106, and a control module 108, which are shown only as examples of parts of the vehicle 102.

[0036] The receiver module 104 is designed to receive spoken user input from a vehicle occupant 110. Subsequently, the user input always corresponds to a vehicle function requested by the vehicle occupant 110, i.e., a vehicle function to be executed by the vehicle 102.

[0037] Furthermore, the user input always includes at least one parameter from a parameter class belonging to the desired vehicle function. To receive the user input in spoken form, the receiver module 104 can be configured to receive the user input as audio data from a microphone, for example, a microphone 112 of the vehicle 102 or a microphone of a mobile device paired with the vehicle 102. The receiver module 104 can generate audio data from the user input or convert the spoken user input into text and make it available for further processing by the device 100.

[0038] 25-0655104 is shown purely as an example of part of a processing unit 114 of the vehicle 102, for example a central vehicle computer.

[0039] The processing module 106 is trained to operate a language model 116 and a first classification layer 117. By way of example, the processing module is further trained to operate a second classification layer 118. This means that the processing module 106 is trained to load and execute the language model 116, the first classification layer 117, and the second classification layer 118 from a memory element, for example, a memory element 122 of the processing unit 114 of the vehicle 102. The language model 116 has been trained to generate an embedding for each word of the user input, for example, based on a generic text corpus. The embeddings are, for example, vectors in a high-dimensional embedding space and correspond to the semantic meaning of the word in the context of the semantic meaning of the user input.The embeddings are further processed by the first classification layer 117, which has been trained to generate an initial output for each word of the user input based on the embeddings. This output indicates, for each parameter class from a predefined list of parameter classes, the probability that the user input word is encompassed by that parameter class. For example, the initial output could be an ordered list of numerical values ​​indicating, for each parameter class in the predefined list, the probability that the word is encompassed by that parameter class.The embeddings are further processed by the second classification layer 118, which has been trained to generate a second output for each word of user input based on the embeddings. This output indicates which of the vehicle functions controllable by the control module 108 most likely corresponds to the desired vehicle function. The second output could, for example, be an ordered list of numerical values ​​for each of the vehicle functions.

[0040] 25-0655 indicates how likely it is that this will be targeted by the user input.

[0041] The processing module 106 is also shown, purely by way of example, as part of the processing unit 114 of the vehicle 102. In other embodiments, however, the processing module 106 can also be formed wholly or partially by a processing unit located remote from the vehicle 102. In particular, the language model 116 can be executed on such a processing unit remote from the vehicle 102. In such an embodiment, the language model 116 is preferably stored on a memory element of the processing unit remote from the vehicle 102.

[0042] The control module 108 is designed to determine a vehicle function based on the user input and the second output, and to control the vehicle function using the first output and the parameter associated with that function. For example, the control module 108 determines, based on the second output, which of the vehicle functions is most likely to be activated and then activates it. If the control module 108 cannot unambiguously determine the desired vehicle function based on the second output—for example, if none of the vehicle functions has been clearly classified as the most likely—the control module 108 can, for example, activate an output unit 120 of the vehicle 102 to prompt the vehicle occupant 110 to repeat the user input.Like the receiver module 104 and the processing module 106, the control module 108 is also shown purely as an example as part of the processing unit 114 of the vehicle 102.

[0043] Figure 2 shows a flowchart of a method for controlling vehicle functions according to one embodiment. The method can

[0044] 25-0655 for example, using the device 100 according to Figure 1.

[0045] In step S200, the procedure is started. In step S202, the spoken user input from vehicle occupant 110 is received, which corresponds to a vehicle function desired by vehicle occupant 110 and includes at least one parameter from a parameter class belonging to the desired vehicle function.

[0046] With their user input, the vehicle occupant specifies which vehicle function should be activated and provides a parameter that defines how the vehicle function should be executed. For example, the vehicle occupant says "open the window halfway" or "open the window fully" if they want a window of vehicle 102 to be opened to a specific position, such as halfway or fully. In this case, the desired vehicle function corresponds, for example, to the vehicle function FN_FENSTER_ÖFFNEN_BIS( <fensterposition>) and the parameters from the vehicle function FN_FENSTER_ÖFFNEN_BIS( <fensterposition>) associated parameter class <fensterposition>They are either "half" or "complete". For example, user input is received by receiver module 104 of device 100.

[0047] In step S204, using language model 116 and based on the user input, an embedding is generated for each word of the user input. This embedding corresponds to the semantic meaning of the word within the context of the user input's semantic meaning. The embedding is a mathematical representation of the semantic content of each individual word of the user input within the context of the semantic content of all other words in the user input. The output of language model 116 is, for example, a vector for each word in a high-dimensional vector space. The embeddings of user inputs with the same

[0048] 25-0655 or similar meanings have a smaller spacing in this vector space than the embeddings of user inputs with different meanings.

[0049] This allows for a mathematical definition of "similar meaning." For example, the embedding of the words "half" and "complete" in the user inputs "open the window halfway" and "open the window completely" would have a small difference in meaning, since both words describe a window position in the context of the user input and therefore have a similar meaning. Step S204 is performed, for example, by the processing module 106. Training the language model 116 can be performed as an optional step within the procedure. Alternatively, a previously trained language model 116 can be used.

[0050] In step S206, using the first classification layer 117 and based on the embeddings, an initial output is generated for each word of the user input. This initial output indicates, for each parameter class from a predefined list of parameter classes, the probability that the user input word is encompassed by that parameter class. For example, the first classification layer 117 generates the ordered list of numeric values ​​from the high-dimensional vector in the embedding space—that is, another vector with a significantly smaller dimension. Thus, the first classification layer 117 assigns a parameter class from the predefined list of parameter classes to the meaning of each word, as determined by the language model 116 within the context of the user input meaning. The predefined list of parameters includes, for example, the parameter class... <unbekannt>, which includes the words in the user input that are not covered by any parameter class belonging to a vehicle function. Referring to the example user inputs "open the window halfway" and "open the window completely," the first output would indicate that the words "half" and "completely" are highly likely to be covered by the parameter class <fensterposition>are included. Furthermore, the first output would, for example, indicate that the remaining words of the user input from 25-0655 belong to the parameter class. <unbekannt>are included. Step S206, like step S204, is performed, for example, by processing module 106.

[0051] The first classification layer 117 has been trained to generate the first output based on the embeddings for each word of the user input. This output indicates, for each parameter class from a predefined list of parameter classes, the probability that the user input word is encompassed by that parameter class. For example, to train the first classification layer 117, a training dataset is created containing sample user inputs, each word of which is labeled with one of the parameter classes from the predefined list. This training dataset is then fed into the language model 116 as input to generate an embedding for each word of the sample user input, which is then fed back into the first classification layer 117 as training input.The variables of the first classification layer 117 are varied until the first output indicates the correct parameter classes to which the words of the example user input correspond. The variables of the language model 116 remain unchanged. This training can be performed as an optional step within the procedure. In particular, the training of the first classification layer 117 can be repeated as part of the procedure if, for example, new vehicle functions become available in the vehicle 102.

[0052] In step S208, at least one vehicle function is determined based on the user input, which most likely corresponds to the desired vehicle function. This is done, for example, using the embedding for each word and the second classification layer 118 described in relation to Figure 1. Alternatively, the vehicle function can also be determined based on the user input using rule-based models or statistical models that do not use SLMs or LLMs. Furthermore, in step S208, based on the determined vehicle function and the first output, at least one word from the user input is used that corresponds to the determined vehicle function.

[0053] The parameter most likely belonging to the parameter class 25-0655 is determined as a parameter of the vehicle function. The determined vehicle function is then controlled by incorporating at least the determined parameter of the vehicle function. Step S208 is, for example, carried out by control module 108. The procedure is then completed in step S210.

[0054] In the embodiments described with reference to Figures 1 and 2, at least the receiving module 104, the processing module 106, and the control module 108 form the device 100 of a vehicle 102 for controlling vehicle functions. Further elements and features shown in Figures 1 and 2 and mentioned in the preceding description may be part of the device 100. Likewise, method steps described with reference to the device 100 may be part of the claimed method.

[0055] 25-0655Reference symbol list

[0056] 100 Device

[0057] 102 vehicles

[0058] 104 Receiving module 106 Processing module 108 Control module

[0059] 110 vehicle occupants 112 microphone

[0060] 114 Processing unit 116 Language model 117, 118 Classification layer 120 Output unit 122 Storage element

[0061] 25-0655< / unbekannt> < / fensterposition> < / unbekannt> < / fensterposition> < / fensterposition> < / fensterposition> < / fensterposition> < / fensterposition> < / fensterposition> < / fensterposition> < / fensterposition> < / unbekannt> < / fensterposition> < / fensterposition> < / fensterposition> < / stadtname> < / stadtname> < / stadtname> < / stadtname> < / stadtname> < / stadtname> < / unbekannt> < / stadtname>

Claims

Claims 1. Device (100) of a vehicle for controlling vehicle functions comprising a receiving module (104) configured to receive spoken user input from a vehicle occupant (110) that corresponds to a vehicle function desired by the vehicle occupant (110) and includes at least one parameter from a parameter class belonging to the desired vehicle function, a storage element (122) on which a language model (116) and a first classification layer (117) are stored, wherein the language model (116) is trained to generate, based on the user input, for each word of the user input, an embedding that corresponds to the semantic meaning of the word in the context of the semantic meaning of the user input, and wherein the first classification layer (117) is trained to generate, based on the embeddings for each word of the user input, a first output that indicates, for each parameter class of a predefined list of parameter classes, the probability with which the word of the user input is covered by the parameter class, a processing module (106) that is trained to load and execute the language model (116) and the first classification layer (117) from the storage element (122), and a control module (108) which is configured to determine, based on the user input, a vehicle function which most likely corresponds to the desired vehicle function, and which is configured, based on the determined vehicle function and the first output, to determine at least one word of the 25-0655 user input which is most likely to be included in the parameter class associated with the determined vehicle function, as a parameter of the vehicle function, and which is configured to control the vehicle function taking into account the determined parameter of the vehicle function.

2. Device (100) according to claim 1, wherein the first classification layer (117) has been trained using a training data set comprising exemplary user inputs, the words of which are each labelled with one of the parameter classes of the predefined list, and using the language model (116) by inputting the exemplary user inputs to the language model (116).

3. Device according to claim 1 or 2, wherein a second classification layer (118) is stored on the storage element (122), which is trained to generate a second output based on the embeddings for each word of the user input, indicating which of the vehicle functions controllable by the control module (108) most likely corresponds to the desired vehicle function, wherein the processing module (106) is configured to load and execute the second classification layer (118) from the storage element (122), wherein the control module (108) is configured to determine, based on the second output, the vehicle function which is most likely to correspond to the desired vehicle function.

4. Device (100) according to claim 3, wherein the second output comprises an ordered list which includes, for each of the vehicle functions controllable by the control module (108), a numerical value indicating the probability that the vehicle function corresponds to the desired vehicle function.

5. Device (100) according to one of claims 3 or 4, wherein the second classification layer (118) has been trained using a training data set comprising exemplary user inputs, each labelled with a corresponding vehicle function, and using the language model (116) by inputting the exemplary user inputs to the language model (116).

6. Device (100) according to one of claims 3 to 5, wherein the second classification layer (118) comprises at least one neural network.

7. Device (100) according to one of the preceding claims, wherein the language model (116) has been trained using a generic text corpus.

8. Device (100) according to any one of the preceding claims, wherein the language model (116) is a Large Language Model or a Small Language Model.

9. Device (100) according to one of the preceding claims, wherein the receiving module (104) is configured to convert the user input into a text form that can be processed by the language model (116).

10. Device (100) according to any one of the preceding claims, wherein the first output for each word of the user input comprises an ordered list which includes, for each of the parameter classes of the predefined list, a numerical value indicating the probability with which the word is included in the parameter class.

11. Device (100) according to any one of the preceding claims, wherein the processing module (106) is part of the vehicle (102).

12. Device (100) according to any one of claims 1 to 10, wherein the processing module (106) is part of a processing unit located away from the vehicle (102).

13. Method for controlling vehicle functions of a vehicle (102), wherein a) a spoken user input is received from a vehicle occupant (110) which corresponds to a vehicle function desired by the vehicle occupant (110) and includes at least one parameter from a parameter class belonging to the desired vehicle function; b) using a language model (116) and based on the user input, an embedding is created for each word of the user input that corresponds to the semantic meaning of the word in the context of the semantic meaning of the user input; c) using a first classification layer (117) based on the embeddings for each word of user input, a first output is generated which indicates for each parameter class from a predefined list of parameter classes the probability with which the word of user input is covered by the parameter class; 25-0655d) at least on the basis of user input, a vehicle function is determined which is most likely to correspond to the desired vehicle function; e) based on the determined vehicle function and the first output, at least one word of the user input, which is most likely to be included in the parameter class belonging to the determined vehicle function, is determined as a parameter of the vehicle function; f) the vehicle function is controlled by taking into account at least the determined parameter of the vehicle function.

14. The method of claim 13, wherein, using a second classification layer (118) and based on the embeddings for each word of the user input, a second output is generated indicating which of the vehicle functions controllable by the control module (108) most likely corresponds to the desired vehicle function; and where, based on the second output, the vehicle function is determined that most likely corresponds to the desired vehicle function.

15. Method according to claim 13 or 14, wherein the first classification layer (117) and / or the second classification layer (118) are retrained when a vehicle function has changed, when a previously available vehicle function is no longer available and / or when a new vehicle function is available. 25-0655