Device and method for controlling vehicle functions

WO2026201350A1PCT designated stage Publication Date: 2026-10-01BAYERISCHE MOTOREN WERKE AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/052252
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2026-01-29
Publication Date
2026-10-01

Smart Images

  • Figure EP2026052252_01102026_PF_FP_ABST
    Figure EP2026052252_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a device (100) of a vehicle (102) for controlling vehicle functions, comprising a receiving module (104) which is designed to receive spoken user input from a vehicle occupant (112), which corresponds to a vehicle function desired by the vehicle occupant (112) and comprises at least one parameter. The device (100) also comprises a storage element (106) on which a tokenizer (118), a language model (120) and at least one first classification layer (122a) are stored. The tokenizer (118) is designed to carry out a segmentation of the user input into tokens. The language model (120) is trained such that, based on the segmented user input, it generates an embedding for each token of the user input, which corresponds to the semantic meaning of the token in the context of the semantic meaning of the user input. The first classification layer (122a) is trained such that, on the basis of the embeddings for each token of the user input, it generates a first output which, for each parameter characteristic in a predefined list of parameter characteristics, specifies the probability that the token corresponds to that parameter characteristic. A processing module (108) comprised by the device (100) is designed to load and execute the tokenizer (118), the language model (120) and the first classification layer (122a) from the storage element (106). The device (100) also comprises a control module (110) which is designed to determine, on the basis of the user input, a vehicle function that most likely corresponds to the desired vehicle function, which is designed to determine, on the basis of the user input, whether a token corresponds to a parameter of the determined vehicle function, which is designed to determine a parameter characteristic of the parameter for tokens which correspond to a parameter of the determined vehicle function on the basis of the first output, and which is designed to control the determined vehicle function while taking into account the determined parameter characteristic.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Device and method for controlling vehicle functions

[0002] The invention relates to a device for a vehicle for controlling vehicle functions. The invention further relates to a method for controlling vehicle functions.

[0003] Many vehicle functions in modern vehicles can be activated with spoken user input. This allows, for example, a driver to operate vehicle functions without taking their hands off the steering wheel or their eyes off the road. However, so-called voice assistants are limited by their language understanding and can often only correctly interpret specific, predefined commands. Variations in voice input that deviate from these predefined commands pose a challenge. Since modern vehicles have a multitude of functions, for example, 1000, the challenge of selecting and correctly activating the right function is particularly significant. Users can express their commands in very different ways, which increases the complexity of correctly classifying and assigning voice inputs to the corresponding vehicle functions.

[0004] From WO 2021 / 086645 Al, a method is known in which an utterance in natural language is received as user input. A language model is used for processing the user input.

[0005] The object of the invention is to provide a device for a vehicle and a method for controlling vehicle functions, including a parameter configuration that can implement spoken user input better than known devices and methods.

[0006] 25-0653 This problem is solved by a device having the features of claim 1 and by the subject matter of the dependent claim. Further developments are specified in the dependent claims.

[0007] The proposed device for controlling vehicle functions comprises a receiver module configured to receive spoken user input from a vehicle occupant, corresponding to a vehicle function requested by the occupant and including at least one parameter. The device also includes a memory element on which a tokenizer, a language model, and at least one first classification layer are stored. The tokenizer is configured to segment the user input into tokens. The language model is trained to generate an embedding for each token of the user input, based on the segmented user input, that corresponds to the semantic meaning of the token within the context of the semantic meaning of the user input.The first classification layer is trained to generate an initial output for each user input token, based on the embeddings. This output indicates, for each parameter value in a predefined list of parameter values, the probability that the token corresponds to that parameter value. A processing module included in the device is trained to load and execute the tokenizer, the language model, and the first classification layer from the storage element.The device further comprises a control module configured to determine, based on user input, a vehicle function that most likely corresponds to the desired vehicle function, configured to determine, based on user input, whether a token corresponds to a parameter of the determined vehicle function, configured to determine, for tokens corresponding to a parameter of the determined vehicle function, a parameter value of the parameter based on the first output, and configured to control the determined vehicle function taking into account the determined parameter value.

[0008] 25-0653 The proposed device implements a voice assistant for the vehicle, which controls a vehicle function based on spoken user input. The user input is intended to control a vehicle function that includes a parameter from a parameter class with a finite number of parameter values. For example, the driver's side window is to be opened. The user input could be: "Open the front left window" or "Open the driver's window." The corresponding vehicle function is, for example, FN_OPEN_WINDOW and has a parameter <position>, which specifies the position of the window to be opened. The words "front left" and "driver's window" in the example user inputs each refer to a specific value of the parameter. <position>, which is determined by the device from the user input.

[0009] First, the user input is segmented into different tokens using the tokenizer. These tokens can include, for example, a word, part of a word (such as a prefix), or punctuation marks. The segmented user input is then processed by the language model. The output of the language model consists of embeddings of the tokens, i.e., mathematical representations of the semantic content of the respective tokens, for example, as vectors in a high-dimensional vector space, the embedding space. Based on these embeddings, the control module then determines, based on the segmented user input, which vehicle function most likely corresponds to the vehicle function desired by the vehicle occupant. Furthermore, the embeddings enable the appropriately trained first classification layer to assign different words related to the same parameter values ​​to the same parameter values.In the aforementioned example, the first classification layer assigns a value of the parameter to the words "front left" and "driver's window" respectively. <position>25-0653zu. For this purpose, the first classification layer determines, for example, a probability vector that specifies, for each possible parameter value, the probability with which a token corresponds to that parameter value. The predefined list of parameter values ​​of the parameter <position>includes, for example, the values<vorne links> ,<vorne rechts> ,<hinten links> and<hinten rechts> In the aforementioned example, the first classification layer assigns a high probability to the words "front left" and "driver's window" respectively, according to the parameter value.<vorne lin ks> to correspond, since the embeddings are located, for example, in a limited neighborhood of the embedding space that corresponds to the parameter value<vorne lin ks> is assigned. The classification of the parameters according to parameter values ​​by the specialized first classification layer enables the voice assistant formed by the device to reliably determine parameter values ​​required to control a vehicle function, even on the basis of complex and variably formulated user inputs, and to control the determined vehicle function taking the parameter value into account.

[0010] In one embodiment, the first classification layer is implemented as an attention layer of the language model or as a machine learning model that can be operated independently of the language model. The attention layer is a layer of a neural network, for example, the language model. Attention layers are capable of reliably recognizing complex patterns even in high-dimensional datasets, such as the embedding space. The aforementioned models enable the first classification layer to correctly assign complex variations in user input to parameter values ​​that are so different that they do not lie within a geometrically simple neighborhood within the embedding space.

[0011] 25-0653 In one embodiment, the first classification layer is trained using a first training dataset comprising sample user inputs whose words corresponding to a parameter are labeled with the corresponding parameter value, and using the language model by feeding the sample user inputs to the tokenizer. The first classification layer is trained using a supervised learning approach. During the training of the first classification layer, the language model remains unchanged. For example, one element of the training dataset consists of a sample user input comprising the words "front left" and a parameter value, such as the parameter value...<vorne links> from the list of all possible window positions. The parameter value is used as a label.This training enables the first classification layer to robustly assign tokens of segmented user input to a parameter value, even if these were not part of the training dataset. Alternatively or additionally, a training dataset can be used that includes the corresponding output of the language model instead of the example user inputs.

[0012] In one embodiment, several first classification layers are stored on the memory element, each assigned to one of the parameter classes. The processing module is configured to load and execute the first classification layers from the memory element. The control module is configured to select one of the first classification layers for each token corresponding to a parameter, based on the second output, and to control the processing module to operate the selected first classification layer using the token as input. In this embodiment, one of the first classification layers is used as a separate, specific classifier for each parameter class. This has the advantage that each first classification layer can be trained on the specific characteristics of the respective parameter class. This reduces the complexity of the first classification layers.

[0013] 25-0653, which can ensure low latency, especially on limited vehicle hardware.

[0014] In one embodiment, a second classification layer is stored on the memory element. This layer is trained to generate a second output for each user input token based on its embeddings. This second output indicates, for each parameter class from a predefined list of parameter classes, each associated with a vehicle function controllable by the control unit, the probability that the token belongs to that parameter class. The processing module is configured to load and execute the second classification layer from the memory element. The control module is configured to determine, based on the second output, whether each user input token is a parameter of the identified vehicle function. The second classification layer classifies the tokens according to their parameter class affiliation. For example, the words "front left" are classified as belonging to the parameter class... <position>The parameters of the functions FN_OPEN_WINDOW, which opens a vehicle window, and FN_CLOSE_WININDOW, which closes a vehicle window, are classified as belonging to the same parameter class. When classifying tokens, the second classification layer leverages the fact that token embeddings belonging to the same parameter class reside in the same part of the embedding space. This allows the second classification layer to determine the parameter classes particularly easily and reliably based on the embeddings and generate the corresponding secondary outputs.

[0015] In one embodiment, the second classification layer is implemented as an attention layer of the language model or as a machine learning model that can be operated independently of the language model. The aforementioned models enable the second classification layer to correctly assign complex variations in user input to a parameter class, even if these variations are so different that they do not lie within a geometrically simple neighborhood within the embedding space.

[0016] In one embodiment, the second classification layer is trained using a second training dataset containing sample user inputs. Words corresponding to a parameter in this dataset are labeled with the appropriate parameter classes from a predefined list. This training is performed using the language model by feeding the sample user inputs to the tokenizer. The training of the second classification layer can also be performed using a supervised learning approach. During the training of the second classification layer, the language model remains unchanged. For example, one element of the training dataset consists of a sample user input containing the words "front left" and a parameter class, such as parameter class [parameter class name missing in original text]. <position>, encompassing all possible window positions. The parameter value is used as a label. This training enables the second classification layer to robustly assign tokens of the segmented user input to a parameter class, even if these were not part of the training dataset. Alternatively or additionally, a training dataset can be used that includes the corresponding output of the language model instead of the sample user input.

[0017] In one embodiment, a third classification layer is stored on the memory element. This layer is trained to generate a third output based on the embeddings, indicating which of the vehicle functions controllable by the control module most likely corresponds to the desired vehicle function. The processing module is configured to load and execute the third classification layer from the memory element. The control module is configured to determine, based on the third output, the vehicle function that most likely corresponds to the desired vehicle function. The embeddings 25-0653 of tokens with similar semantic content, generated by the language model, are located closer together in the embedding space than the embeddings of tokens with different semantic content.This allows the appropriately trained third classification layer to assign different user inputs, but related to the same vehicle function, to the same vehicle function. For example, the user input might be to open a vehicle window. The user input could be: "Open the window." However, the vehicle occupant could also use alternative phrases, such as: "Open the window" or "Lower the window." The language model then assigns embeddings to all these user inputs, all of which lie within a limited neighborhood of the embedding space. The third classification layer then assigns the vehicle function "Open window" to all embeddings in this neighborhood. The control module then executes the vehicle function determined by the third classification layer.This means that the voice assistant formed by the device is able to control vehicle functions of the vehicle even on the basis of complex and variably formulated user inputs.

[0018] In one embodiment, the third classification layer comprises at least one neural network. Neural networks are capable of reliably recognizing complex patterns even in high-dimensional datasets, such as the embedding space. This enables the neural network to correctly assign embeddings to the same vehicle function, even if they relate to the same vehicle function but have such different semantic content that they do not lie within a geometrically simple neighborhood within the embedding space, for example, "Open the window" and "The air is very bad." Alternatively or additionally, the third classification layer can include further elements, such as elements of a transformer architecture like one or more attentionheads.

[0019] 25-0653 In one embodiment, the output of the third classification layer comprises an ordered list containing, for each of the vehicle functions controllable by the control module, a numerical value indicating the probability that the vehicle function corresponds to one of the desired vehicle functions. In such an embodiment, the third classification layer resolves the embeddings by reducing the high-dimensional embeddings to a vector whose dimensionality corresponds to the number of controllable vehicle functions. Each entry in such a vector corresponds to the probability that one of the vehicle functions is one of the desired vehicle functions. If one or more desired vehicle functions cannot be uniquely determined, the control module can, for example, activate an output unit of the vehicle to prompt the vehicle occupant to repeat the user input.

[0020] In one embodiment, the third classification layer is trained using a third training dataset comprising exemplary user inputs, each labeled with a corresponding vehicle function, and using the language model. This training is achieved by feeding the exemplary user inputs to the tokenizer. Each exemplary user input is labeled with one or more vehicle functions that are to be controlled by the exemplary user input. The training of the third classification layer can also be performed using a supervised learning approach.

[0021] During the training of the third classification layer, the language model remains unchanged. For example, one element of the training dataset consists of the user input "Open the window" and the vehicle function FN_OPEN_WINDOW as its label. The tokenizer segments the example user inputs and feeds them to the language model. The language model generates embeds for each segmented user input, which are then fed to the third classification layer as training input. The labeled output is the vehicle function or functions corresponding to the respective

[0022] 25-0653 corresponds to exemplary user input. This training enables the third classification layer to robustly assign variations in user input to their respective vehicle functions, even if these were not part of the training dataset. For example, the following exemplary user inputs for the vehicle function FN_OPEN_WINDOW were used as part of the training dataset: "Open the window" and "Open the window". After training, variations of these user inputs, such as "Window down", "Open window", "Please open the window", will also have a high probability of assignment to the vehicle function FN_OPEN_WINDOW.

[0023] In one embodiment, the language model is designed as an encoder-decoder model. Encoder-decoder models, such as BERT, are particularly good at capturing complex patterns in input sequences and converting them into meaningful output. They can process the entire input before generating an ordered and coherent output, which is especially advantageous for capturing user input.

[0024] In one embodiment, the language model is a Large Language Model or a Small Language Model. Large Language Models (LLM) and Small Language Models (SLM) are classes of language models that differ primarily in the size of the training dataset used to train them. SLMs, in particular, can be optimized for a specific task. LLMs are typically trained with a text corpus that can be several hundred gigabytes in size. SLMs are typically trained with a text corpus that is only a few gigabytes in size. LLMs and SLMs also differ in the number of variables and thus the size of the model itself. An LLM can have one hundred billion variables; for example, GPT-3 has 175 billion variables, while an SLM typically has no more than one billion variables; for example, BERT has 340 million variables.By appropriately selecting the model size, low latency can be ensured, and the 25-0653 speech model can also be run on the vehicle's limited hardware. For example, BERT can be operated with an inference time of 20 ms, which is imperceptible to humans. The speech model can be either BERT or one of its many successors and enhancements, such as DistilBERT, ALBERT, roBERTa, ELECTRA, and T5.

[0025] In one embodiment, the language model has been trained using at least a generic text corpus. This generic text corpus comprises texts that are not limited to a specific subject area. For example, the Toronto Book Corpus or a filtered version of Wikipedia can be used as the generic text corpus. Training with the generic text corpus gives the language model a general understanding of language and a broad knowledge base. This general knowledge is also referred to as world knowledge. This world knowledge enables the language model to assign different tokens with the same semantic content to an embedding in the same neighborhood of the embedding space, even if it has not been specifically trained on those tokens.

[0026] In one embodiment, the receiving module is configured to convert the user input into a text format that can be processed by the tokenizer. For example, the receiving module can be configured to generate a string based on the spoken user input, which corresponds to the user input. The tokenizer can be kept particularly simple if the input to the tokenizer is in text format.

[0027] In one embodiment, the processing module is part of the vehicle or part of a processing unit located remote from the vehicle. For example, the processing module is part of a processing unit within the vehicle, such as a central vehicle computer. Alternatively, the processing module can be implemented, at least partially, by a processing unit located remote from the vehicle, such as a server located remote from the vehicle, or in a cloud computing environment. For example, the language model is executed on a server located remote from the vehicle or in a cloud computing environment. In such an embodiment, the language model is preferably stored on a memory element of the processing unit located remote from the vehicle. This allows the use of a language model that might not be able to run on the limited hardware of the vehicle, or only with very high latency.The processing module can also be part of a mobile device that can be installed in the vehicle, for example, a smartphone or a tablet computer belonging to a vehicle occupant. In such an embodiment, the language model is preferably stored on a memory element of the mobile device.

[0028] The invention further relates to a method for controlling vehicle functions of a vehicle. The method includes at least the following steps: a) A spoken user input is received from a vehicle occupant, corresponding to a vehicle function desired by the occupant and comprising at least one parameter; b) The user input is segmented into tokens; c) Using a language model and based on the segmented user input, an embedding is created for each token of the user input, corresponding to the semantic meaning of the word within the context of the semantic meaning of the user input; d) Based on the user input, a vehicle function is determined that most likely corresponds to the desired vehicle function; e) Based on the user input, it is determined whether a token corresponds to a parameter of the determined vehicle function.f) Using a first classification layer and based on the embeddings for each token corresponding to a parameter of the determined vehicle function, an initial output is generated that indicates, for each parameter value from a predefined list of 25-0653 parameter values, the probability that the token corresponds to the parameter value. g) The determined vehicle function is controlled using the determined parameter value.

[0029] The method has the same advantages as the claimed device.

[0030] In particular, the method can be further developed with features described in this document in connection with the device. Furthermore, the claimed device can be further developed with features described in this document in connection with the method.

[0031] In one embodiment, a second classification layer is used, and based on the embeddings, a second output is generated for each user input token. This output indicates, for each parameter class from a predefined list of parameter classes, each assigned to a vehicle function controllable by the control unit, the probability that the token belongs to that parameter class. Based on this second output, it is determined for each user input token whether the word is a parameter of the identified vehicle function. The second classification layer allows differently worded user inputs, but all referring to the same parameter class (e.g., window position), to be categorized within the same parameter class.This makes it possible for a voice assistant operated according to this embodiment to control vehicle functions of the vehicle even on the basis of complex and variably formulated user inputs.

[0032] In one embodiment, a third output is generated using a third classification layer and based on the embeddings. This third output indicates which of the vehicle functions controllable by the control module most likely corresponds to the desired vehicle function. Based on this third output, the vehicle function that most closely matches the desired vehicle function is determined.

[0033] 25-0653 most likely corresponds to this. The third classification layer assigns different user inputs, but relating to the same vehicle function, to the same vehicle function. This enables a voice assistant operating according to this embodiment to control vehicle functions even based on complex and variably formulated user inputs.

[0034] In one embodiment, the first, second, and / or third classification layer are retrained when a vehicle function changes, when a previously available vehicle function is no longer available, and / or when a new vehicle function becomes available. In particular, in such an embodiment, the language model is not retrained. This saves the considerable computational effort required for training the language model.

[0035] Exemplary embodiments of the invention are explained in more detail below with reference to the figures. These show:

[0036] Figure 1 shows a schematic representation of a device of a vehicle for controlling vehicle functions according to one embodiment; and

[0037] Figure 2 shows a flowchart of a method for controlling vehicle functions according to one embodiment.

[0038] Figure 1 shows a schematic representation of a device 100 of a vehicle 102 for controlling vehicle functions according to an exemplary embodiment. The device 100 implements a voice assistant for the vehicle 102, which controls vehicle functions of the vehicle 102 based on the spoken user input. In particular, the user input comprises a vehicle function with a parameter from a parameter class that includes a finite number of

[0039] The device 100 has parameter values ​​(25-0653). Vehicle functions controllable by the device 100 include, for example, opening a window at a specific position, activating seat heating for a specific seat, activating ventilation at a specific position, initiating a call for a contact from a list of contacts, lighting control, control of an entertainment system of the vehicle 102, and control of a vehicle mode. The device 100 comprises a receiver module 104, a storage element 106, a processing module, and a control module 110, which are shown only as examples of parts of the vehicle 102.

[0040] The receiver module 104 is configured to receive spoken user input from a vehicle occupant 112. The user input comprises a vehicle function requested by the vehicle occupant 112, i.e., a vehicle function to be executed by the vehicle 102, and at least one parameter. The parameter can only assume a finite number of values, for example, one of four window positions.<vorne links> ,<vorne rechts> ,<hinten lin ks> and<hinten rechts> In order to receive user input in spoken form, the receiving module 104 can be configured to receive user input in the form of audio data from a microphone, for example a microphone 114 of the vehicle 102 or a microphone of a mobile device paired with the vehicle 102.From the user input, the receiver module 104 can generate audio data or convert the spoken user input into text and make it available for further processing by the device 100. The receiver module 104 is shown purely as an example of part of a processing unit 116 of the vehicle 102, for example, a central vehicle computer.

[0041] Memory element 106 is implemented purely as an example of a memory element 106 of the processing unit 116 of the vehicle 102. A tokenizer 118, a language model 120, and a first classification layer 25-0653122a are stored on memory element 106. Also purely as examples, a second classification layer 122b and a third classification layer 122c are stored on memory element 106. Tokenizer 118 is configured to generate tokens based on user input by segmenting the user input into semantic sections, such as words and punctuation marks. Tokenizer 118 can also generate a token for individual word parts, such as prefixes. Language model 120 has been trained, for example, on a generic text corpus, to generate embeddings from the segmented user input.The embedding of a token is, for example, a vector in a high-dimensional embedding space and corresponds to the semantic meaning of the token in the context of the user input. The first classification layer 122a and the second classification layer 122b are trained to classify tokens corresponding to parameters of the user input based on their embeddings. The first classification layer 122a determines which parameter value the token most likely corresponds to and generates a corresponding first output. The second classification layer 122b determines which parameter class the token most likely corresponds to and generates a corresponding second output. The parameter class specifies which parameter values ​​are possible. Each parameter of a vehicle function is encompassed by a parameter class.Parameters of different vehicle functions can belong to the same parameter class, while each parameter is always assigned to one vehicle function. For example, there is a parameter class that specifies all possible window positions. The parameters of the functions FN_OPEN_WINDOW, which opens a window of vehicle 102, and FN_CLOSE_WININDOW, which closes a window of vehicle 102, can each belong to this parameter class. The third classification layer, 122c, is trained to determine, based on the embeddings, which vehicle function is most likely to be executed with the user output and to generate a corresponding third output.

[0042] 25-0653 The processing module 108 is configured to operate the tokenizer 118, the language model 120, the first classification layer 122a, the second classification layer 122b, and the third classification layer 122c. This means that the processing module 108 is configured to load and execute the tokenizer 118, the language model 120, the first classification layer 122a, the second classification layer 122b, and the third classification layer 122c from the memory element 106. The processing module 108 is also shown, purely by way of example, as part of the processing unit 116 of the vehicle 102. In other embodiments, however, the processing module 108 can also be formed wholly or partially by a processing unit located away from the vehicle 102.In particular, the language model 120 can be executed on such a processing unit remote from the vehicle 102, for example a backend server or a processing unit implemented in a cloud computing environment. In such an embodiment, the language model 120 is preferably stored on a memory element 106 of the processing unit remote from the vehicle 102.

[0043] Based on the segmented user input, control module 110 determines which of the executable vehicle functions should be activated. For example, control module 110 uses the third output of the third classification layer 122c for this purpose. Furthermore, control module 110 determines for each token of the segmented user input whether it corresponds to a parameter of the identified vehicle function. For this, control module 110 uses, for example, the second outputs of the second classification layer 122b. Then, based on the first outputs, control module 110 determines the parameter value from a list of parameter values ​​for each token identified as a parameter. Finally, control module 110 activates the identified vehicle function, taking the determined parameter values ​​into account.Can the control module 110 not determine a parameter value without doubt, or can the control module 110 not determine which one?

[0044] 25-0653Vehicle functions If the desired vehicle function is, the control module 110 can, for example, control an output unit 124 of the vehicle 102 to prompt the vehicle occupant 112 to repeat the user input.

[0045] Figure 2 shows a flowchart of a method for controlling vehicle functions according to one embodiment. The method implements voice control for the vehicle functions of vehicle 102. The method is described purely by way of example with reference to the device 100 according to Figure 1.

[0046] In step S200, the procedure is started. In step S202, the spoken user input from vehicle occupant 112 is received, which corresponds to the vehicle function desired by vehicle occupant 112 and at least one parameter. With this user input, vehicle occupant 112 specifies which vehicle function should be activated. For example, vehicle occupant 112 says "Open the window on the driver's side!" to activate the function FN_OPEN_WINDOW with the parameter<vorne lin ks> to control. Optionally, in this step, a text format can be generated from the spoken user input, which is available, for example, in the form of audio data, making it easier to process further. For example, using a trained audio-to-text model. The user input is received and processed, for example, by the microphone of the receiver module 104 of the device 100.

[0047] In step S204, the user input is segmented into tokens. For example, a corresponding token is generated for each word in the user input. Tokens can also be generated for punctuation marks and / or word parts, such as prefixes. The result of the segmentation is, for example, an ordered list of strings, each corresponding to a word, word part, or punctuation mark, or an ordered list of numeric identifiers. For example, the user input "Open the window on the 25-0653 driver's side!" becomes the list ["Open", "the", "window", "on", "the", "driver's side"]. The user input segmentation is generated, for example, by processing module 108 using tokenizer 118. The segmented user input is processed in step S206 using language model 120. An embed is created for each token.The embeddings are each a mathematical representation, for example a high-dimensional vector in an embedding space, that corresponds to the semantic meaning of the token in the context of the user input. Step S206 is performed, for example, by the processing unit 116 running the language model 120 to generate the embeddings.

[0048] Language Model 120, for example, is an encoder-decoder model or another suitable machine learning model trained to generate embeddings from tokens, for instance, based on a generic text corpus. Training Language Model 120 can be performed as an optional step within the process. Alternatively, a pre-trained Language Model 120 can be used.

[0049] In step S208, the vehicle function that most likely corresponds to the desired vehicle function is determined based on the user input.

[0050] For example, for each of the controllable vehicle functions, the probability is determined that this vehicle function corresponds to one of the desired vehicle functions. Determining the desired vehicle function is done, for example, using the third classification layer 122c. For instance, the third classification layer 122c is trained to generate an ordered list of numerical values ​​from the embeddings, indicating for each controllable vehicle function the probability that it should be controlled by the user input. Alternatively, the third classification layer 122c can also output the most probable vehicle functions, such as the two, three, or ten most probable vehicle functions.

[0051] 25-0653 The third classification layer 122c generates a further vector with a significantly smaller dimension, the ordered list, from the high-dimensional vectors in the embedding space. Thus, the third classification layer 122c assigns concrete vehicle functions, likely intended to be controlled, to the meaning of the user input determined by the language model 120.

[0052] For training the third classification layer 122c, a training dataset is generated containing sample user inputs, each labeled with corresponding vehicle functions. This training dataset is first processed by the tokenizer 118 to segment the sample user inputs. The segmented sample user inputs are then fed into the language model 120 to generate embeddings for each sample user input. These embeddings are then fed back into the third classification layer 122c as training input. The parameters of the third classification layer 122c are varied until its output corresponds to the correct vehicle functions. The parameters of the language model 120 remain unchanged. This training can be performed as an optional step within the overall process.

[0053] In particular, the training of the third classification layer 122c can be repeated as part of the procedure if, for example, new vehicle functions are available in the vehicle 102.

[0054] In step S210, it is determined for each token whether it corresponds to a parameter of the previously determined vehicle function. This can be done, in particular, using the second outputs of the second classification layer 122b. For example, each vehicle function is assigned at least one parameter from a parameter class. Based on the second outputs, it can now be determined whether a token corresponds to the parameter class assigned to the previously determined vehicle function.

[0055] 25-0653 An exemplary training dataset for the second classification layer 122b comprises tokens corresponding to parameters of the controllable vehicle functions and labeled according to their associated parameter class. During training, the parameters of the second classification layer 122b are varied until the output of the second classification layer 122b matches the correct labels. In particular, the language model 120 is not modified during the training of the second classification layer 122b. Training the second classification layer 122b can be performed as an optional step within the procedure. Specifically, training the second classification layer 122b can be repeated as part of the procedure if, for example, new vehicle functions become available in the vehicle 102, parameter classes of parameters change, or vehicle functions accept new parameters.

[0056] In step S212, an initial output is generated for each token corresponding to a parameter of the vehicle function determined in step S208. This initial output indicates, for each parameter value in a predefined list of parameter values, the probability that the token corresponds to that parameter value. The initial output is generated using the first classification layer 122a. For this purpose, the first classification layer 122a can, for example, be trained to generate an ordered list of numerical values ​​from the token embedding, indicating for each parameter value in the predefined list the probability that the token corresponds to that parameter value. Alternatively, the first classification layer 122a can also output the most probable parameter values, such as the two, three, or ten most probable parameter values.

[0057] An example training dataset for the first classification layer 122a comprises tokens corresponding to parameters of the controllable vehicle functions and labeled according to their assigned parameter values. During the 25-0653 training, the parameters of the first classification layer 122a are varied until the output of the first classification layer 122a corresponds to the correct labels. Specifically, the language model 120 is not modified during the training of the first classification layer 122a. The training of the first classification layer 122a can be performed as an optional step within the procedure. In particular, the training of the first classification layer 122a can be repeated as part of the procedure if, for example, new vehicle functions become available in the vehicle 102 or if vehicle functions accept new parameter values ​​as parameters.

[0058] In step S214, the vehicle function determined in step S208 is activated, taking into account the parameter values ​​determined in step S212. The process then concludes in step S216.

[0059] In the embodiments described with reference to Figures 1 and 2, at least the receiver module 104, the storage element 106, the processing module 108, and the control module 110 constitute the device 100 of a vehicle 102 for controlling vehicle functions. Further elements and features shown in Figures 1 and 2 and mentioned in the preceding description may be part of the device 100. Likewise, method steps described with reference to the device 100 may be part of the claimed method.

[0060] 25-0653 Reference number list

[0061] 100 Device

[0062] 102 vehicles

[0063] 104 Receiving module 106 Storage element 108 Processing module 110 Control module

[0064] 112 Vehicle occupants 114 Microphone

[0065] 116 Processing unit 118 Tokenizer

[0066] 120 Language model 122a, 122b, 122c Classification layer 124 Output unit< / position> < / position> < / position> < / position> < / position> < / position>

Claims

Claims 1. Device (100) of a vehicle (102) for controlling vehicle functions comprising a receiving module (104) designed to receive spoken user input from a vehicle occupant (112) that corresponds to a vehicle function desired by the vehicle occupant (112) and includes at least one parameter, a storage element (106) on which a tokenizer (118), a language model (120) and at least one first classification layer (122a) are stored, wherein the tokenizer (118) is trained to perform a segmentation of the user input into tokens, wherein the language model (120) is trained to generate, based on the segmented user input, for each token of the user input, an embedding that corresponds to the semantic meaning of the token in the context of the semantic meaning of the user input, and wherein the first classification layer (122a) is trained, based on the embeddings, to generate a first output for each token of the user input that indicates, for each parameter value of a predefined list of parameter values, the probability with which the token corresponds to the parameter value. a processing module (108) trained to load and execute the tokenizer (118), the language model (120) and the first classification layer (122a) from the storage element (106), and a control module (110) that is trained to determine, based on user input, a vehicle function that most likely corresponds to the desired vehicle function, that is trained, based on user input, to determine whether a token corresponds to a parameter of the determined vehicle function, that is trained, for tokens that correspond to a parameter of the determined vehicle function, to determine a parameter value of the parameter based on the first output, and that is trained to control the determined vehicle function taking into account the determined parameter value.

2. Device (100) according to claim 1, wherein the first classification layer (122a) has been trained using a first training data set comprising exemplary user inputs, the words of which correspond to a parameter are labelled with the corresponding parameter value, and using the language model (120) by inputting the exemplary user inputs to the tokenizer (118) as input.

3. Device (100) according to one of claims 1 or 2, wherein several first classification layers (122a) are stored on the storage element (106), each of which is assigned to one of the parameter classes; wherein the processing module (108) is configured to load and execute the first classification layers (122a) from the storage element (106); and wherein the control module (110) is configured to select one of the first classification layers (122a) based on the second output for each token corresponding to a parameter and to control the processing module (108) to operate the selected first classification layer (122a) with the token as input.

4. Device according to one of the preceding claims, wherein a second classification layer (122b) is stored on the storage element (106), which is trained to generate a second output based on the embeddings for each token of the user input, which indicates for each parameter class from a predefined list of parameter classes, each of which is assigned to a vehicle function that can be controlled by the control unit, the probability with which the token is included in the parameter class; wherein the processing module (108) is configured to load and execute the second classification layer (122b) from the storage element (106); and wherein the control module (110) is designed to determine, based on the second output, for each token of the user input, whether the token is a parameter of the determined vehicle function.

5. Device (100) according to claim 4, wherein the second classification layer (122b) is trained using a second training data set comprising exemplary user inputs, the words of which correspond to a parameter are labelled with the corresponding parameter classes of the predefined list, and is trained using the language model (120) by providing the exemplary user inputs to the tokenizer (118) as input.

6. Device (100) according to one of the preceding claims, wherein a third classification layer (122c) is stored on the storage element (106), which is trained to generate a third output based on the embeddings, indicating which of the controllable by the control module (110) 25-0653Vehicle functions most likely corresponds to the desired vehicle function; wherein the processing module (108) is configured to load and execute the third classification layer (122c) from the storage element (106); and wherein the control module (110) is designed to determine, based on the third output, the vehicle function that most likely corresponds to the desired vehicle function.

7. Device (100) according to claim 6, wherein the third classification layer (122c) has been trained using a third training data set comprising exemplary user inputs, each labelled with a corresponding vehicle function, and using the language model (120) by inputting the exemplary user inputs to the tokenizer (118).

8. Method for controlling vehicle functions of a vehicle (102), wherein a) a spoken user input is received from a vehicle occupant (112) which corresponds to a vehicle function desired by the vehicle occupant (112) and includes at least one parameter; b) user input is segmented into tokens; c) using a language model (120) and based on the segmented user input, for each token of the user input a 25-0653Embedment is created that corresponds to the semantic meaning of the word in the context of the semantic meaning of the user input; d) based on user input, a vehicle function is determined which most likely corresponds to the desired vehicle function; e) based on user input, it is determined whether a token corresponds to a parameter of the determined vehicle function; f) using a first classification layer (122a) and based on the embeddings, for each token corresponding to a parameter of the determined vehicle function, a first output is generated which indicates, for each parameter value from a predefined list of parameter values, the probability that the token corresponds to the parameter value; and g) the determined vehicle function is controlled taking into account the determined parameter values.

9. The method of claim 8, wherein, using a second classification layer (122b) and based on the embeddings for each user input token, a second output is generated which indicates, for each parameter class from a predefined list of parameter classes, each of which is assigned to a vehicle function that can be controlled by the control unit, the probability with which the token is included in the parameter class; and Based on the second output, it is determined for each token of the user input whether the word is a parameter of the determined vehicle function. 25-065310. Method according to claim 8 or 9, wherein, using a third classification layer (122c) and based on the embeddings, a third output is generated which indicates which of the vehicle functions controllable by the control module (110) most likely corresponds to the desired vehicle function; and where, based on the third output, the vehicle function is determined that most likely corresponds to the desired vehicle function.

11. Method according to any one of claims 8 to 10, wherein the first classification layer (122a), the second classification layer (122b) and / or the third classification layer (122c) are retrained when a vehicle function has changed, when a previously available vehicle function is no longer available and / or when a new vehicle function is available. 25-0653