Device and method for controlling vehicle functions

WO2026201349A1PCT designated stage Publication Date: 2026-10-01BAYERISCHE MOTOREN WERKE AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/052250
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2026-01-29
Publication Date
2026-10-01

Smart Images

  • Figure EP2026052250_01102026_PF_FP_ABST
    Figure EP2026052250_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a device (100) of a vehicle (102) for controlling vehicle functions, comprising a receiving module (104) is designed to receive spoken user input from a vehicle occupant (112), which corresponds to one or more vehicle functions desired by the vehicle occupant (112). The device (100) also comprises a storage element (106) on which a tokenizer (118), a language model (120) and a first classification layer (122a) are stored, wherein the tokenizer (118) is designed to carry out a segmentation of the user input into tokens. The language model (120) is trained such that, based on the user input, it generates an embedding for each token of the user input, which corresponds to the semantic meaning of the token in the context of the semantic meaning of the user input. The first classification layer (122a) is trained to generate a first output based on the embeddings for each token. A processing module (108) comprised by the device is designed to load and execute the tokenizer (118), the language model (120) and the first classification layer (122a) from the storage element (106). The device also comprises a control module (110), which is designed to determine, on the basis of the first output, a number of vehicle functions to be controlled, which is designed to determine, on the basis of the segmented user input, a probability for each vehicle function that can be controlled by the control module (110), said vehicle function corresponding to one of the desired vehicle functions, and which is designed to control a number of vehicle functions which are most likely to correspond to one of the desired vehicle functions. The number of controlled vehicle functions is equal to the determined number of vehicle functions to be controlled.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Device and method for controlling vehicle functions

[0002] The invention relates to a device for a vehicle for controlling vehicle functions. The invention further relates to a method for controlling vehicle functions.

[0003] Many vehicle functions in modern vehicles can be activated with spoken user input. This allows, for example, a driver to operate vehicle functions without taking their hands off the steering wheel or their eyes off the road. However, so-called voice assistants are limited by their language understanding and can often only correctly interpret specific, predefined commands. Variations in voice input that deviate from these predefined commands pose a challenge. Since modern vehicles have a multitude of functions, for example, 1000, the challenge of selecting and correctly activating the right function is particularly significant. Users can express their commands in very different ways, which increases the complexity of correctly classifying and assigning voice inputs to the corresponding vehicle functions.In particular, voice inputs that are intended to control several vehicle functions simultaneously can be difficult to classify.

[0004] From US patent 11,658,835 B2, a method is known in which expressions in a user request are recognized, each separating several parameters of a call function. This method initiates a group call.

[0005] The object of the invention is to provide a device for a vehicle and a method for controlling vehicle functions that can implement spoken user inputs better than known devices and methods, especially when the spoken user inputs relate to several vehicle functions simultaneously.

[0006] 25-0654 This problem is solved by a device having the features of claim 1 and by the subject matter of the dependent claim. Further developments are specified in the dependent claims.

[0007] The proposed device for controlling vehicle functions comprises a receiver module configured to receive spoken user input from a vehicle occupant. This spoken user input corresponds to one or more vehicle functions requested by the occupant. The device further includes a memory element containing a tokenizer, a language model, and a first classification layer. The tokenizer is configured to segment the user input into tokens, and the language model is trained to generate an embedding for each token based on the user input, corresponding to the token's semantic meaning within the context of the user input. The first classification layer is trained to generate an initial output based on these embeddings for each token.Furthermore, the device includes a processing module configured to load and execute the tokenizer, the language model, and the first classification layer from the storage element. The device also includes a control module configured to determine a number of vehicle functions to be controlled based on the initial outputs. The control module is further configured, based on the segmented user input, to determine for each vehicle function controllable by the control module the probability that the vehicle function corresponds to one of the desired vehicle functions. The control module is further configured to control a number of the vehicle functions that have the highest probabilities of corresponding to one of the desired vehicle functions, wherein the number of controlled vehicle functions is equal to the determined number of vehicle functions to be controlled.

[0008] 25-0654 The proposed device implements a voice assistant for the vehicle, which controls one or more vehicle functions based on spoken user input. In particular, the user input is intended to control multiple vehicle functions simultaneously.

[0009] For example, the user input is: "Close the window and activate the heating," which is intended to control the two vehicle functions FN_CLOSE_WINDOW and FN_ACTIVATE_HEATING. However, these two vehicle functions can be controlled in different ways by a single user input. For instance, a user might formulate a very concise input: "Window closed, heating on!" The proposed device recognizes the user's intention to control multiple vehicle functions with one input and controls the vehicle functions accordingly.

[0010] First, the user input is segmented into different tokens using the tokenizer. These tokens can include, for example, a word, part of a word (such as a prefix), or punctuation marks (such as a comma in a list). The segmented user input is then processed by the language model. The output of the language model consists of embeddings of the tokens, i.e., mathematical representations of the semantic content of the respective tokens, for example, as vectors in a high-dimensional vector space, the embedding space. Embeddings of tokens with similar semantic content will be located closer together in the embedding space than embeddings of tokens with different semantic content. In particular, tokens that refer to a separator (such as a comma) or a separator word (such as "and") will be located close together in the embedding space.This allows the appropriately trained first classification layer to generate the first outputs, which can be used to determine how many vehicle functions the driver wants to control.

[0011] 25-0654 The control module determines the number of vehicle functions to be controlled based on the initial output. For example, the user input is "Open the window and turn on the heater." In this example, the control module determines that two vehicle functions should be controlled.

[0012] Furthermore, based on user input, for example using embeddings generated by the language model or other methods, the control module determines several vehicle functions that most likely correspond to the vehicle functions desired by the vehicle occupant. For example, based on the user input "Open the window and turn on the heating," the control module determines a probability for each vehicle function controllable by the control module that the respective vehicle function corresponds to the desired vehicle function. Thus, a probability of 0.4 is determined for the vehicle functions FN_OPEN_WINDOW and FN_ACTIVATE_HEATING, and a probability of 0.1 is determined for the vehicle functions FN_ACTIVATE_RADIO and FN_OPEN_TAILGATE.In particular, the determined probabilities can be normalized according to the number of vehicle functions to be controlled, so that the sum of the probabilities equals the number of vehicle functions to be controlled. In the example above, this means that for the probabilities of the vehicle functions FN_OPEN_WINDOW and FN_ACTIVATE_HEATING, for example, a probability of 0.8 is determined for each, and for the vehicle functions FN_ACTIVATE_RADIO and FN_OPEN_TAILGATE, a probability of 0.2 is determined for each.

[0013] The control module then activates a number of vehicle functions corresponding to the previously determined number. In this example, two vehicle functions are activated. The activated vehicle functions are those with the highest probability. Based on the previously determined number...

[0014] Referring to the example described in 25-0654, the vehicle functions FN_OPEN_WINDOW and FN_ACTIVATE_HEATING are controlled, as these have the highest probability of corresponding to a desired vehicle function.

[0015] The classification of user input words by the first classification layer enables the voice assistant generated by the device to reliably determine how many vehicle functions are to be controlled by the user input. This allows the device to reliably process user inputs that relate to multiple vehicle functions simultaneously. Furthermore, the classification of the tokens can be performed in parallel with other operations for processing the user input. This enables a more efficient implementation of spoken user input, especially on limited hardware such as that typically found in vehicles, than with known devices and methods.

[0016] In one embodiment, the first classification layer is trained to determine, based on the embedding, a corresponding intention for each token of the user input, to index all tokens for which the same intention was determined with the same index, and to generate the initial outputs such that they each include the index with which the corresponding token was indexed. In such an embodiment, the initial outputs can be used to count how many different intentions the user input encompasses. To index the intentions, the first classification layer exploits the fact that the embeddings of tokens belonging to the same intention will be located close to each other in the embedding space. It can be assumed that each intention corresponds to a vehicle function to be controlled. Thus, based on the embeddings, the first classification layer can determine the number of actions to be executed.

[0017] 25-0654 Determine vehicle functions particularly easily and reliably and generate the first outputs accordingly.

[0018] In one embodiment, the first classification layer is trained using a first training dataset containing sample user inputs, each token labeled according to an intention, and using the language model by feeding the sample user inputs into the language model. The first classification layer is trained using a supervised learning approach. During the training of the first classification layer, the language model remains unchanged. For example, an element of the training dataset consists of a token corresponding to a word from a sample user input and an index that assigns the token to one of several intentions encompassed by the sample user input.

[0019] For example, the user input is "open the window and turn on the heating." A token belonging to the word "window" could be indexed with 1 to assign the token to a first intention ("open the window"). Another token belonging to the word "heating" could be indexed with 2 to assign the token to a second intention ("turn on the heating"). This training enables the first classification layer to robustly assign tokens from the segmented user input to an intention, even if these tokens were not part of the training dataset. Alternatively or additionally, a training dataset can be used that includes the corresponding output of the language model instead of the example tokens.

[0020] In one embodiment, the spoken user input includes at least one separator word or character that semantically separates the desired vehicle functions. The first output indicates a probability for each token that it corresponds to a separator word or character. In such an embodiment, the first classification layer counts the number of separators, such as "and," and separators, such as commas, in the user input. Based on this count, the number of vehicle functions to be controlled by the user input can be determined particularly easily, since it can be assumed that each of the vehicle functions to be controlled corresponds to a sentence fragment separated from the other sentence fragments by a separator word or character.

[0021] In one embodiment, the first classification layer is trained using a first training dataset containing sample user inputs with delimiters and / or separators, each appropriately labeled, and using the language model by feeding the sample user inputs to the language model. In this embodiment, the first classification layer is also trained using a supervised learning approach. The language model remains unchanged during the training of the first classification layer. For example, an element of the training dataset consists of a token corresponding to a delimiter or separator and a label identifying the token as a delimiter or separator.The training enables the first classification layer to robustly classify tokens from segmented user input as separators or delimiters, even if these tokens were not part of the training dataset. Alternatively or additionally, a training dataset can be used that contains the corresponding output of the language model instead of the example tokens.

[0022] In one embodiment, the first classification layer comprises at least one neural network. Neural networks are capable of reliably recognizing complex patterns even in high-dimensional datasets, such as the embedding space. Thus, the neural network enables the first classification layer to reliably classify the tokens. Alternatively or additionally

[0023] 25-0654 the second classification layer may include further elements, for example elements of a transformer architecture such as one or more attentionheads.

[0024] In one embodiment, the control module is configured to divide the user input into a number of parts corresponding to the number of vehicle functions desired by the vehicle occupant, based on the initial output. The control module is further configured to determine, based on each part of the user input, the probability that the vehicle function corresponds to one of the desired vehicle functions. In this embodiment, the user input is divided into parts, each corresponding to a vehicle function to be controlled. For example, tokens classified as separators or delimiters are used to divide the user input into parts. These parts are then processed again by the language model to determine which vehicle function is to be controlled by which part.This approach leverages the language model's ability to classify intentions exceptionally well, even when these intentions are expressed in previously untrained linguistic variations. This allows the various components of user input to be reliably assigned to specific vehicle functions.

[0025] In one embodiment, the control module is configured to generate at least one embedding for each part of the user input using the language model, corresponding to the semantic meaning of that part. A second classification layer can be stored on the memory element, which is trained to generate a second output based on the embeddings for each part of the user input. This second output indicates, for each vehicle function controllable by the control module, the probability that the user input part relates to that vehicle function. Furthermore, the processing module can be configured to load and execute the second classification layer from the memory element. The control module is then configured, for example, as shown in 25-0654, to determine, based on the second outputs for each vehicle function, the probability that the vehicle function corresponds to one of the desired vehicle functions.The output of the second classification layer includes, for example, an ordered list containing a numerical value for each of the vehicle functions controllable by the control module. This value indicates the probability that the vehicle function corresponds to the desired vehicle function. In such an embodiment, the second classification layer resolves the embeddings of the user input components by reducing the high-dimensional embeddings to a single vector whose dimensionality corresponds to the number of controllable vehicle functions. Each entry in this vector represents the probability that one of the vehicle functions is the vehicle function to be controlled by that component. The control module can then, for example, determine the vehicle function with the highest probability for each vector and control it accordingly.If one or more desired vehicle functions cannot be clearly determined, the control module can, for example, activate an output unit of the vehicle to prompt the vehicle occupant to repeat the user input.

[0026] In one embodiment, a second classification layer is stored on the memory element. This second layer is trained to generate a second output based on the embeddings, indicating a probability for each vehicle function controllable by the control module that the vehicle function corresponds to one of the desired vehicle functions. The control module is configured to determine, based on this second output, a probability for each controllable vehicle function that the vehicle function corresponds to one of the desired vehicle functions. The embeddings of tokens with similar semantic content generated by the language model are located closer together in the embedding space than the embeddings of tokens with different semantic content. This enables the appropriately trained second layer to...

[0027] 25-0654 Classified! The language model assigns different user inputs related to the same vehicle function to the corresponding vehicle function. For example, a user input might be to open a vehicle window. The user input could be: "Open the window." However, the vehicle occupant could also use alternative phrases, such as: "Open the window" or "Lower the window." The language model assigns embeddings to all these user inputs, all of which lie within a limited neighborhood of the embedding space. The second classification layer then assigns the vehicle function "Open window" to all embeddings in this neighborhood. The control module then executes the vehicle function determined by the classification layer. Thus, the voice assistant created by the device is able to control vehicle functions even based on complex and variably formulated user inputs.

[0028] In one embodiment, the output of the second classification layer comprises an ordered list containing a numerical value for each of the vehicle functions controllable by the control module. This value indicates the probability that the vehicle function corresponds to one of the desired vehicle functions. In such an embodiment, the second classification layer resolves the embeddings by reducing the high-dimensional embeddings to a vector whose dimensionality corresponds to the number of controllable vehicle functions. Each entry in such a vector corresponds to the probability that one of the vehicle functions is one of the desired vehicle functions. If one or more desired vehicle functions cannot be uniquely determined, the control module can, for example, activate an output unit in the vehicle to prompt the vehicle occupant to repeat the user input.

[0029] In one embodiment, the second classification layer is trained using a second training dataset containing exemplary user inputs and the language model (25-0654) by feeding these exemplary user inputs into the language model. Each exemplary user input is labeled with one or more vehicle functions that are to be controlled by the exemplary user input. The training of the second classification layer can also be performed using a supervised learning approach. During the training of the second classification layer, the language model remains unchanged. For example, one element of the training dataset consists of the user input "Open the window" and the vehicle function FN_OPEN_WINDOW as its label. The tokenizer segments the exemplary user inputs and feeds them into the language model.The language model generates embeddings for each segmented user input, which are then fed into the classification layer as training input. The labeled output is the vehicle function or functions corresponding to the respective example user input. This training enables the classification layer to robustly assign variations in user input to their respective vehicle functions, even if these variations were not part of the training dataset. For example, the following example user inputs were used as part of the training dataset for the vehicle function FN_OPEN_WINDOW: "Open the window" and "Open the window." After training, variations of these user inputs, such as "Window down," "Open window," and "Please open the window," will also have a high probability of being correctly assigned to the vehicle function FN_OPEN_WINDOW.

[0030] In one embodiment, the control module is configured to determine a sequence of vehicle functions to be controlled based on the embedded data. The control module is further configured to control the vehicle functions according to this determined sequence. Certain vehicle functions must be executed before others. For example, a route destination must first be defined before points of interest at destination 25-0654 can be determined. In this embodiment, the control module determines a sequence in which the vehicle functions to be executed should be controlled to avoid conflicts. For example, a predefined sequence of all controllable vehicle functions can be used.

[0031] Alternatively, the second classification layer and / or the language model can be trained to determine the order in which the vehicle functions to be controlled are activated. For example, the second classification layer is trained using a training dataset containing sample user inputs and using the language model by feeding the sample user inputs into the language model. The sample user inputs are each labeled with at least a number of vehicle functions to be controlled by the sample user input and their logical execution order. Alternatively or additionally, the sample user inputs are each labeled with at least a number of vehicle functions to be controlled by the sample user input.Additionally, each token of the respective user input, which indicates the position of a vehicle function in the sequence of vehicle functions to be controlled, is labeled accordingly. The tokenizer segments the sample user inputs and feeds them to the language model. The language model generates embeddings for each segmented user input, which are then fed to the second classification layer as training input. This training of the second classification layer also occurs unsupervised and preferably without modifying the language model. The training enables the second classification layer to determine the sequence independently.

[0032] In one embodiment, the second classification layer comprises at least one neural network. Neural networks are capable of reliably recognizing complex patterns even in high-dimensional datasets, such as the embedding space. This enables the neural network to correctly assign embeddings to the same vehicle function, even if they relate to the same vehicle function but have semantic content that is so different they do not lie within a geometrically simple neighborhood within the embedding space, for example, "Open the window" and "The air is very bad." Alternatively or additionally, the second classification layer can include further elements, such as elements of a transformer architecture like one or more attentionheads.

[0033] In one embodiment, the language model is designed as an encoder-decoder model. Encoder-decoder models, such as BERT, are particularly good at capturing complex patterns in input sequences and converting them into meaningful output. They can process the entire input before generating an ordered and coherent output, which is especially advantageous for capturing user input.

[0034] In one embodiment, the language model is a Large Language Model or a Small Language Model. Large Language Models (LLM) and Small Language Models (SLM) are classes of language models that differ primarily in the size of the training dataset used to train them. SLMs, in particular, can be optimized for a specific task. LLMs are typically trained with a text corpus that can be several hundred gigabytes in size. SLMs are typically trained with a text corpus that is only a few gigabytes in size. LLMs and SLMs also differ in the number of variables and thus the size of the model itself. An LLM can have one hundred billion variables; for example, GPT-3 has 175 billion variables, while an SLM typically has no more than one billion variables; for example, BERT has 340 million variables.By appropriately selecting the model size, low latency can be ensured, and the 25-0654 speech model can also be run on the vehicle's limited hardware. For example, BERT can be operated with an inference time of 20 ms, which is imperceptible to humans. The speech model can be either BERT or one of its many successors and enhancements, such as DistilBERT, ALBERT, roBERTa, ELECTRA, and T5.

[0035] In one embodiment, the language model has been trained using at least a generic text corpus. This generic text corpus comprises texts that are not limited to a specific subject area. For example, the Toronto Book Corpus or a filtered version of Wikipedia can be used as the generic text corpus. Training with the generic text corpus gives the language model a general understanding of language and a broad knowledge base. This general knowledge is also referred to as world knowledge. This world knowledge enables the language model to assign different tokens with the same semantic content to an embedding in the same neighborhood of the embedding space, even if it has not been specifically trained on those tokens.

[0036] In one embodiment, the receiving module is configured to convert the user input into a text format that can be processed by the tokenizer. For example, the receiving module can be configured to generate a string based on the spoken user input, which corresponds to the user input. The tokenizer can be kept particularly simple if the input to the tokenizer is in text format.

[0037] In one embodiment, the processing module is part of the vehicle or part of a processing unit located remote from the vehicle. For example, the processing module is part of a processing unit within the vehicle, such as a central vehicle computer. Alternatively, the processing module can be implemented, at least partially, by a processing unit located remote from the vehicle, such as a server located remote from the vehicle, or in a cloud computing environment. For example, the language model is executed on a server located remote from the vehicle or in a cloud computing environment. In such an embodiment, the language model is preferably stored on a memory element of the processing unit located remote from the vehicle. This allows the use of a language model that might not be able to run on the limited hardware of the vehicle, or only with very high latency.The processing module can also be part of a mobile device that can be installed in the vehicle, for example, a smartphone or a tablet computer belonging to a vehicle occupant. In such an embodiment, the language model is preferably stored on a memory element of the mobile device.

[0038] The invention further relates to a method for controlling vehicle functions of a vehicle. The method includes at least the following steps: a) A spoken user input is received from a vehicle occupant, corresponding to one or more vehicle functions desired by the occupant; b) Using a tokenizer, the user input is segmented into tokens; c) Using a language model, an embedding is created for each token of the user input, corresponding to the semantic meaning of the token within the context of the semantic meaning of the user input; d) Using a first classification layer, a first output is generated for each token based on the embeddings; e) Based on the first output, a number of vehicle functions to be controlled is determined.f) At least based on the segmented user input, a probability is determined for each vehicle function that the vehicle function corresponds to one of the desired vehicle functions. g) A number of the vehicle functions with the highest probabilities of corresponding to one of the desired vehicle functions are activated. The number of activated,

[0039] 25-0654 Vehicle function is equal to the determined number of vehicle functions to be controlled.

[0040] The method has the same advantages as the claimed device.

[0041] In particular, the method can be further developed with features described in this document in connection with the device. Furthermore, the claimed device can be further developed with features described in this document in connection with the method.

[0042] In one embodiment, a second classification layer is used to generate a second output based on the embedded user input. This output indicates the probability that each vehicle function controllable by the control module corresponds to one of the desired vehicle functions. Based on this second output, a probability is determined for each vehicle function that it corresponds to one of the desired vehicle functions. The second classification layer assigns different user inputs, but related to the same vehicle function, to the same vehicle function. This enables a voice assistant operating according to this embodiment to control vehicle functions even based on complex and variably formulated user inputs.

[0043] In one embodiment, the second classification layer is retrained if a vehicle function has changed, if a previously available vehicle function is no longer available, and / or if a new vehicle function becomes available. Specifically, in such an embodiment, only the training of the second classification layer is repeated. This saves the considerable computational effort required for training the language model.

[0044] 25-0654 Exemplary embodiments of the invention are explained in more detail below with reference to the figures. These show:

[0045] Figure 1 shows a schematic representation of a device of a vehicle for controlling vehicle functions according to one embodiment; and

[0046] Figure 2 shows a flowchart of a method for controlling vehicle functions according to one embodiment.

[0047] Figure 1 shows a schematic representation of a device 100 of a vehicle 102 for controlling vehicle functions according to an exemplary embodiment. The device 100 implements a voice assistant for the vehicle 102, which controls vehicle functions of the vehicle 102 based on spoken user input, in particular based on user input that is intended to control more than one vehicle function simultaneously. Vehicle functions that can be controlled by the device 100 include, for example, opening a window of the vehicle 102, starting route guidance, activating seat heating in the vehicle 102, activating ventilation in the vehicle 102, initiating a call, controlling the lighting, controlling an entertainment system in the vehicle 102, controlling a vehicle mode, providing an operating aid, querying the status, and providing a support function for the vehicle 102.The device 100 comprises a receiver module 104, a storage element 106, a processing module 108 and a control module 110, which are shown only by way of example as part of the vehicle 102.

[0048] The receiver module 104 is designed to receive spoken user input from a vehicle occupant 112. The user input always includes several vehicle functions requested by the vehicle occupant 112, i.e., vehicle functions to be executed by the vehicle 102. To receive the user input in spoken form, the

[0049] The receiver module 104 (25-0654) is designed to receive user input in the form of audio data from a microphone 114, for example, a microphone of the vehicle 102 or a microphone of a mobile device paired with the vehicle 102. From the user input, the receiver module 104 can generate audio data or convert the spoken user input into text and make it available for further processing by the device 100. The receiver module 104 is shown purely as an example of part of a processing unit 116 of the vehicle 102, for example, a central vehicle computer.

[0050] Memory element 106 is implemented purely as an example of a memory element 106 of the processing unit 116 of the vehicle 102. A tokenizer 118, a language model 120, and a first classification layer 122a are stored on memory element 106. A second classification layer 122b is also stored on memory element 106, purely as an example. Tokenizer 118 is configured to generate tokens based on user input by segmenting the user input into semantic sections, such as words and punctuation marks. Tokenizer 118 can also generate a token for individual word parts, such as prefixes. Language model 120 has been trained, for example, on a generic text corpus, to generate embeddings from the segmented user input.The embedding of a token is, for example, a vector in a high-dimensional embedding space and corresponds to the semantic meaning of the token in the context of the user input. The first classification layer, 122a, is trained to classify the tokens based on the embeddings in order to determine which tokens correspond to separators or delimiters. The first classification layer, 122a, generates an initial output for each token, corresponding to the result of this classification. Based on this initial output, it is possible to determine how many desired vehicle functions the user input includes. The second classification layer, 122b, is trained to determine, at least for a portion of the user input, which functions are included.

[0051] This part most likely corresponds to the vehicle functions controllable under 25-0654. The second classification layer, 122b, generates a corresponding second output.

[0052] The processing module 108 is configured to operate the tokenizer 118, the language model 120, the first classification layer 122a, and the second classification layer 122b. This means that the processing module 108 is configured to load and execute the tokenizer 118, the language model 120, the first classification layer 122a, and the second classification layer 122b from the memory element 106. The processing module 108 is also shown, purely by way of example, as part of the processing unit 116 of the vehicle 102. In other embodiments, however, the processing module 108 can also be formed wholly or partially by a processing unit 116 located remotely from the vehicle 102. In particular, the language model 120 can be executed on such a processing unit 116 located remotely from the vehicle 102, for example, a backend server or a processing unit 116 implemented in a cloud computing environment.In such an embodiment, the language model 120 is preferably stored on a storage element 106 of the processing unit 116 located away from the vehicle 102.

[0053] Based on the initial outputs, control module 110 determines a number of vehicle functions to be controlled that are encompassed by the user input. Furthermore, based on the segmented user input, control module 110 determines a probability for each controllable vehicle function. The probability determined for a controllable vehicle function indicates how likely this vehicle function is to correspond to one of the desired vehicle functions encompassed by the user input. For example, control module 110 uses the second output generated by the second classification layer 122b for this purpose. Control module 110 then controls a number of vehicle functions that have been determined to be most likely to correspond to one of the desired vehicle functions. The number of vehicle functions controlled corresponds to the previously determined number of vehicle functions to be controlled.If the control module 110 cannot determine without doubt which vehicle functions are the desired vehicle functions, the control module 110 can, for example, control an output unit of the vehicle 102 to prompt the vehicle occupant 112 to repeat the user input.

[0054] Figure 2 shows a flowchart of a method for controlling vehicle functions according to one embodiment. The method implements voice control for the vehicle functions of vehicle 102. The method is described purely by way of example with reference to the device 100 according to Figure 1.

[0055] In step S200, the process is initiated. In step S202, the spoken user input from vehicle occupant 112 is received, corresponding to the vehicle functions desired by the occupant. With this user input, the vehicle occupant specifies which vehicle functions should be activated. For example, the vehicle occupant says: "Close the window and activate the heating" to activate the functions FN_CLOSE_WINDOW and FN_ACTIVATE_HEATING, which close a window of vehicle 102 and turn on the heating, respectively. Optionally, in this step, a text format can be generated from the spoken user input, which is, for example, in the form of audio data, making it easier to process further. This can be done, for example, using a trained audio-to-text model. The user input is received and processed, for example, by the receiver module 104 of the device 100 using microphone 114.

[0056] In step S204, the user input is segmented into tokens. For example, a corresponding token is generated for each word and each punctuation mark in the user input. The result of the segmentation is, for example, an ordered list of strings, each corresponding to a word or punctuation mark 25-0654, or an ordered list of numeric identifiers. For example, the user input "Close the window and activate the heating" becomes the list ["Close", "the", "window", "and", "activate", "the", "heating"]. The user input segmentation is generated, for example, by the processing module 108 using the tokenizer 118. The segmented user input is processed in step S206 using the language model 120. An embedding is created for each token.The embeddings are each a mathematical representation, for example a high-dimensional vector in an embedding space, that corresponds to the semantic meaning of the token in the context of the user input. Step S206 is performed, for example, by the processing unit 116 running the language model 120 to generate the embeddings.

[0057] Language Model 120, for example, is an encoder-decoder model or another suitable machine learning model trained to generate embeddings from tokens, for instance, based on a generic text corpus. Training Language Model 120 can be performed as an optional step within the process. Alternatively, a pre-trained Language Model 120 can be used.

[0058] In step S208, the first outputs are generated using the first classification layer 122a and based on the embeddings created in step S206. These initial outputs allow the system to determine how many desired vehicle functions the user input encompasses. The first classification layer 122a is, for example, an attention layer of the language model 120 or a machine learning model that can be operated independently of the language model 120. Step S208 is performed, for example, by the processing module 108 operating the first classification layer 122a. In step S210, the system then determines, based on these initial outputs, how many desired vehicle functions the user input encompasses. For example, step S210 is performed by the control module 110 25-0654.The following are various examples of how the first output can be generated by the first classification layer 122a and how the number of desired vehicle functions can be determined from these first outputs.

[0059] In a first example, the first classification layer, 122a, is trained to determine an intention for each token and assign tokens with the same intention the same index. This allows the first output to determine how many different intentions are encompassed by the user input. Since it can be assumed that each intention corresponds to one of the desired vehicle functions, the number of different indices corresponds to the number of desired vehicle functions. An example training dataset for the first classification layer, 122a, includes tokens that correspond to words from example user inputs and are each indexed according to an assigned intention. Tokens belonging to the same intention have the same index. These indices are used as labels during this training.

[0060] In a second example, the first classification layer 122a is trained to determine the probability for each token that it is a separator word or a delimiter. This allows the system to determine, based on the initial output, how many sentence parts the user input contains. It can be assumed that each sentence part relates to a different desired vehicle function. Based on the sentence parts, the number of desired vehicle functions can then be determined. In this example, a sample training dataset for the first classification layer 122a includes tokens that correspond to separators and delimiters and are labeled accordingly.

[0061] 25-0654 During training, the parameters of the first classification layer 122a are varied until the output of the first classification layer 122a corresponds to the correct labels. In particular, the language model 120 is not changed during the training of the first classification layer 122a. The training of the first classification layer 122a can be performed as an optional step within the procedure. Specifically, the training of the first classification layer 122a can be repeated as part of the procedure if, for example, new vehicle functions become available in the vehicle 102.

[0062] In step S212, based on the segmented user input, it is determined which of the controllable vehicle functions are most likely encompassed by the user input. For this purpose, the probability is calculated for each of the controllable vehicle functions that this vehicle function corresponds to one of the desired vehicle functions. This is done, for example, using the second classification layer 122b. The following examples describe how the second classification layer 122b can be used in step S212.

[0063] The second classification layer 122b can, for example, be trained to generate an ordered list of numerical values ​​from the embeddings, indicating for each controllable vehicle function the probability that it should be controlled by the user input. Alternatively, the second classification layer 122b can also output the most probable vehicle functions, for example, the two, three, or ten most probable vehicle functions. In doing so, the second classification layer 122b generates another vector with a significantly smaller dimension from the high-dimensional vectors in the embedding space—the ordered list. Thus, the second classification layer 122b assigns to the meaning of the user input, determined by the language model 120, concrete vehicle functions that are likely to be controlled.For training the second classification layer 122b, a training dataset is generated containing sample user inputs, each labeled with corresponding vehicle functions. This training dataset is first processed by the tokenizer 118 to segment the sample user inputs. The segmented sample user inputs are then fed into the language model 120 to generate embeddings for each sample user input. These embeddings are then fed back into the second classification layer 122b as training input. The parameters of the second classification layer 122b are varied until its output corresponds to the correct vehicle functions. The parameters of the language model 120 remain unchanged. This training can be performed as an optional step within the overall process.In particular, the training of the second classification layer 122b can be repeated as part of the procedure if, for example, new vehicle functions are available in the vehicle 102.

[0064] In step S212, the user input can be further divided into parts, each corresponding to one of the desired vehicle functions. For example, the user input is separated by separators or delimiters that can be determined based on the initial output. Each part is then processed separately by language model 120 to generate new embeddings. These new embeddings are then further processed, for example using the second classification layer 122b, to determine the most probable vehicle function for each part.

[0065] Step S212 can be performed, for example, by having the processing module 108 operate the language model 120 and / or the second classification layer 122b. Steps S210 and S212 can be performed simultaneously or in any order. In step S214, the vehicle functions that were previously determined to be most likely to be covered by the user input are then activated. The number of vehicle functions activated in step S214 corresponds to the number of vehicle functions determined in step S210. The procedure is then terminated in step S216.

[0066] In the embodiments described with reference to Figures 1 and 2, at least the receiver module 104, the storage element 106, the processing module 108, and the control module 110 form the device 100 of a vehicle 102 for controlling vehicle functions. Further elements and features shown in Figures 1 and 2 and mentioned in the preceding description may be part of the device 100. Likewise, method steps described with reference to the device 100 may be part of the claimed method. List of reference numerals

[0067] 100 Device

[0068] 102 vehicles

[0069] 104 Receiving module 106 Storage element 108 Processing module 110 Control module

[0070] 112 Vehicle occupants 114 Microphone

[0071] 116 Processing unit 118 Tokenizer

[0072] 120 Language model 122a, 122b Classification layer 124 Output unit

Claims

Claims 1. Device (100) of a vehicle (102) for controlling vehicle functions comprising a receiving module (104) which is configured to receive spoken user input from a vehicle occupant (112) corresponding to one or more vehicle functions desired by the vehicle occupant (112), a storage element (106) on which a tokenizer (118), a language model (120) and a first classification layer (122a) are stored, wherein the tokenizer (118) is trained to perform a segmentation of the user input into tokens, wherein the language model (120) is trained to generate, based on the user input, for each token of the user input, an embedding that corresponds to the semantic meaning of the token in the context of the semantic meaning of the user input, and wherein the first classification layer (122a) is trained to generate a first output based on the embeddings for each token, a processing module (108) that is trained to load and execute the tokenizer (118), the language model (120) and the first classification layer (122a) from the storage element (106), and a control module (110) which is trained to determine, based on the first output, a number of vehicle functions to be controlled, which is trained, based on the segmented user input, for each vehicle function controllable by the control module (110), to determine a probability that the vehicle function corresponds to one of the desired vehicle functions, and which is trained to control a number of the 25-0654 vehicle functions which have the highest probabilities of corresponding to one of the desired vehicle functions, wherein the number of controlled vehicle functions is equal to the determined number of vehicle functions to be controlled.

2. Device (100) according to claim 1, wherein the first classification layer (122a) is trained to determine, based on the embedding, a corresponding intention for each token of the user input, to index all tokens for which the same intention has been determined with the same index, and to generate the first outputs such that they each include the index with which the corresponding token was indexed.

3. Device (100) according to claim 2, wherein the first classification layer (122a) has been trained using a first training data set comprising exemplary user inputs, each token of which is labelled according to an intention, and using the language model (120) by inputting the exemplary user inputs to the language model (120).

4. Device (100) according to claim 1, wherein the spoken user input comprises at least one separator word or separator character that semantically separates the desired vehicle functions from one another, and where the first output for each token gives a probability that the token corresponds to a separator word or a separator character.

5. Device (100) according to claim 4, wherein the first classification layer (122a) has been trained using a first training data set comprising exemplary user inputs with separators and / or delimiters, each of which is appropriately labelled, and using the language model (120) 25-0654 by inputting the exemplary user inputs to the language model (120).

6. Device (100) according to claim 4 or 5, wherein the control module (110) is configured to divide the user input into a number of parts corresponding to the number of vehicle functions desired by the vehicle occupant (112) based on the first output, and is configured to determine, based on each part of the user input, for each vehicle function a probability that the vehicle function corresponds to one of the desired vehicle functions.

7. Device (100) according to claim 6, wherein the control module (110) is configured to generate at least one embedding for each part of the user input using the language model (120), which corresponds to the semantic meaning of the part of the user input, and wherein a second classification layer (122b) is stored on the storage element (106), which is trained to generate a second output based on the embeddings for each part of the user input, which indicates for each vehicle function controllable by the control module (110) a probability that the part of the user input relates to this vehicle function, wherein the processing module (108) is configured to load and execute the second classification layer (122b) from the storage element (106), and wherein the control module (110) is configured to determine, based on the second outputs for each vehicle function, a probability that the vehicle function corresponds to one of the desired vehicle functions. 25-06548. Device (100) according to one of claims 1 to 6, wherein a second classification layer (122b) is stored on the storage element (106), which is trained to generate a second output based on the embeddings, which indicates for each vehicle function controllable by the control module (110) a probability that the vehicle function corresponds to one of the desired vehicle functions, wherein the control module (110) is configured to determine, based on the second output, for each controllable vehicle function a probability that the vehicle function corresponds to one of the desired vehicle functions.

9. Device (100) according to claim 6 or 7, wherein the second classification layer (122b) has been trained using a second training data set comprising exemplary user inputs and using the language model (120) by inputting the exemplary user inputs to the language model (120), where each of the example user inputs is labelled with a number of vehicle functions that are to be controlled by the example user input.

10. Device (100) according to one of claims 6 to 8, wherein the control module (110) is configured to determine a sequence of the vehicle functions to be controlled based on the embeddings, and wherein the control module (110) is configured to control the vehicle functions to be controlled in the determined sequence. 25-065411. Device (100) according to claim 9, wherein the second classification layer (122b) has been trained using a training data set comprising exemplary user inputs and using the language model (120) by inputting the exemplary user inputs to the language model (120), wherein the exemplary user inputs are each labelled with at least a number of vehicle functions to be controlled by the exemplary user input and their logical execution sequence accordingly, and / or wherein the exemplary user inputs are each labelled with at least a number of vehicle functions to be controlled by the exemplary user input, and additionally each token of the respective user input that indicates a position of a vehicle function in the sequence of vehicle functions to be controlled is labelled accordingly.

12. Device (100) according to one of the preceding claims, wherein the language model (120) has been trained at least using a generic text corpus.

13. Method for controlling vehicle functions of a vehicle (102), wherein a) a spoken user input is received from a vehicle occupant (112) which corresponds to one or more vehicle functions desired by the vehicle occupant (112); b) using a tokenizer (118) to segment the user input into tokens; 25-0654c) using a language model (120) for each user input token, an embedding is created that corresponds to the semantic meaning of the token in the context of the semantic meaning of the user input; d) using a first classification layer (122a) based on the embeddings for each token, a first output is issued; e) based on the first output, a number of vehicle functions to be controlled is determined; f) at least on the basis of the segmented user input, a probability is determined for each vehicle function that the vehicle function corresponds to one of the desired vehicle functions; and g) a number of the vehicle functions are controlled which have the highest probability of corresponding to one of the desired vehicle functions, wherein the number of controlled vehicle functions is equal to the determined number of vehicle functions to be controlled.

14. The method of claim 13, wherein, using a second classification layer (122b) and based on the embeddings of the user input, a second output is generated which indicates, for each vehicle function controllable by the control module (110), a probability that the vehicle function corresponds to one of the desired vehicle functions, and 25-0654, in which, based on the second edition, a probability is determined for each vehicle function that the vehicle function corresponds to one of the desired vehicle functions.

15. Method according to claim 14, wherein the second classification layer (122b) is retrained when a vehicle function has changed, when a previously available vehicle function is no longer available and / or when a new vehicle function is available. 25-0654