Device and method for controlling vehicle functions of a vehicle
Patent Information
- Application Number
- PCT/EP2026/052260
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-01-29
- Publication Date
- 2026-10-01
Smart Images

Figure EP2026052260_01102026_PF_FP_ABST
Abstract
Description
[0001] Device and method for controlling vehicle functions of a vehicle
[0002] The invention relates to a device of a vehicle for controlling vehicle functions. The invention further relates to a method for controlling vehicle functions of a vehicle.
[0003] Many vehicle functions in modern vehicles can be activated with spoken user input. This allows, for example, a driver to operate vehicle functions without taking their hands off the steering wheel or their eyes off the road. However, so-called voice assistants are limited by their language understanding and can often only correctly interpret specific, predefined commands. Variations in voice input that deviate from these predefined commands pose a challenge. Since modern vehicles have a multitude of functions, for example, 1000, the challenge of selecting and correctly activating the right function is particularly significant. Users can express their commands in very different ways, which increases the complexity of correctly classifying and assigning voice inputs to the corresponding vehicle functions.Number words in particular can be pronounced in different ways and in many regional variations.
[0004] It is therefore an object of the invention to provide a device and a method for controlling vehicle functions of a vehicle that can implement spoken user inputs better than known devices and methods.
[0005] This problem is solved by a device having the features of claim 1 and by the subject matter of the dependent claim. Advantageous embodiments are specified in the dependent claims.
[0006] 25-0899 The proposed device for controlling vehicle functions comprises a receiver module configured to receive spoken user input from a vehicle occupant, corresponding to a vehicle function desired by the occupant and including at least one numeric word as a parameter of the desired vehicle function. The device also includes a memory element on which a tokenizer, a language model, and a text-to-number model are stored, the tokenizer being configured to segment the user input into tokens. The language model is trained to generate an embedding for each token based on the segmented user input, corresponding to the semantic meaning of the token in the context of the user input.The text-to-number model is trained to generate output, based on the token corresponding to the number word or the embedding of the token corresponding to the number word, that includes a numeric value of the number word in the form of a numeric data type. The device further includes a processing module configured to load and execute the language model and the text-to-number model from the memory element. The device also includes a control module configured to determine, based on user input, a vehicle function that most likely corresponds to the desired vehicle function, and, based on the determined vehicle function and the output of the text-to-number model, to control the determined vehicle function, taking into account the numeric value of the number word.
[0007] The proposed device implements a voice assistant for the vehicle that controls a vehicle function based on spoken user input. The user input is intended to control a vehicle function that requires a numerical value as a parameter. For example, the vehicle function FN_INCREASE_TEMP, which increases the temperature in the vehicle's interior, should be controlled. The FN_INCREASE_TEMP function requires a temperature offset as a parameter.
[0008] Parameter 25-0899 specifies the temperature by which the interior temperature of the vehicle should be increased. For example, the function FN_INCREASE_TEMP can only be called with numerical values within a specific range, such as between 0 °C and 16.0 °C, and with a maximum accuracy, such as 0.5 °C. A valid call to the function would therefore be FN_INCREASE_TEMP(temp_offset=1.5). However, in the user input, the desired increase can be expressed in various ways, such as "Please make it 1.5 degrees warmer," "I would like it 1.5 degrees warmer," or "Temperature 2 degrees warmer." The device then recognizes numerical words in the user input and translates them into a machine-readable number format by converting the numerical word into its corresponding numerical value in the numeric data type, such as an integer or a floating-point number.
[0009] First, the user input is segmented into different tokens using the tokenizer. These tokens might include a word, part of a word (such as a prefix), or punctuation. The segmented user input is then processed by the language model. The language model's output consists of embeddings of the tokens, i.e., mathematical representations of the semantic content of each token, for example, as vectors in a high-dimensional vector space, the embedding space. To further process the number word, the device uses the text-to-number model, a small model specifically trained to generate a numeric value in the machine-readable number format. The text-to-number model takes as its input either the embedding of the number word (i.e., the high-dimensional vector in the embedding space) or the token of the number word itself.If the text-to-number model uses tokens, then the text-to-number model can create an embedding of the token in a separate, smaller embedding space for digit recognition only.
[0010] 25-0899 Regardless of whether the text-to-number model uses language model embeddings or tokens, the language model does not need to be called again to generate the numeric value of the number word in the form of the numeric data type. Since calling the language model can sometimes be a time-consuming and computationally intensive step, the device performs the conversion of the number word into a machine-readable format in a particularly computationally efficient manner. Furthermore, the use of a machine learning model allows for more reliable recognition of pronunciation variants that deviate from a standard pronunciation than, for example, a rule-based approach. Additionally, the conversion of the number word can be performed in parallel with other operations for processing user input, such as the classification of other tokens.This allows for a more efficient implementation of spoken user input, especially on limited hardware such as that typically found in vehicles, than with known devices and methods.
[0011] In one embodiment, the text-to-number model is trained using a training dataset comprising exemplary number words, each labeled with a corresponding numeric value in the form of a numeric data type. The exemplary number words are fed into the text-to-number model as input. Preferably, number words in a variety of pronunciation variants, such as different regional and dialectal variants, are used for training the text-to-number model. A list of such words can be generated, for example, using a Large Language Model, such as GPT-4 or BERT. In particular, only number words from a range of numbers and at a resolution that can be accepted as parameters by the controllable vehicle functions can be used for training the text-to-number model.For example, the number range for the function FN_INCREASE_TEMP is 0 °C to 16.0 °C and has a resolution of 0.5 °C. This is used for training the text-to-number model, which only requires parameters.
[0012] 25-0899, to convert this function into a machine-readable form, can be restricted to this range of numbers.
[0013] In one embodiment, the text-to-number model is trained to perform an autoregressive procedure for generating the output of the text-to-number model. In an autoregressive procedure, the numerical value of the number word is built up digit by digit. The final output of the text-to-number model for a number word is, in particular, an end-of-sequence character indicating that the autoregressive procedure is terminated. The text-to-number model has the ability to store the previous sequence in memory (hidden states) and use this information to generate the next digit or punctuation mark. Thus, the text-to-number model produces consistent output even if, for example, parts of the number word are unclear.
[0014] In one embodiment, the text-to-number model is implemented as an attention layer of the language model or as a machine learning model that can be operated independently of the language model, in particular as one of the following machine learning models: a Long Short-Term Memory (LSTM) model, a transformer, a Generalist Reasoning Optimizer (GRO), or a recurrent neural network (RNN). The attention layer is a layer of a neural network, for example, the language model. Attention layers are capable of reliably recognizing complex patterns even in high-dimensional datasets, such as the embedding space. LSTM models have memory cells that retain information across long sequences. RNNs are also particularly good at reliably recognizing long sequences. Therefore, these two models are especially well-suited for recognizing long number words, such as telephone numbers.Transformer and GRO can recognize complex patterns faster than other models. This makes them well-suited for deciphering complex number words, such as Danish or French number words. The aforementioned models enable the text-to-number model to correctly assign a numerical value to complex, long number words, including those that are so different that they do not lie within a geometrically simple neighborhood within the embedding space, for example, "eleven" and "11f".
[0015] In one embodiment, a classification layer is stored on the memory element. This layer is trained to generate an output, based on the embeddings, indicating which of the vehicle functions controllable by the control module most likely corresponds to the desired vehicle function. The processing module is configured to load and execute the classification layer from the memory element. The control module is configured to determine, based on the output of the classification layer, the vehicle function that most likely corresponds to the desired vehicle function. The embeddings of tokens with similar semantic content, generated by the language model, are located closer together in the embedding space than the embeddings of tokens with different semantic content.This allows the appropriately trained classification layer to assign different user inputs, but related to the same vehicle function, to the same vehicle function. For example, the user input might be to open a vehicle window. The user input could be: "Open the window." However, the vehicle occupant could also use alternative phrases, such as: "Open the window" or "Lower the window." The language model then assigns embeddings to all these user inputs, all of which lie within a limited neighborhood of the embedding space. The classification layer then assigns the vehicle function "Open window" to all embeddings in this neighborhood. The control module then executes the vehicle function determined by the classification layer.This means that the voice assistant formed by the device is able to control vehicle functions of the vehicle even on the basis of complex and variably formulated user inputs.
[0016] 25-0899 In one embodiment, the classification layer is trained using a training dataset containing sample user inputs, each labeled with a corresponding vehicle function, and using the tokenizer and the language model. The sample user inputs are fed into the tokenizer as input. The classification layer is trained using a supervised learning approach. During the training of the classification layer, the language model remains unchanged. For example, one element of the training dataset consists of the user input "Please make it 1.5 degrees warmer" and the associated vehicle function FN_INCREASE_TEMP. The tokenizer segments the sample user inputs and feeds them into the language model.The language model generates embeddings for each segmented user input, which are then fed into the classification layer as training input. The labeled output is the vehicle function corresponding to the respective example user input. This training enables the classification layer to robustly assign variations in user input to the associated vehicle function, even if these variations were not part of the training dataset. For example, the following example user inputs were used as part of the training dataset for the vehicle function FN_INCREASE_TEMP: "Temperature two degrees warmer" and "I would like it one and a half degrees warmer." After training, variations of these user inputs, such as "two degrees warmer, please," will also have a high probability of being correctly assigned to the vehicle function FN_INCREASE_TEMP.Alternatively or additionally, a training dataset can be used which includes the corresponding output of the language model instead of the example user inputs.
[0017] In one embodiment, the output of the classification layer includes an ordered list which, for each of the vehicle functions controllable by the control module, includes a numerical value indicating the probability with which the vehicle function corresponds to the desired vehicle function.
[0018] 25-0899 corresponds to this. In such an embodiment, the classification layer resolves the embeddings by reducing the high-dimensional embeddings to a vector whose dimensionality corresponds to the number of controllable vehicle functions. Each entry in this vector corresponds to the probability of one of the vehicle functions being the desired vehicle function. The control module can then, for example, determine the vehicle function with the highest probability and control it. If the desired vehicle function cannot be uniquely determined, the control module can, for example, control an output unit of the vehicle to prompt the vehicle occupant to repeat the user input.
[0019] In one embodiment, the classification layer comprises at least one neural network. Neural networks are capable of reliably recognizing complex patterns even in high-dimensional datasets, such as the embedding space. This enables the neural network to correctly assign embeddings to the same vehicle function, even if they relate to the same vehicle function but have semantic content that is so different they do not lie within a geometrically simple neighborhood within the embedding space, for example, "Open the window" and "The air is very bad." Alternatively or additionally, the classification layer can include further elements, such as elements of a transformer architecture like one or more attentionheads.
[0020] In one embodiment, the language model is designed as an encoder-decoder model. Encoder-decoder models, such as BERT, are particularly good at capturing complex patterns in input sequences and converting them into meaningful output. They can process the entire input before generating an ordered and coherent output, which is especially advantageous for capturing user input.
[0021] 25-0899 In one embodiment, the language model is a Large Language Model or a Small Language Model. Large Language Models (LLM) and Small Language Models (SLM) are classes of language models that differ primarily in the size of the training dataset used to train them. SLMs, in particular, can be optimized for a specific task. LLMs are typically trained with a text corpus that can be several hundred gigabytes in size. SLMs are typically trained with a text corpus that is only a few gigabytes in size. LLMs and SLMs also differ in the number of variables and thus the size of the model itself. An LLM can have one hundred billion variables; for example, GPT-3 has 175 billion variables, while an SLM typically has no more than one billion variables; for example, BERT has 340 million variables.By appropriately selecting the model size, low latency can be ensured, and the speech model can also be run on the vehicle's limited hardware. For example, BERT can be operated with an inference time of 20 ms, which is imperceptible to humans. The speech model can be either BERT or one of its many successors and enhancements, such as DistilBERT, ALBERT, roBERTa, ELECTRA, and T5.
[0022] In one embodiment, the language model is trained using a generic text corpus. This generic text corpus comprises texts that are not limited to a specific subject area. For example, the Toronto Book Corpus or a filtered version of Wikipedia can be used as the generic text corpus. Training with the generic text corpus gives the language model a general understanding of language and a broad knowledge base. This general knowledge is also referred to as world knowledge. This world knowledge enables the language model to embed different tokens with the same semantic content.
[0023] 25-0899 in the same neighborhood of the embedding space, even if it has not been specifically trained on these tokens.
[0024] In one embodiment, the receiving module is configured to convert the user input into a text format that can be processed by the tokenizer. For example, the receiving module can be configured to generate a string based on the spoken user input, which corresponds to the user input. The tokenizer can be kept particularly simple if the input to the tokenizer is in text format.
[0025] In one embodiment, the processing module is part of the vehicle or part of a processing unit located remote from the vehicle. For example, the processing module is part of a processing unit within the vehicle, such as a central vehicle computer. Alternatively, the processing module can be implemented, at least partially, by a processing unit located remote from the vehicle, such as a server located remote from the vehicle, or in a cloud computing environment. For example, the language model is executed on a server located remote from the vehicle or in a cloud computing environment. In such an embodiment, the language model is preferably stored on a memory element of the processing unit located remote from the vehicle. This allows the use of a language model that might not be able to run on the limited hardware of the vehicle, or only with very high latency.The processing module can also be part of a mobile device that can be installed in the vehicle, for example, a smartphone or a tablet computer belonging to a vehicle occupant. In such an embodiment, the language model is preferably stored on a memory element of the mobile device.
[0026] The invention further relates to a method for controlling vehicle functions of a vehicle. The method includes at least the following steps.
[0027] 25-0899 performed: a) A spoken user input is received from a vehicle occupant, corresponding to a vehicle function desired by the occupant and including at least one numeric word as a parameter of the desired vehicle function; b) The user input is segmented into tokens; c) Using a language model and based on the segmented user input, an embedding is created for each token, corresponding to the semantic meaning of the token in the context of the user input; d) Using a text-to-number model and based on the token corresponding to the numeric word, or the embedding of the token corresponding to the numeric word, an output is generated that includes a numeric value of the numeric word in the form of a numeric data type; e) Based on the user input, a vehicle function is determined that most likely corresponds to the desired vehicle function.f) Based on the determined vehicle function and the output of the text-to-number model, the determined vehicle function is controlled taking into account the numerical value of the number word.
[0028] The method has the same advantages as the claimed device.
[0029] In particular, the method can be further developed with features described in this document in connection with the device. Furthermore, the claimed device can be further developed with features described in this document in connection with the method.
[0030] In one embodiment, a training dataset for training the text-to-number model is generated using a Large Language Model (LLM). This LLM is used to instruct the LLM to output a multitude of corresponding number words for a given numerical value. Using the LLM, a list of number words in a variety of pronunciations, particularly different regional and dialectal variants, can be generated for each numerical value to be trained on the text-to-number model. This enables the text-to-number model to reliably recognize number words with a wide range of pronunciations and convert them into a machine-readable format. This, in turn, allows for robust control of the vehicle function.
[0031] In one embodiment, a classification layer and the embedded data generate an output indicating which of the vehicle functions controllable by the control module most likely corresponds to the desired vehicle function. Based on the output of the classification layer, the vehicle function that most likely corresponds to the desired vehicle function is determined. The classification layer assigns different user inputs, but related to the same vehicle function, to the same vehicle function. This enables a voice assistant operating according to this embodiment to control vehicle functions even based on complex and variably formulated user inputs.
[0032] In one embodiment, the text-to-number model and / or the classification layer are retrained when a vehicle function changes, when a previously available vehicle function is no longer available, and / or when a new vehicle function becomes available. Specifically, in such an embodiment, only the training of the text-to-number model or the classification layer is repeated. This saves the considerable computational effort required for training the language model.
[0033] Exemplary embodiments of the invention are explained in more detail below with reference to the figures. These show:
[0034] Figure 1 shows a schematic representation of a vehicle device for controlling vehicle functions; and
[0035] 25-0899 Figure 2 shows a flowchart of a procedure for controlling vehicle functions of a vehicle.
[0036] Figure 1 shows a schematic representation of a device 100 of a vehicle for controlling vehicle functions according to an exemplary embodiment. The device 100 implements a voice assistant for the vehicle 102, which controls a vehicle function of the vehicle 102 based on the spoken user input. Vehicle functions that can be controlled by the device 100 are, in particular, those that have one or more numerical values as parameters, for example, an air conditioning system, a temperature control system, a seat heater, each of which has a temperature as a numerical parameter; a cruise control system, which has a speed as a numerical parameter; an adaptive cruise control system, which has a distance as a numerical parameter; and a media system, which, for example, has a volume as a numerical parameter.The device 100 comprises a receiver module 104, a storage element 106, a processing module 108 and a control module 110, which are shown only by way of example as part of the vehicle 102.
[0037] The receiver module 104 is configured to receive spoken user input from a vehicle occupant 112. The user input always comprises a vehicle function requested by the vehicle occupant 112, i.e., a vehicle function to be executed by the vehicle 102, and at least one numerical word corresponding to a numerical parameter of this vehicle function to be executed. To receive the user input in spoken form, the receiver module 104 can be configured to receive the user input as audio data from a microphone, for example, a microphone 114 of the vehicle 102 or a microphone 114 of a mobile device connected to the vehicle 102. From the user input, the receiver module 104 can generate audio data or process the spoken input.
[0038] 25-0899 Convert user input into text format and make it available for further processing by the device 100. The receiving module 104 is shown purely as an example of part of a processing unit 116 of the vehicle 102, for example, a central vehicle computer.
[0039] Memory element 106 is purely exemplary and represents a memory element 106 of the processing unit 116 of the vehicle 102. A tokenizer 118, a language model 120, and a text-to-number model 122 are stored on memory element 106. Tokenizer 118 is configured to generate tokens based on user input by segmenting the input into semantic sections, such as words and punctuation marks. Tokenizer 118 can also generate a token for individual word parts, such as prefixes. Language model 120 has been trained, for example, on a generic text corpus, to generate embeddings from the segmented user input. The embedding of a token is, for example, a vector in a high-dimensional embedding space and corresponds to the semantic meaning of the token in the context of the user input.The Text-to-Number Model 122 has been trained to generate machine-readable output indicating the numeric value corresponding to the number word in the user input. To achieve this, the Text-to-Number Model 122 produces the output in the form of a numeric data type, such as an integer or a floating-point number. As input, the Text-to-Number Model 122 accepts either the token of the number word or the embedding that corresponds to that token.
[0040] The processing module 108 is configured to operate the tokenizer 118, the language model 120, and the text-to-number model 122. This means that the processing module 108 is configured to load and execute the tokenizer 118, the language model 120, and the text-to-number model 122 from the memory element 106. The processing module 108 is also shown purely as an example, as part of the processing unit 116 of the vehicle 102. In other cases...
[0041] 25-0899In one of the implementation forms, the processing module 108 can also be formed wholly or partially by a processing unit located away from the vehicle 102.
[0042] In particular, the language model 120 can be executed on such a processing unit remote from the vehicle 102, for example a backend server or a processing unit implemented in a cloud computing environment. In such an embodiment, the language model 120 is preferably stored on a memory element of the processing unit remote from the vehicle 102.
[0043] The control module 110 first determines, based on user input, a vehicle function that most likely corresponds to the desired vehicle function. For example, the processing module 108 operates a classification layer that, based on the embedded values, generates an output indicating which vehicle function from a list of controllable vehicle functions most likely corresponds to the desired vehicle function. The output of the classification layer can then be used by the control module 110 to determine the desired vehicle function.If the control module 110 cannot unambiguously determine which vehicle function is the desired one, for example, if none of the vehicle functions has been clearly classified as the most likely, the control module 110 can, for instance, activate an output unit 124 of the vehicle 102 to prompt the vehicle occupant 112 to repeat the user input. Taking into account the output of the text-to-number model 122, this desired vehicle function is then activated. Like the receiver module 104, the storage element 106, and the processing module 108, the control module 110 is also shown, purely by way of example, as part of the processing unit 116 of the vehicle 102.
[0044] Figure 2 shows a flowchart of a method for controlling vehicle functions according to one embodiment. The method implements voice control for the vehicle functions of vehicle 102. The 25-0899 method is described purely by way of example with reference to the device 100 according to Figure 1.
[0045] In step S200, the procedure is initiated. In step S202, the spoken user input from vehicle occupant 112 is received. This input corresponds to the vehicle function desired by vehicle occupant 112 and includes at least one numerical word as a parameter of this vehicle function. With this user input, vehicle occupant 112 specifies which vehicle function should be activated. For example, vehicle occupant 112 says: "Please make it one point five degrees warmer" or "I would like it one and a half degrees warmer" to activate a temperature control function of vehicle 102, specifically to increase the temperature in the interior of vehicle 102 by 1.5 °C. The vehicle occupant 112 could further say: “Turn on cruise control! 100 km / h!”, “Set the cruise control to 100 km / h” or “Please drive at 100 km / h” to control the cruise control of vehicle 102, to set a target speed of 100 km / h.Optionally, in this step, a text format can be generated from the spoken user input, which is available, for example, in the form of audio data, making it easier to process further.
[0046] For example, using a trained audio-to-text model. User input is received and processed, for example, by the receiver module 104 of the device 100 using microphone 114.
[0047] In step S204, the user input is segmented into tokens. For example, a token corresponding to each word in the user input is generated. The result of the segmentation is, for example, an ordered list of strings, each corresponding to a word, or an ordered list of numeric identifiers. For example, the user input "I would like it one and a half degrees warmer" generates the list ["I", "would like", "it", "like", "one and a half", "degrees", "warmer"]. The user input segmentation is performed, for example, by processing module 108 using tokenizer 118. The segmented user input is processed in step S206 using language model 120. An embedding is created for each token.The embeddings are each a mathematical representation, for example a high-dimensional vector in an embedding space, that corresponds to the semantic meaning of the token in the context of the user input. Step S206 is performed, for example, by the processing unit 116 running the language model 120 to generate the embeddings.
[0048] Language Model 120, for example, is an encoder-decoder model or another suitable machine learning model trained to generate embeddings from tokens, for instance, based on a generic text corpus. Training Language Model 120 can be performed as an optional step within the process. Alternatively, a pre-trained Language Model 120 can be used.
[0049] In step S208, the Text-to-Number Model 122 generates the output, which comprises the numeric value of the number word in the form of the numeric data type. For example, the number word "one and a half" is converted to the floating-point number "1.5". As a basis for the conversion, the Text-to-Number Model 122 can use either the token(s) corresponding to the number word or the embeddings of the tokens corresponding to the number word. In particular, the Text-to-Number Model 122 can perform an autoregressive procedure to generate the output. For example, the Text-to-Number Model 122 sequentially generates the outputs "1", ".", "5", "EOS", where "EOS" is an end-of-sequence character and indicates the end of the sequence. The outputs are then concatenated to form the floating-point number "1.5". Step S208 is performed, for example, by the processing unit 116 operating the text-to-number model 122 accordingly.
[0050] 25-0899 The text-to-number model 122 is, for example, an attention layer of the language model 120 or a machine learning model that can be operated independently of the language model 120, such as a long short-term memory model or a recurrent neural network. An example training dataset for the text-to-number model 122 comprises example number words, each labeled with a corresponding numeric value in the form of a numeric data type. This training dataset can be generated, in particular, by a large language model, especially the language model 120, by instructing the large language model to output a multitude of corresponding number words for a given numeric value. The generation of the training dataset for the text-to-number model 122 can be performed as an optional step within the procedure.The text-to-number model 122 is trained by inputting example number words into it. Alternatively, the text-to-number model 122 can be trained using the language model 120. In this case, the example number words are first inputted to the language model 120 to generate corresponding embeddings, which are then inputted to the text-to-number model 122. The parameters of the text-to-number model 122 are then varied until the output of the text-to-number model 122 consistently matches the correct numerical values. Training the text-to-number model 122 can be performed as an optional step within the procedure.In particular, the training of the text-to-number model 122 can be repeated as part of the procedure if, for example, new vehicle functions are available in the vehicle 102 that have a different parameter space than the one trained so far.
[0051] In step S210, a vehicle function is determined based on the user input, which most likely corresponds to the desired vehicle function. For example, the classification layer generates an ordered list of numerical values from the embeddings, indicating for each controllable vehicle function, 25-0899, how likely it is that this function should be controlled by the user input. Alternatively, the classification layer can also output the most probable vehicle functions, for example, the two, three, or ten most probable vehicle functions. In doing so, the classification layer generates another vector with a significantly smaller dimension from the high-dimensional vectors in the embedding space: the ordered list. Thus, the classification layer assigns one or more concrete vehicle functions, which are likely to be controlled, to the meaning of the user input determined by the language model 120.Step S210, for example, is performed by processing unit 116.
[0052] The classification layer has been trained to generate output based on the embeddings, indicating which of the vehicle functions controllable by the control module 110 most likely corresponds to the desired vehicle function. For training the classification layer, a training dataset is created, containing example user inputs, each labeled with a corresponding vehicle function. This training dataset is then processed by the tokenizer 118 to segment the example user inputs. The segmented example user inputs are then fed into the language model 120 as input to generate embeddings for each example user input, which are in turn fed back into the classification layer as training input. The parameters of the classification layer are varied until its output consistently corresponds to the correct vehicle functions.The parameters of language model 120 are not changed. This training can be performed as an optional step within the procedure.
[0053] In particular, the training of the classification layer can be repeated as part of the procedure if, for example, new vehicle functions are available in vehicle 102.
[0054] 25-0899 In step S212, based on the determined vehicle function and the output of the text-to-number model 122, the determined vehicle function is controlled, taking into account the numerical value of the number word. If, for example, several vehicle functions are equally probable, or none of the vehicle functions is more probable than a predetermined limit, such as 50%, or if the numerical value of the parameter cannot be determined, the vehicle occupant 112 may also be prompted to repeat their user input. Step S212 is performed, for example, by the control module 110. The procedure is then terminated in step S214.
[0055] In the embodiments described with reference to Figures 1 and 2, at least the receiver module 104, the storage element 106, the processing module 108, and the control module 110 constitute the device 100 of the vehicle 102 for controlling vehicle functions. Further elements and features shown in the figures and mentioned in the preceding description may be part of the claimed device 100. Likewise, method steps described with reference to the device 100 may be part of the claimed method.
[0056] 25-0899 Reference number list
[0057] 100 Device
[0058] 102 vehicles
[0059] 104 Receiving module 106 Storage element 108 Processing module 110 Control module
[0060] 112 Vehicle occupants 114 Microphone
[0061] 116 Processing unit 118 Tokenizer
[0062] 120 language model
[0063] 122 Text-to-number model 124 Output unit
[0064] 25-0899
Claims
Claims 1. Device (100) of a vehicle (102) for controlling vehicle functions, comprising a receiving module (104) designed to receive spoken user input from a vehicle occupant (112) that corresponds to a vehicle function desired by the vehicle occupant (112) and includes at least one numeric word as a parameter of the desired vehicle function, a storage element (106) on which a tokenizer (118), a language model (120) and a text-to-number model (122) are stored, wherein the tokenizer (118) is trained to perform a segmentation of the user input into tokens, wherein the language model (120) is trained to generate an embedding for each token based on the segmented user input, which corresponds to the semantic meaning of the token in the context of the user input, and wherein the text-to-number model (122) is trained to generate an output based on the token corresponding to the number word, or the embedding of the token corresponding to the number word, which includes a numeric value of the number word in the form of a numeric data type, a processing module (108) trained to load and execute the language model (120) and the text-to-number model (122) from the storage element (106), and a control module (110) which is trained to determine, based on user input, a vehicle function which most likely corresponds to the desired vehicle function, and, based on the vehicle function determined in 25-0899 and the output of the text-to-number model (122), to control the determined vehicle function taking into account the numerical value of the number word.
2. Device (100) according to claim 1, wherein the text-to-number model (122) has been trained using a training data set comprising exemplary number words, each labelled with a corresponding numeric value in the form of the numeric data type, by inputting the exemplary number words to the text-to-number model (122) as input.
3. Device (100) according to claim 1 or 2, wherein the text-to-number model (122) is trained to perform an autoregressive method for generating the output of the text-to-number model (122).
4. Device (100) according to one of the preceding claims, wherein the text-to-number model (122) is configured as an attention layer of the language model (120) or as a machine learning model that can be operated independently of the language model (120), in particular as one of the following machine learning models: a Long Short-Term Memory model, a Transformer, a Generalist Reasoning Optimizer or a Recurrent Neural Network (RNN).
5. Device (100) according to one of the preceding claims, wherein a classification layer is stored on the storage element (106) which is trained to generate an output based on the embeddings indicating which of the vehicle functions controllable by the control module (110) most likely corresponds to the desired vehicle function, wherein the processing module (108) is configured to load and execute the classification layer from the storage element (106), wherein the control module (110) is designed to determine, based on the output of the classification layer, the vehicle function that most likely corresponds to the desired vehicle function.
6. Device (100) according to claim 5, wherein the classification layer has been trained using a training data set comprising exemplary user inputs, each labelled with a corresponding vehicle function, and using the tokenizer (118) and the language model (120) by inputting the exemplary user inputs to the tokenizer (118).
7. Device (100) according to claim 5 or 6, wherein the output of the classification layer comprises an ordered list which includes, for each of the vehicle functions controllable by the control module (110), a numerical value indicating the probability with which the vehicle function corresponds to the desired vehicle function.
8. Device (100) according to one of the preceding claims, wherein the language model (120) is configured as an encoder-decoder model.
9. Device (100) according to any one of the preceding claims, wherein the language model (120) has been trained using a generic text corpus.
10. Device (100) according to any one of the preceding claims, wherein the receiving module (104) is configured to convert the user input into a text format that can be processed by the tokenizer (118).
11. Device (100) according to one of the preceding claims, wherein the processing module (108) is part of the vehicle (102) or part of a processing unit (116) located away from the vehicle (102).
12. Method for controlling vehicle functions of a vehicle (102) wherein: a) a spoken user input is received from a vehicle occupant (112) which corresponds to a vehicle function desired by the vehicle occupant (112) and includes at least one numeric word as a parameter of the desired vehicle function; b) user input is segmented into tokens; c) using a language model (120) and based on the segmented user input, an embedding is created for each token that corresponds to the semantic meaning of the token in the context of the user input; d) using a text-to-number model (122) and based on the token corresponding to the number word, or the embedding of the token corresponding to the number word, an output is generated that includes a numeric value of the number word in the form of a numeric data type; e) based on the user input, a vehicle function is determined that most likely corresponds to the desired vehicle function; and f) based on the determined vehicle function and the output of the text-to-number model (122), the determined vehicle function is controlled taking into account the numerical value of the number word.
13. Method according to claim 12, wherein a training data set for training the text-to-number model (122) is generated using a Large Language Model in which the Large Language Model is controlled to output a plurality of corresponding number words for a numerical value.
14. The method of claim 13, wherein, using a classification layer and based on the embeddings, an output is generated indicating which of the vehicle functions controllable by the control module (110) most likely corresponds to the desired vehicle function; and where, based on the output of the classification layer, the vehicle function that most likely corresponds to the desired vehicle function is determined.
15. Method according to any one of claims 12 to 14, wherein the text-to-number model (122) and / or the classification layer are retrained when a vehicle function has changed, when a previously available vehicle function is no longer available and / or when a new vehicle function is available. 25-0899