Computer-implemented method for providing a personalized virtual assistant and vehicle
A pre-trained language learning model in vehicles generates context-dependent responses through a vector database, addressing the lack of effective personalized assistants in vehicles, enhancing user interaction and reducing driver distraction.
Patent Information
- Application Number
- DE102024130970
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-24
- Publication Date
- 2026-04-30
AI Technical Summary
Existing vehicle systems lack an effective personalized virtual assistant that can provide context-dependent, knowledge-based responses to user inputs, reducing driver distraction by utilizing voice commands effectively.
A computer-implemented method using a pre-trained language learning model, such as a neural network, generates and embeds text data from various documents, creates a vector database, and processes user inputs to provide context-dependent, knowledge-based responses through a personalized virtual assistant, controlling output devices for audio or visual display.
The method enables a personalized virtual assistant to offer context-dependent, knowledge-based interactions, enhancing user engagement and reducing driver distraction by providing relevant and specific responses.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a computer-implemented method for providing a personalized virtual assistant and a vehicle.
[0002] Vehicles can provide a variety of functions, which can be implemented, for example, by a computer. Entertainment and / or infotainment systems can be controlled by the computer to output specific audio and / or video signals. Operating the computer can reduce the driver's attention to the road. To minimize this reduction in attention, it is known that the driver can use voice commands to input commands into the computer.
[0003] Furthermore, CN 117455003 A discloses the use of a machine learning system to conduct a continuous dialogue with a vehicle occupant as a virtual assistant. This can be achieved by performing vector mappings of speech, which can then be provided to the machine learning system.
[0004] The object of the invention is to provide a computer-implemented method and a vehicle that provide an improved personalized virtual assistant.
[0005] The problem is solved by the features of the independent claims. Advantageous further developments are the subject of the dependent claims and the following description.
[0006] According to a first aspect, a computer-implemented method for providing a personalized virtual assistant is described, using a pre-trained language learning model that has been trained to provide text data generated at an output level from content data provided at an input level. The method comprises at least the following steps: generating text data from the content data of a multitude of documents using the at least one pre-trained language learning model; embedding the multitude of documents in the text data; creating at least one vector database from the text data with the embedded documents; generating conversation text data from at least one user input using the at least one pre-trained language learning model; and embedding the at least one user input in the conversation text data.Generating response text data as a response from the personalized virtual assistant to the user's at least one input, based on the conversation text data with the at least one embedded user input, using the at least one pre-trained language learning model and the at least one vector database; and providing a control signal displaying the response text data to control an output device.
[0007] The computer-implemented method provides a personalized virtual assistant that can offer knowledge-based responses to user statements. For example, if the user asks a question, the personalized virtual assistant can provide a knowledge-based answer. Furthermore, the personalized virtual assistant can ask a knowledge-based question that is context-dependent and relevant to a user's statement. This is achieved using a pre-trained language learning model, which may include a neural network, particularly an artificial intelligence. The language learning model is further trained to generate text data from content data provided at its input level, drawn from a variety of documents, and to make this text available at the output level.The provided documents can belong to a specific topic, such as a particular genre of music, a specific artist, a specific learning situation, and / or a specific consultant, etc. According to the computer-implemented procedure, a large number of documents are provided to the pre-trained language learning model to generate text data from these documents. This text data is preferably context-dependent with respect to the documents. Furthermore, the text data can preferably be generated uniquely for each document or group of documents. The documents used to generate the generated text data are embedded within it.
[0008] A vector database is created from the text data containing the embedded documents. The embedded documents can be identified by querying the text data within this vector database. User input can then be received, which can be based on audio data of the user's spoken language or on written input. Conversation text data can be generated from this user input by providing it to the pre-trained language learning model at an input level. The language learning model can then provide this conversation text data at an output level. The user input is then embedded within this conversation text data. Response text data can then be generated from this conversation text data, based on the vector database.The response text data can be output as a reply from the personalized virtual assistant to the user's input via a control signal. In particular, the response text data is knowledge-based, as it is based on the vector database whose entries are derived from a multitude of documents. The control signal can then activate an output device to display the response text data. Thus, this computer-implemented method provides an improved personalized virtual assistant, since the response provided by the personalized virtual assistant is knowledge-based.
[0009] According to some embodiments, it is conceivable that the method may further include at least the following step between the step: generating conversation text data, and the step: generating response text data: performing at least one similarity search with the conversation text data using at least one embedded user input based on the vector database to find text data similar to the conversation text data; wherein the step: generating response text data is based on at least one pre-trained language learning model and at least one document embedded in the text data found by the similarity search.
[0010] Before generating the response text data, a similarity search can be performed in the vector database. The conversation text data can be used for this search within the vector database. The search results can include text data from the records that comprise the vector database. The response text data can then be generated from this text data and the documents embedded within it, using the language learning model.
[0011] According to some embodiments, it is conceivable that the method may further include at least the following step between the step: Performing at least one similarity search with the conversation text data, and the step: Generating response text data: Creating at least one query term with the content data of at least one document that is embedded in the text data found by the similarity search; wherein, in the step: Generating response text data, the at least one query term of the at least one input level of the at least one language learning model is provided as a response of the personalized virtual assistant to the at least one user input of the user.
[0012] Using the at least one embedded document obtained from the vector database through similarity research, a query term, also known as a prompt or AI prompt, can be created and provided to an input level of the language learning model. By using the query term, the at least one relevant document can be summarized and presented to the input level of the language learning model in a simplified manner. The creation of the query term can be performed, for example, by a computer.
[0013] According to some embodiments, it is conceivable that in the step: generating text data from content data of a large number of documents, the large number of documents may contain consulting information, in particular driver training information and / or performance training information; music information and / or audio information; and / or text creation information.
[0014] The documents can thus be assigned to different topic areas, so that the text data generated with the language model can also be assigned to different topic areas. In this way, data records can be created for the vector database that can be specifically assigned to a topic area. This allows for the creation of different vector databases that can be assigned to different topic areas, or an additional link can be provided within the vector database to filter by topic area. The response text data of the personalized virtual assistant can thus be generated topic-specifically. The knowledge-based response text data of the personalized virtual assistant can therefore provide a basis for natural and specific conversational communication.
[0015] According to some embodiments, it is conceivable that the method, prior to the step of generating conversation text data, may further comprise at least the following steps: displaying at least two images for the user to select, each image representing and / or symbolizing a different personalized virtual assistant; and determining a user selection, which indicates a choice made by the user between the at least two images.
[0016] Each icon can symbolize a personalized virtual assistant, which can be assigned to a specific topic. By displaying at least two icons, at least two different topic areas can be selected, for which a personalized virtual assistant should be provided. The different personalized virtual assistants can use different vector databases or different datasets within a vector database to generate the response text data. Any similarity search performed can then also be carried out with increased accuracy, since only specific datasets or a specific vector database can be considered.
[0017] According to some examples, it is conceivable that the at least one pre-trained language learning model could be a pre-trained Large Language Model.
[0018] According to a second aspect, a computer program product is described, comprising instructions that, when the program is executed by a computer, cause it to perform the steps of the procedure according to the preceding description.
[0019] The advantages, effects, and further developments of the computer program product result from the advantages, effects, and further developments of the method described above. Therefore, reference is made to the preceding description in this regard. A computer program product can be understood, for example, as a data carrier on which a computer program element is stored, containing instructions executable by a computer. Alternatively or additionally, a computer program product can also be understood, for example, as a permanent or volatile data storage medium, such as flash memory or main memory, that contains the computer program element. However, this does not exclude other types of data storage media that contain the computer program element.
[0020] According to a third aspect, a vehicle is described comprising at least one computer that is trained to carry out the procedure according to the preceding description.
[0021] The advantages, effects, and further developments of the vehicle result from the advantages, effects, and further developments of the procedure described above. To avoid repetition, reference is therefore made to the preceding description in this regard.
[0022] According to some embodiments, it is conceivable that the vehicle may further have at least one input device, in particular a microphone, for detecting a user's voice input as user input, wherein the computer may further be configured to convert the voice input into a text input.
[0023] Converting speech input to text input is optional. In some implementations, an audio signal can be provided via the microphone, the underlying data of which can be directly supplied to an input layer of the correspondingly pre-trained language learning model. Similarly, the text data can also be converted into audio data before being stored in the vector database. In other implementations, text input can be generated from the speech input, which can then be used for similarity searches, if necessary.
[0024] According to some embodiments, it is conceivable that the vehicle may further have at least one output device for outputting an audio signal based on the response text data.
[0025] The output device can then be controlled by the control signal such that the response text data is output, for example, as an audio signal. In particular, the output device can include a loudspeaker. Alternatively or additionally, in some other embodiments, the output device can include a screen on which the response text data can be displayed.
[0026] The invention is described below with reference to an exemplary embodiment and the accompanying drawing. The drawing shows: Fig. 1. A schematic representation of the procedure; Fig. 2. A flowchart of the process: and Fig. 3 a schematic representation of a vehicle.
[0027] The procedure for providing a personalized virtual assistant to control a vehicle function is shown schematically in Fig. 1 is represented and is referenced in its entirety with the reference symbol 100.
[0028] Method 100 uses at least one pre-trained language learning model 20. The pre-trained language learning model 20 can be implemented as an artificial neural network that may have an input level 21 and an output level 23. At least one hidden level can link the input level 21 with the output level 23. All levels of the neural network can contain a multitude of artificial neural nodes, which can be linked to artificial neural nodes of other levels or even within the same level. Data or signals provided at the input level 21 can traverse the hidden levels and are modified or processed in the process. The modified or processed signals are then provided at the output level 23.
[0029] In particular, the pretrained language learning model 20 can be configured as a pretrained Large Language Model (PLLM). The pretrained language learning model 20 can be configured such that it provides content data, especially from documents 18, for example, written and / or pictorial information, which is provided at input level 21, as text data 22 corresponding to the content data at output level 23. The text data 22 can thus be uniquely related to the provided content data.
[0030] This can be according to Fig. 2 in one step 102 by generating text data from content data of a large number of documents 18 using at least one pre-trained language model 20.
[0031] In accordance with step 104, the numerous documents 18 are embedded into the generated text data 22. This means that the text data 22 each contain the documents 18, at least in embedded form, from which they were generated using the language learning model 20.
[0032] In a further step 106, at least one vector database 24 is created from the generated text data 22 with the embedded documents 18. A data record of the vector database 24 can, for example, comprise the generated text data 22 of a document 18 and the document 18 itself, with which the text data 22 stored in the data record was generated. The data records are structured as vectors, whereby the text data 22 and / or the documents 18 can be defined as dimensions of the vector.
[0033] Steps 102, 104 and 106 can be summarized as preparatory step 101.
[0034] The documents 18 used to create the vector database 24 can cover various content topics. For example, the documents 18 can contain consulting information, music information (i.e., in particular information about musical pieces, musicians and / or genres), audio information (in particular songs, chants and musical pieces), or text production information (i.e., in particular information about how to write texts), and / or information about stories, etc.
[0035] Different vector databases can be created for each of these different types of information. Alternatively or additionally, a vector database can be provided in which the data records can also be queried according to the types of information explained above.
[0036] Therefore, a user can initially be provided with various options regarding what kind of information the personalized virtual assistant should provide or in which subject area the personalized virtual assistant should work.
[0037] In an optional step 120, the user can be shown the options explained above, for example, as icons. Different icons can represent different personalized virtual assistants.
[0038] The user can then select one of the personalized virtual assistants by choosing the corresponding icon. This user selection can be determined in a further optional step 122.
[0039] In a further step 108, a conversation with user 26 can be initiated by user 26 performing a user input 28. The user input 28 can be provided to input level 21 of the language learning model 20. At output level 23 of the language learning model 20, generated conversation text data 30 can then be provided.
[0040] The user input data 28 can then be embedded into this conversation text data 30 according to step 110. This involves embedding the user input data 28 that was used to create the conversation text data 30 with the pre-trained language model 20.
[0041] In a further optional step 116, a similarity search 32 can be performed in the vector database 24 using the conversation text data 30 in which the respective user input 28 is embedded. The similarity search 32 can then identify records in which the text data 22 stored therein is similar to the conversation text data 30, for example in word choice, the order of the words used, etc.
[0042] These data sets can then be provided, and the documents 18 embedded in the corresponding text data 22 can also be provided.
[0043] Using these documents 18 identified in this way, a query term 34 can then be created. The content data of each document 18 can be used to create the query term 34. The query term 34 can be generated and provided by a computer.
[0044] The query term 34 can be provided according to step 112 of the input level 21 of the pretrained language learning model 20. The pretrained language learning model 20 then provides response text data 36, which can be used as a response from the personalized virtual assistant to the at least one user input 28 from the user 26.
[0045] The response text data 36 obtained in this way can be considered a knowledge-based response. This means that it explicitly addresses the questions or statements in the user input 28 from user 26. For example, if user input 28 contains a question, the response text data 36 can contain an answer to this question, further information, or even additional questions. If user input 28 contains a statement from user 26 or an answer to a question from the personalized virtual assistant, the response text data can, for example, provide a suitable counter-question, follow-up question, or counter-statement.
[0046] Using the document types described above, the personalized virtual assistant can cover 18 different subject areas. For example, the personalized virtual assistant can offer driver training and provide 36 instructions for the route to be driven via the response text data. Information from other vehicle systems, such as the navigation system, driver assistance systems, and speed measurement systems, can also be incorporated.
[0047] If the personalized virtual assistant is trained as an entertainment assistant, for example, it can ask a child passenger in vehicle 10 what kind of stories the child would like to hear. Based on this, and possibly multi-stage, inquiry, a new story can then be generated—one that most likely didn't exist before—and read aloud to the passenger.
[0048] In another embodiment, the personalized virtual assistant could, for example, be a consultant offering advice, such as coaching, on a specific topic. Method 100 could then be used to conduct an interactive consultation with knowledge-based answers.
[0049] In some other embodiments, the personalized virtual assistant can also be a musician who is knowledgeable about the history of music and can provide relevant information, as well as relevant pieces from a specific genre or period. Alternatively or additionally, a specific type of music can be provided, tailored to the user's mood 26 as determined by the entertainment text data 30 and / or the current conditions in and / or around the vehicle.
[0050] In a further step 128, the response text data 36 can be provided as a control signal, which can be used to control an output device 16, including switching it on and off. The control signal can control the output device 16 in such a way that it outputs the response text data 36. The output can be in text form, for example on a screen, or in speech form, for example by means of a loudspeaker.
[0051] Steps 108 to 122 can be collectively referred to as usage step 103. Usage step 103 can be performed multiple times after at least one execution of step 101.
[0052] The procedure 100 can be carried out in a vehicle 10 according to Fig.3. For this purpose, the vehicle 10 may, for example, have a computer 12 which may be configured to carry out the procedure 100. The computer 12 may be a control computer for the entire vehicle 12 or a separate computer.
[0053] In particular, the computer 12 can have an input device 14, for example a microphone, with which a user 26 can provide, for example, voice input 28. Alternatively, a keyboard can be used instead of a microphone to provide user input 28. Furthermore, it can be provided that voice input can be converted into text input by the computer 12. This can be done, for example, by a transcription module.
[0054] Furthermore, the computer 12 can have at least one output device 16 with which the response text data 36 can be output to the user 26. The output device 16 can, for example, be configured as a loudspeaker and / or as a screen.
[0055] The example described above does not in any way limit the invention. Rather, the invention can be modified in numerous ways. All features of the invention described above can be essential to the invention, either alone or in combination. Reference symbol list 10 vehicles 12 computers 14 Input device 16 Output device 18 documents 20 pre-trained language learning model 21 Input level 22 text data 23 Output level 24 vector database 26 users 28 User input 30 entertainment text data 32 Similarity search 34 query terms 36 response text data QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] CN 117455003 A
[0003]
Claims
[1] Computer-implemented method (100) for providing a personalized virtual assistant for controlling a vehicle function, using a pre-trained language learning model (20) that was subsequently trained to provide text data (22) generated from content data provided at an input level (21) at an output level (23), wherein the method (100) comprises at least the following steps: a. Generating (102) text data (22) from content data of a large number of documents (18) using at least one pre-trained language learning model (20); b. Embedding (104) the multitude of documents (18) into the text data (22); c. Create (106) at least one vector database (24) from the text data (22) with the embedded documents (18); d. Generating (108) entertainment text data (30) from at least one user input (28) of a user (26) using the at least one pre-trained language learning model (20); e. Embedding (110) the at least one user input (28) into the conversation text data (30); f. Generating (112) response text data (36) as a response of the personalized virtual assistant to the at least one user input (28) of the user (26) based on the conversation text data (30) with the at least one embedded user input (28) using the at least one pre-trained language learning model (20) and the at least one vector database (24); and g. Providing (128) a control signal displaying the response text data (36) to control an output device (16). [2] Computer-implemented method (100) according to claim 1, characterized by, that the procedure (100) further includes at least the following step between the step: Generating (108) conversation text data (30), and the step: Generating (112) response text data (36): a. Performing (116) at least one similarity search (32) with the conversation text data (30) with the at least one embedded user input (28) based on the vector database (24) to find text data (22) that are similar to the conversation text data (30); wherein the step: generating (112) response text data (36) based on at least one pre-trained language learning model and at least one document (18) embedded in the text data (22) found by the similarity search (32). [3] Computer-implemented method (100) according to claim 2, characterized by, that the procedure (100) further includes at least the following step between the step: Performing (116) at least one similarity search (32) with the conversation text data (30), and the step: Generating (112) response text data (36): a. Creating (118) at least one query term (34) with the content data of at least one document (18) embedded in the text data (22) found by the similarity search (32); wherein in the step: Generating (112) response text data (36) as a response of the personalized virtual assistant to the at least one user input (28) of the user (26), the at least one query term (34) of the at least one input level (21) of the at least one language learning model (20) is provided. [4] Computer-implemented method (100) according to any one of claims 1 to 3, characterized by, that in the step: Generating (102) text data (22) from content data of a multitude of documents (18), the multitude of documents (18) contain consulting information, in particular driver training information and / or performance training information; music information and / or audio information; and / or text creation information. [5] Computer-implemented method (100) according to any one of claims 1 to 4, characterized by , that the procedure (100) before the step: Generating (108) entertainment text data (30), further comprising at least the following steps: a. Display (120) at least two pictograms for the user to select (26), each pictogram displaying and / or symbolizing a different personalized virtual assistant; and b. Determining (122) a user selection that displays a user selection (26) between the at least two pictograms. [6] Computer-implemented method (100) according to any one of claims 1 to 5, characterized by , that at least one pre-trained language learning model (20) is a pre-trained Large Language Model. [7] Computer program product comprising instructions which, when the program is executed by a computer (12), cause it to perform the steps of the method (100) according to any one of claims 1 to 6. [8] Vehicle (10), comprising at least one computer (12) configured to perform the method (100) according to any one of claims 1 to 6. [9] Vehicle (10) according to claim 8 characterized by , that the vehicle (10) further comprises at least one input device (14), in particular a microphone, for detecting a voice input from a user (26) as user input (28), wherein the computer is further configured to convert the voice input into a text input. [10] Vehicle (10) according to claim 8 or 9, characterized by, that the vehicle (10) further comprises at least one output device (16) for outputting an audio signal based on the response text data (36).
Citation Information
Patent Citations
Machine learning system and method for continuous conversation in vehicle machine, medium and electronic equipment
CN117455003A
System and procedure for a dialogue with a user
DE102020100638A1
DIALOGUE SYSTEMS USING KNOWLEDGE BASES AND LANGUAGE MODELS FOR MOTOR VEHICLE SYSTEMS AND APPLICATIONS
DE102023124878A1
Assistance system for controlling a vehicle component using a large language module
DE102024117097A1
Using large language model(s) in generating automated assistant response(s
US20230074406A1