Method for the automated filling of a database with user information by means of a spoken dialog system, and spoken dialog system and motor vehicle having the spoken dialog system
The speech dialogue system uses LLMs to automatically collect user information through natural language queries, addressing the limitations of existing methods by ensuring accurate and relevant data collection for personalized vehicle services.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2026-03-26
AI Technical Summary
Existing methods for populating user information in vehicle systems lack the ability to collect accurate and relevant data about user relationships, preferences, and shared activities, often requiring lengthy and uncontrolled data collection that does not account for individual user habits and group dynamics.
A method using a speech dialogue system to automatically populate a database with user information by generating natural language queries based on conversation phases, employing Large Language Models (LLMs) for pattern recognition and feedback extraction, ensuring data relevance and accuracy.
Enables efficient, automated collection of user preferences and relationships, allowing personalized service recommendations and conflict-free decision-making in group settings, with scalable and organized data management.
Smart Images

Figure EP2025076304_26032026_PF_FP_ABST
Abstract
Description
[0001] CARIAD SE 2023P00371
[0002] patent application
[0003] DESCRIPTION
[0004] Method for automatically populating a database with user information using a voice dialogue system, as well as a voice dialogue system and a motor vehicle, comprising the voice dialogue system
[0005] The invention relates to a method for automatically filling a database with user information using a speech dialogue system.
[0006] Mobile services often require user information or data about one or more occupants of a vehicle, their relationships to one another, and their preferences to enable the services they use, such as navigation and / or in-car infotainment systems. For example, when searching for a restaurant, the composition of the occupants and / or their eating habits and / or any food intolerances are crucial. Similarly, for selecting entertainment such as music and / or audiobooks, knowledge of the preferences of at least one occupant in the vehicle is important. Although the technical implementation of such services is generally relatively simple, adapting them to individual users and / or the (specific) circumstances, such as the composition of the occupants in the vehicle, in order to provide appropriate recommendations, is complex.The necessary skills are diverse and cannot be easily acquired through predefined scripts or standardized questionnaires. Additionally, at least one user expects natural language interaction, allowing them to speak freely and / or engage in open dialogues. With the advent of Large Language Models (LLMs) or machine language models, and especially services like Chat-GPT®, a technology is now available that can meet these requirements. This technology can process information about the vehicle user in real time and offer personalized recommendations and / or services based on this information.
[0007] US 2023 / 029 0 342 A1 describes a dialogue system comprising: a database, a speech recognition module configured to convert a user's utterance in a vehicle into text, an intention determination module configured to identify the user's intention based on the text, an emotion determination module configured to identify the user's emotional state based on the identified user intention, and a controller configured to compare data to compare the user's intention and emotional state with rules stored in the database and to determine whether to issue a response to the user's utterance based on the result of the comparison.
[0008] German patent DE 102019217 751 A1 discloses a method for operating a speech dialogue system. Speech input is captured, a first response output is generated based on the speech input using non-goal-directed dialogue analysis, and a second response output is generated based on the speech input using goal-directed dialogue analysis. A first relevance probability is determined for the first response output and a second relevance probability is determined for the second response output, and a speech output is generated based on the response output with the highest relevance probability.The speech dialogue system comprises a capture unit configured to capture speech input, a first and a second dialogue analysis unit configured to generate a first and a second response output based on the speech input, a control unit configured to determine a first relevance probability for the first response output and a second relevance probability for the second response output, and an output unit configured to generate a speech output based on the response output with the highest relevance probability. EP 2 140 341 A1 discloses an emotive consultation system and method.
[0009] The known methods require a lengthy and uncontrolled collection of user data. They do not collect information that corresponds to a predefined category, but only information derived from user behavior. In particular, this approach lacks information about relationships with other people or users, as well as shared preferences, interests, and activities. The collected information is not always accurate or relevant because users, especially within a group, deviate from their usual habits and do things that are irrelevant for later system decisions or controls.
[0010] The invention is based on the objective of providing a method for the automated collection of user information.
[0011] The problem is solved by the subject matter of the independent patent claims. Advantageous embodiments of the invention are described by the dependent patent claims, the following description, and the figures.
[0012] The invention relates to a method for automatically populating a database with user information using a speech dialogue system. In other words, the method relates to the automated acquisition of user information. The following steps are performed: a. Providing a database, in particular in a plain text format such as JSON (JavaScript Object Notation) and / or XML (Extensible Markup Language), comprising one or more data fields that are assigned to one or more predefined categories; b. Based on a recognized category of a conversation phase: generating and submitting a query in the form of natural language information to the user by the speech dialogue system for collecting or storing user information, wherein the query is at least partially related to the recognized category (e.g., at least 30 to 60 percent in percentage points, e.g., realized by natural language pattern recognition); c.Receiving feedback from the user and extracting user information from the feedback, i.e., storing the user information extracted by the speech dialog system in at least one data field.
[0013] In other words, step b initiates and / or conducts a conversation with the user in natural language. The query can initiate this conversation. The conversation can then be completed by receiving the response in step c.
[0014] In other words, the query can be a natural-language conversation with the user. The query is therefore not just a question, but is embedded in a conversation with the user.
[0015] "Automated" means that the process of filling the database with user information by the voice dialogue system is automatic and / or at least partially without manual intervention or human intervention.
[0016] In other words, the process can involve a dynamic description of the initially empty data fields. These data fields can then be populated with content, particularly with the preferences of the driver or user and / or other users.
[0017] The term "conversation phase" can refer to a (predefined) segment within an ongoing conversation between at least two users and / or between at least one user and the speech dialogue system. For example, if two users are discussing music, the speech dialogue system can categorically classify this conversation phase as "music preference" and ask and / or generate corresponding queries to extract user information.
[0018] Overall, it may be possible for the vehicle interior and / or the driver to be (continuously) monitored by the voice dialogue system. As soon as voice activity is detected, e.g., via hotword recognition, the voice dialogue system can populate one or at least one data field with user information by identifying and / or assigning recognized driver preferences through appropriate queries.
[0019] A data field, which can also be called a "database field", refers to a single component within a record in a database.
[0020] "Natural language" means that the query and the entire dialogue between the user and the speech dialogue system is written or expressed in a language used by people in everyday life, without specific adaptations or formalisms for technical systems.
[0021] Step b. can stipulate that the query is only generated and presented to the driver when or as soon as an information criterion is recognized by the voice dialogue system. This information criterion, in the context of a conversation phase, includes the verbal recognition of specific words that correspond at least partially, literally and / or semantically (e.g., implemented using word embeddings), to the categories assigned to the data fields. Additionally or alternatively, the information criterion can include the detection of more than one person in the vehicle and / or the selection of a specific route. Examples of such words are "restaurant," "food," "activity," "sports," and / or "music." The list of words is not exhaustive and / or can be expanded, for example, manually.
[0022] In step b., the system can further stipulate that, after the query is posed, the voice dialog system stores the user's feedback in a ring buffer for a predetermined time, e.g., 10 seconds to 3 minutes, and performs pattern recognition to extract relevant user information. The voice dialog system can then inform the user about the success of the extraction or, if necessary, request additional information and subsequently ask the same or a similar query. An example query could be structured as follows: Voice dialog system: "I love movies, do you?"
[0023] User: "Yes, absolutely."
[0024] Voice dialogue system: "What's your favorite movie? Mine is The Big Lebowski." User: "Yeah, cool. But I'm more of a Rocky fan."
[0025] Information to be extracted: Favorite film "Rocky".
[0026] If an initial query yields insufficient user information, where "insufficient" means a quantity of tokens or words below a certain threshold (e.g., less than one to ten tokens), a similar query can be posed to the driver. This similar query may involve generating and posing a query that is at least partially syntactically and / or semantically similar.
[0027] The "similar query" can be achieved by generating a probability distribution for different queries within a speech dialog query. The speech dialog system can then output the query with the highest probability first. This can be implemented using a specially configured Naive Bayes model.
[0028] By using similar queries, the speech dialogue system can refine the dialogue and / or ensure that the information provided is relevant and useful.
[0029] Extraction can be achieved using at least one of the aforementioned natural language pattern recognition techniques. The intended approach is to generate the feedback as text using automatic speech recognition (ASR). Extraction is then performed using a machine learning model (LLM). This ensures that complex dialogues between the speech dialogue system and the user are captured. Alternatively, the feedback can undergo preprocessing, specifically tokenization, lemmatization, and part-of-speech tagging, to retain only relevant tokens. Named Entity Recognition (NER) has proven particularly advantageous in this regard. Such an NER model analyzes the preprocessed text and identifies entities that correspond to specific categories, such as people, places, organizations, and / or times.This is typically achieved through the use of machine learning algorithms such as Conditional Random Fields (CRF), Hidden Markov Models (HMM), or approaches like transformer-based models. The identified entities can be classified (using the LLM) according to their category. For example, personal names are assigned to the category "Person" and / or place names to the category "Place." This classification allows the extracted information to be structured in a categorized form. The extracted entities and their categories can then be stored in the database, which contains corresponding data fields and / or tables and / or records to hold this information. Thus, each data field and / or record can be assigned to a category or entity.
[0030] The term "filling" refers to filling or replenishing the database with user information.
[0031] The term "voice dialogue system" can refer to, for example, a voice assistant and / or an in-car voice control system.
[0032] "User information" can refer in particular to preferences and / or personal data and / or behavioral data and / or interests and / or location data and / or voice data and / or relationship data and / or device and usage data and / or user status data. "Relationship data" refers to a relationship or social connection, e.g., friendly and / or professional and / or familial, between at least two users, which can be derived, for example, from the telephone directory.
[0033] Providing a database in a plain text format such as JSON or XML offers several advantages. These formats enable a flexible data structure that can be easily adapted to new requirements without changing the database schema. Their readability and interoperability facilitate development and integration with other systems. After the user provides feedback, the voice dialogue system automatically extracts the user information from the response. This automated extraction can improve the quality of the user information already collected. Finally, the extracted user information can be stored in at least one categorically appropriate data field. This enables structured and organized data management and / or improves data integrity. Additionally, the invention offers the advantage of scalability and / or unlimited extensibility with regard to data collection.
[0034] The invention also includes further developments that result in additional advantages.
[0035] According to a further development, it is provided that a control or action is assigned to the user information and that the control is triggered upon recognition of a completeness criterion. First, the speech dialogue system can actively extract user information using the inventive method, i.e., by querying the system. For example, the speech dialogue system records the culinary preferences of each user in the vehicle, and as soon as user information regarding the culinary preference has been recorded for all relevant users, the control can be activated. The control can then include filtering restaurants based on the group's shared preferences. Restaurants that do not match the specified criteria can be excluded, while suitable options are displayed. Further specifications can include a (predefined) dietary restriction (e.g.,This may include (e.g., gluten-free) and / or budget limits and / or preferred location. It may be necessary to first ask the driver or group whether and what action should be taken, so that the action is only initiated upon positive feedback from the user.
[0036] Overall, the training can include the possibility that, for example, a web crawler, based on user information extracted from at least one vehicle occupant, automatically filters relevant activities and / or locations and / or suggestss them as route stops. If there is more than one vehicle occupant, the system can also consider an overlap of their preferences.
[0037] This simplifies decision-making processes in groups, as the voice dialogue system automatically offers a selection of options that correspond to the shared preferences of all group members. This can, for example, enable a conflict-free journey.
[0038] A further development provides that a data field is assigned to a predefined category or dynamically assigned to a newly created category, with the query being formulated based on at least one category. The category can, in particular, comprise one or more preferences of the driver or another vehicle occupant, such as "favorite food" and / or "hobbies" and / or "taste in music." Such categories can either be predefined or dynamically created, for example, by initiating a conversation with the user (using the query) in the vehicle and categorizing the conversation (as a conversation phase). The database according to the invention can thus be considered a dynamically managed list whose contents in the data fields can be overwritten and / or supplemented. This can be implemented using natural language pattern recognition from the prior art.
[0039] One advanced feature envisions the voice dialogue system receiving a request from the driver or a vehicle occupant and generating a query based on that request, which is then presented to the user. In other words, it could be implemented, for example, that the driver initiates the dialogue with the voice dialogue system, and the system then generates a query based on their request. This would allow the driver to independently add and / or overwrite content, such as information related to their preferences.
[0040] A request can refer to the statement or question that a user addresses to the speech dialogue system via spoken language (direct request) and / or to at least one other user (indirect or implicit request).
[0041] When a user submits a request via an interface of the speech dialog system, this request can be classified using at least one specially trained language model (LLM) and assigned to at least one data field. This is particularly useful when the request has not yet been assigned to a data field and / or is not yet defined. This offers user-friendliness and fast processing, as the user can formulate their request in a natural way.
[0042] At least one data field or record can therefore be pre-assigned to a class or category, or dynamically assigned to one. For example, if it is detected that the user is talking about music (during a conversation) and the category "Music" is not yet assigned to a data field, or if a data field does not yet have such an assignment, a new data field with this assignment can be created. User information assigned to the category "Music" can then be added as content to the "Music" data field.
[0043] The speech dialogue system can therefore respond to the user's request by asking for additional details to better understand the request and / or provide further information.
[0044] A training program stipulates that data storage is initiated or triggered when a predefined set of user information and / or a set of topic-specific (or categorically recognized and assigned) user information is detected and / or received. "Predefined set" could mean, for example, that a certain minimum number of data points or tokens, or a complete set of information, must be present before the information is stored in one or more data fields. The set might be considered complete, for example, when at least 2 to 80 data points have been detected and / or received, where "data points" refers to individual extracted entities and / or tokens. This approach offers the advantage that user information is only collected and / or stored once a predefined set of data points has been reached.This avoids the need to create and store a data field and / or a data record and / or a database for only a small amount of user information, which may not be relevant anyway.
[0045] Further training stipulates that the speech dialogue system must have at least a Large Language Model (LLM) that includes at least a prompt manager and a dialogue manager, whereby the prompt manager uses pattern recognition to identify a category of the (currently necessary or appropriate) conversation phase and, depending on the identified category, produces corresponding context for the dialogue manager, whereby the dialogue manager includes a communication unit visible to the user and creates a query based on the identified category.
[0046] The LLM is the essential component of a speech dialogue system capable of understanding and generating natural language. Examples of LLMs include GPT-4 (Generative Pre-Trained Transformer) from OpenAI® and / or BERT (Bidirectional Encoder Representations from Transformers) from Google®.
[0047] Depending on the detected conversation phase, the prompt manager can generate appropriate context. For example, if the user asks for a restaurant in their vicinity, the prompt manager can determine the user's location, and this information is then used by the dialogue manager to communicate or provide the requested information to the user.
[0048] The dialog manager can act as an interface to the user and controls interactions. It can ensure that the contexts generated by the prompt manager are presented in a user-friendly way. To this end, the database can store user information about the user's recent queries, and the dialog manager can then use this information to ask follow-up questions (further queries). "Context" can refer to the collection and use of all relevant information that the speech dialog system needs to output an appropriate query via the dialog manager.
[0049] According to a further training, the speech dialog system is designed to create a new database, new data fields, and / or at least one new data record upon recognizing a new user. This can be achieved by monitoring the interactions and input of at least one user. Such a monitoring system can be derived from existing technology. For example, a biometric identification method and / or an authentication system can be implemented. As soon as a new user is recognized, the speech dialog system can automatically initiate the process of creating new data records, new data fields, and / or at least one new database for that user.
[0050] These new data fields can be completely empty or already contain information, such as the user's biometric characteristics, based on which the user can be, or has already been, identified. An example of the technical implementation could be a script file that is executed when a new user is detected. This script file can establish a connection to the database and create new tables or data fields and / or at least one new record for the new user.
[0051] This enables flexible and scalable management of user data and helps to ensure a personalized and user-centric experience for every user.
[0052] A beneficial further development approach involves the speech dialogue system incorporating logic that identifies existing and at least partially similar user information by comparing input data with the user information stored in the data field, thus avoiding data field duplication. "Partially similar" means that user information is similar to a certain percentage. Word embeddings can be used additionally or alternatively to check for semantic similarity.
[0053] Duplicates can be identified by comparing the input data with the user data stored in the database. This can be achieved using the LLM (Language Lifecycle Management). Alternatively or additionally, a similarity analysis (e.g., cosine similarity and / or Jaccard similarity) can be applied to textual data, while numerical or categorical data can be compared directly. Once duplicates are identified, the speech dialog system can take appropriate measures to prevent data field duplication. This might mean that no new records are created, but rather existing records or data fields are updated or extended to ensure data consistency. The speech dialog system can provide the user with feedback on the identification of duplicates and, if necessary, request confirmation before taking further action.Integrating this logic into the speech dialogue system ensures that data consistency is maintained and / or duplications are avoided.
[0054] Additionally or alternatively, it may be provided that one or more data fields are overwritten if it is detected that the driver or another vehicle occupant or user has changed their preference regarding a previously defined category. For example, if the data field with the category "Favorite Food" contains the word or content "Pizza" and the driver or user states that their favorite food is now "Burger," "Burger" could be prioritized and / or the content "Pizza" could be deleted or overwritten.
[0055] Further training stipulates that the voice dialogue system must be connected to at least one other system, so that queries from that other system are initiated via the voice dialogue system. In other words, the voice dialogue system can exchange data with one or more systems. This allows, for example, the testing of in-vehicle improvements by asking the driver questions about a function or feature in the vehicle. Specific examples include: "How do you like the steering wheel?" or "Is the seat comfortable?" In particular, a query about a function or feature can be asked if the driver used it shortly beforehand, for example, 10 seconds to 10 minutes prior.
[0056] By connecting to at least one other system, the voice dialogue system can dynamically capture the user's preferences (regarding vehicle functions).
[0057] According to a training course, the speech dialogue system is designed to include one or at least one language learning module (LLM) trained on at least one text dataset to capture speech patterns across domains. After training, the LLM is fine-tuned to suit one or more specific domains. The text dataset can originate from various sources, such as books, articles, websites, and chat histories. This training allows the LLM to become familiar with a broad spectrum of speech data.
[0058] "Cross-domain" can refer to the LLM's ability to capture and / or assign language patterns from different areas or subject fields.
[0059] The LLM is designed, for example, to:
[0060] • to monitor a dialogue between the speech dialogue system and the user and / or
[0061] • to categorically populate the database with user information and / or
[0062] • to create at least one additional database and / or at least one additional data record and / or at least one additional data field and / or to display the content of the data field and / or the information about which data field was created, and / or
[0063] • to read filled data fields and / or at least to pose a query to the user depending on a recognized category of a conversation phase and / or to receive a request.
[0064] The listed functions of the LLM can lead to an improved user experience and / or enable personalized and / or context-sensitive interaction with the user.
[0065] One advanced training program envisions the speech dialogue system performing state monitoring, quantifying a user's readiness to communicate using a predefined scale and adjusting the interaction accordingly. Depending on the user's readiness, for example, fewer queries can be asked (e.g., only one to three per hour) to make the interaction less intrusive. Conversely, with a high readiness, more questions are asked (e.g., four to thirty per hour) to gather more user information. Readiness to communicate can be quantified, for example, using state-of-the-art emotion recognition. Additionally or alternatively, it can be quantified or determined using state-of-the-art speech analysis and / or by evaluating interaction patterns.Speech analysis can, for example, evaluate changes in tone of voice or speaking speed, which can provide clues about a user's willingness to communicate. "Interaction patterns" can refer to, for example, the time a user takes to respond to a query and / or how many queries the user makes or answers, which can be an indicator of their willingness to communicate. This can increase comfort and / or the user experience, thereby enabling a conflict-free journey.
[0066] Further training stipulates that condition monitoring includes: evaluating:
[0067] Number of users within a specified radius, e.g. up to 3 or 5 meters, around the voice dialogue system, e.g. by means of an infrared sensor and / or facial recognition and / or evaluation of health data of at least one user and / or detection of fatigue of at least one user.
[0068] The voice dialogue system can be connected to wearables such as smartwatches or fitness trackers that measure health data like heart rate and blood pressure to assess the user's stress level. For example, an elevated heart rate (e.g., above 100 bpm) and / or elevated blood pressure (e.g., above 120 / 80 mmHg) can indicate increased stress, potentially reducing the user's willingness to communicate. If the driver's smartwatch reports an elevated heart rate, the voice dialogue system registers this as a possible stress indicator and adopts a calmer and / or less demanding communication style.
[0069] Health data can be transmitted to the voice dialogue system via a wireless connection (e.g. Bluetooth and / or WLAN, Wireless Local Area Network).
[0070] For example, cameras and infrared sensors installed in a vehicle can monitor the eye movements and / or blinking of occupants to detect signs of fatigue. Upon detecting signs of fatigue, the voice dialogue system can adjust its communication by asking less complex and / or shorter questions (e.g., consisting of a maximum of 5 to 20 words) or even suggesting a break.
[0071] If facial recognition detects, for example, that there are three other passengers in the vehicle besides the user, the voice dialogue system can issue queries or information louder than usual (e.g. 5 to 10 decibels louder) to address all users in the vehicle.
[0072] The voice dialogue system can therefore adapt to the current state of at least one user. The state monitoring can include a monitoring LLM that communicates with the
[0073] LLM communicates and / or exchanges data with this.
[0074] The speech dialogue system can use a feedback system to learn whether the user exhibits a low, medium, or high willingness to communicate when stress levels are elevated and / or fatigue is detected.
[0075] This enables the speech dialogue system to respond to the user's willingness to communicate in a variety of ways and to adapt the interaction accordingly.
[0076] The voice dialogue system can be installed in a motor vehicle. The motor vehicle can be a car, in particular a passenger car or truck, or a bus or motorcycle.
[0077] For use cases or application situations that may arise during the procedure and are not explicitly described here, it may be provided that, according to the procedure, an error message and / or a request for user feedback is issued and / or a default setting and / or a predetermined initial state is set.
[0078] The invention also includes the control device for the speech dialogue system. The control device can comprise a data processing device or a processor circuit configured to perform an embodiment of the method according to the invention. In particular, the processor circuit can include the LLM (Language Learning Module). For this purpose, the processor circuit can comprise at least one microprocessor and / or at least one microcontroller and / or at least one FPGA (Field Programmable Gate Array) and / or at least one DSP (Digital Signal Processor). In particular, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or an NPU (Neural Processing Unit) can be used as the microprocessor. Furthermore, the processor circuit can comprise program code configured to perform the embodiment of the method according to the invention when executed by the processor circuit.The program code can be stored in a data memory of the processor device. The processor device can be based, for example, on at least one circuit board and / or on at least one SoC (System on Chip).
[0079] As a further solution, the invention also includes a computer-readable storage medium comprising program code which, when executed by a computer or a computer network, causes it to execute an embodiment of the method according to the invention. The storage medium can be provided at least partially as a non-volatile data storage medium (e.g., as flash memory and / or as an SSD - solid state drive) and / or at least partially as a volatile data storage medium (e.g., as RAM - random access memory). The storage medium can be located within the computer or computer network. However, the storage medium can also be operated, for example, as an app store server and / or cloud server on the internet. The computer or computer network can provide a processor circuit with, for example, at least one microprocessor.The program code can be provided as binary code, assembly code, source code in a programming language (e.g., C), or a program script (e.g., Python). Alternatively, the computer-readable storage medium can be implemented as a signal containing computer-readable data, such as a time-varying voltage signal or a radio signal.
[0080] The invention also includes combinations of the features of the described embodiments. The invention therefore also includes realizations that each exhibit a combination of the features of several of the described embodiments, provided that the embodiments have not been described as mutually exclusive.
[0081] Exemplary embodiments of the invention are described below. Figure 1 shows a schematic representation of one embodiment of the method according to the invention.
[0082] Fig. 2 shows a flowchart of an embodiment of the method according to the invention.
[0083] The exemplary embodiments described below are preferred embodiments of the invention. In these exemplary embodiments, the described components each represent individual features of the invention, which can be considered independently of one another and each further develops the invention independently. Therefore, the disclosure is intended to include combinations of features of the embodiments other than those shown. Furthermore, the described embodiments can also be supplemented by further features of the invention already described.
[0084] In the figures, identical reference symbols denote functionally equivalent elements.
[0085] Figure 1 shows a speech dialog system 10, a database 8, an interface 6, an LLM 11, a state monitor 12, a plaintext 13, a prompt manager 14, a dialog manager 15, a monitoring LLM 16, a monitor 17, a user 2, and one or more systems 4. The respective components are represented symbolically and show only one possible embodiment of the idea.
[0086] According to Figures 1 and 2, a method for automatically populating a database 8 with user information using a speech dialogue system 10 can be carried out as follows. In step a., a database 8 containing one or more data fields can be provided. Then, in step b., a query in the form of natural language information can be generated and presented to the corresponding user 2 by the speech dialogue system 10 to collect user information. This can be implemented according to predefined categories or dynamically, by recognizing a conversation phase or topic, e.g., in the interior of a vehicle. Subsequently, in step c., feedback from user 2 can be received, and user information can be extracted from the feedback, in particular using the LLM (Language Lifecycle Management). Finally, in step d.,The user information extracted by the speech dialogue system 10 is stored in at least one data field.
[0087] Steps a. and b. can be represented by step S10, while step c. can be represented by step S20 and step d. by step S30.
[0088] According to a particular embodiment, a (dynamic) database 8 containing user information can be automatically populated using an (intelligent) speech dialog system 10. Categories for data fields and / or content for data fields can be freely formulated and entered into the database 8 by a user 2, with the system or speech dialog system 10 automatically filling in all data fields at runtime.
[0089] The system can consist of or include the following components:
[0090] - Plain text 13 and database 8 for structured content: A database 8 in a plain text format 13 such as JSON. The database 8 can contain or generate categories for data fields that are already known to the system, as well as further categories for data fields that have been requested by third parties or other systems 4.
[0091] - Interface 6 for defining new (free) data fields or categories for data fields externally: To avoid data field duplication and / or create a uniform structure, logic using a Large Language Modeling component or an LLM 11 can be placed in front of database 8. This logic receives freely formulated queries from the outside and automatically inserts them into database 8. The logic can also receive queries and return the correct data fields to the requester.
[0092] - Speech Dialogue System 10 (LLM-based), which uses the textual description, e.g., existing user information in database 8, to query further user information. The Speech Dialogue System 10 can consist of or comprise at least one Large Language Model 11, at least one Prompt Manager 14, and at least one Dialogue Manager 15 or dialog-leading model. The Prompt Manager 14 can recognize the currently required conversation phase and topic and create corresponding context (prompts) for the dialog-leading model. Additional LLMs or rule-based control systems can be implemented here to further monitor and control the dialog management.
[0093] - The speech dialogue system 10 can proactively store new user information in the dynamic database 8. This can occur, for example, if user 2 accidentally reveals information about themselves, or if a basic set of information is defined in advance, which is to be collected, for example, during an initial familiarization phase between the system and user 2.
[0094] - The speech dialog system 10 can be configured to understand which data is already available and therefore does not query data fields twice.
[0095] The idea therefore includes the intelligent combination of database 8 for user information and speech dialogue system 10, which makes it possible to obtain any user information about the customer or user 2 via a simple interface 6.
[0096] Another embodiment provides the following:
[0097] - Third parties (apps, services, ... ) can access the database 8 for user information via an interface 6.
[0098] - The customer can communicate using the voice dialogue system 10.
[0099] - The customer can be monitored by a condition monitoring system 12, i.e., subjected to monitoring 12.
[0100] In detail, the components can then function as follows:
[0101] - Database 8 for user information can only be accessed externally by third parties via an interface 6. Customer data can be stored in a plain text database 8 (e.g., JSON).
[0102] - A (specially) trained artificial neural network (Large Language Model 11) can be set up to populate and read the database 8 by creating data fields or receiving existing data fields as a prompt (command).
[0103] - The neural network or LLM 11 can also generate information relevant to the speech dialogue system 10 and also monitors the dialogue between system and customer in order to generate new data fields.
[0104] - The speech dialog system 10 can consist of two specially trained Large Language Models 11 in the Kem: a Prompt Manager 14 and a Dialog Manager 15.
[0105] - The Dialog Manager 14 can communicate with Database 8 or User Information Database 8 and creates special prompts accordingly.
[0106] These prompts can be captured by the dialogue generation model (Dialogue Manager 15). The dialogue generation model can be the externally visible communication unit for the customer.
[0107] - The condition monitoring 12 can calculate or quantify a “Willingness to Communicate” or communication readiness and monitors the indoor condition, such as the number of people, health data and fatigue.
[0108] - Another trained Large Language Model can communicate with the monitoring unit or monitoring LLM 16 and also generates prompts for the LLM 11 in the speech dialog system 10 with which the customer interacts, so that the status monitoring 12 also has an influence on the dialog management.
[0109] Overall, the examples demonstrate how an intelligent and generic module for obtaining user information can be provided. Reference list
[0110] 2 users
[0111] 4 systems 6 interfaces
[0112] 8 Database
[0113] 10 Voice Dialogue System
[0114] 11 LLM
[0115] 12 Condition monitoring 13 Plain text
[0116] 14 Prompt Managers
[0117] 15 Dialogue Managers
[0118] 16 Monitoring LLM
[0119] 17 Monitoring
Claims
1. 24 PATENT CLAIMS 1. A method for automatically populating a database (8) with category-related user information using a speech dialogue system (10), comprising the following steps: a. Providing a database (8) having one or more data fields, b. According to a recognized category of a conversation phase: generating and posing a query, in the form of natural language information to the user (2) by the speech dialogue system (10) to collect user information, wherein the query is at least partially related to the recognized category, c. Receiving feedback from the user (2) and extracting user information from the feedback, d. Storing the user information extracted by the speech dialogue system (10) in the at least one data field.
2. Method according to claim 1, wherein a control is assigned to the user information and the control is triggered upon recognition of a completeness criterion of the user information, wherein the control is configured according to the user information.
3. Method according to one of the preceding claims, wherein the query is formulated based on the at least one category.
4. Method according to one of the preceding claims, wherein the voice dialog system receives a request from the user (2) or another vehicle occupant and, depending on the request, generates the query and sends it to the user (2).
5. Method according to one of the preceding claims, wherein the storage is triggered when a predetermined basic set of user information and / or a predetermined basic set of topic-specific User information, which has been defined in advance, is recognized and / or received.
6. A method according to any of the preceding claims, wherein the speech dialog system (10) comprises at least one Large Language Model, LLM, (11) which includes at least one prompt manager (14) and one dialog manager (15), wherein the prompt manager (14) recognizes a category of the conversation phase by means of pattern recognition and produces context for the dialog manager (15) depending on the recognized category, wherein the dialog manager (15) comprises a communication unit visible to the user (2) and creates a query based on the recognized category.
7. Method according to one of the preceding claims, wherein the speech dialog system (10) upon recognizing a new user (2) generates a corresponding new database (8) to be filled and / or new data fields and / or one or at least one new data record.
8. Method according to one of the preceding claims, wherein the speech dialog system (10) comprises logic which, by comparing input data of the feedback with the user information stored in the data field, identifies user information that is already present and at least partially similar in a predefined manner, thereby avoiding data field duplications.
9. Method according to one of the preceding claims, wherein the speech dialog system (10) is connected to at least one further system for the exchange of user information, wherein the query is initiated by the further system via the speech dialog system.
10. Method according to any of the preceding claims, wherein the speech dialogue system comprises one or more LLMs (11) trained on at least one text data set to reproduce cross-domain speech patterns to capture, whereby the LLM (11 ) is appropriately adapted to one or more domains by means of fine-tuning after training.
11. Method according to one of the preceding claims, wherein the speech dialogue system (10) performs a state monitoring (12) which quantifies the communication readiness of a user (2) on the basis of a predetermined scale and adjusts the interaction with the corresponding user (2) depending on the quantified communication readiness.
12. Method according to claim 11, wherein the condition monitoring (12) comprises: evaluating: the number of users (2) within a predetermined radius around the speech dialog system (10) and / or Evaluating health data of at least one user (2) and / or detecting fatigue of at least one user (2).
13. Processor circuit with an LLM, wherein the processor circuit is configured to perform a method according to one of the preceding method claims.
14. Speech dialog system (10) comprising a processor unit comprising program instructions which, when executed by the processor unit, cause it to perform a method according to one of the preceding method claims.
15. Motor vehicle comprising a speech dialogue system (10) according to claim 14.
Citation Information
Patent Citations
Procedures for operating a speech dialogue system and speech dialogue system
DE102019217751A1
Emotive advisory system and method
EP2140341A1
Dialogue system and control method thereof
US20230290342A1
Systems and methods for responding to natural language speech utterance
US20070033005A1
Generating Proactive Content for Assistant Systems
US20210117214A1