Method for interaction between a vehicle and a user, and vehicle

A method using AI-based algorithms to detect and evaluate user interactions in vehicles, generating tailored recommendations to improve user understanding and operation of complex vehicle functions, enhancing safety and satisfaction.

WO2026109251A1PCT designated stage Publication Date: 2026-05-28MERCEDES BENZ GROUP AG
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/080816
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-21
Filing Date
2025-10-24
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Users of modern vehicles with complex functionalities face challenges in understanding and operating available functions, leading to underutilization, incorrect operation, and increased risk of accidents due to distraction, which negatively impacts customer satisfaction.

Method used

A method utilizing a cascade of AI-based algorithms, including machine learning models, to detect user interactions, generate and evaluate action recommendations, and provide tailored assistance to enhance user interaction and functionality understanding.

Benefits of technology

Enhances user experience and safety by providing reliable, user-specific recommendations, reducing the need for manual guidance and minimizing distractions, thereby improving vehicle operation and reducing the risk of accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025080816_28052026_PF_FP_ABST
    Figure EP2025080816_28052026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for interaction between a vehicle (1) and a user (2), wherein the user (2) carries out an interaction action for interacting with the vehicle (1), and wherein the vehicle (1) detects the interaction action with the aid of detection means (3) and generates an action recommendation (4) for the user (2) on the basis of an analysis of the interaction action and outputs said action recommendation to the user (2) by output means (5) and / or directly implements said action recommendation. The method according to the invention is characterized by the following method steps: A: feeding input data (6) generated by the detection means (3) to a first machine learning model (ML1) in order to generate an interaction text (7); B: processing the interaction text (7) by means of a first large language model (LLM1) in order to generate at least one action recommendation (4); C: processing the action recommendation (4) generated by the first large language model (LLM1) by means of a second large language model (LLM2) in order to generate a probability of success (8) for the respective action recommendation (4); D: comparing the probabilities of success (8) to a defined probability-of-success threshold value; and E1: outputting and / or performing the action recommendation (4) having the highest probability of success (8), if the probability of success (8) is greater than the probability-of-success threshold value; or E2: carrying out steps A to D again with adjustment of a random factor for the first machine learning model (ML1) until step E1 is reached.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Mercedes-Benz Group AG

[0002] Methods for interaction between a vehicle and a user as well as vehicle

[0003] The invention relates to a method for interaction between a vehicle and a user according to the type defined in more detail in the preamble of claim 1, and to a vehicle for carrying out the method.

[0004] With the increasing number and complexity of functionalities provided in vehicles, it is becoming ever more difficult for vehicle users to understand which functions are available and how to operate them. For example, modern vehicles have more and more complex driver assistance systems, such as adaptive cruise control, lane departure warning, high beam assist, and the like. In addition, more and more comfort features are being integrated into vehicles, such as playing music, TV series, movies, and the like via the vehicle's infotainment system, interacting with voice assistants, and personalizing settings for the seat position, climate control, and so on.

[0005] Especially for users who are not technically savvy, the risk increases that vehicle functions will not be fully utilized or will not be used at all, even though they are available. Furthermore, the risk of operating these functions incorrectly also increases, which can negatively impact customer satisfaction. Because users may misunderstand functions or not grasp the operating concept, they may have to visit a workshop or dealer to have the function explained or to have a perceived fault rectified. In the worst-case scenario, the risk of an accident can even increase if a user focuses too much attention on operating a function while driving because they are not yet fully familiar with its operation.This can lead to driver distraction or aggression due to a poor user experience, which can affect driving behavior. Various approaches to providing interactive assistance in vehicles to improve user interaction are known from the prior art. For example, DE 10 2006 052 897 A1 and DE 102006 049 965 A1 disclose a device and a method for interactive information output and / or assistance for vehicle functions and their operation. Assistance for vehicle occupants on operating vehicle functions can be displayed via a display device. This assistance can include animations. The vehicle occupant can be detected by a camera system, allowing the vehicle to infer the user's behavior. Various output modalities for the interactive assistance are possible and can be determined automatically.

[0006] Furthermore, US patent 2021 0 023 945 A1 discloses an intelligent user manual system for vehicles, which allows users to query vehicle functions or statuses. The system recognizes the query, searches for the answer in a digital database, and displays it.

[0007] The present invention is based on the objective of providing a method for interaction between a vehicle and a user that is improved compared to the prior art.

[0008] According to the invention, this problem is solved by a method for interaction between a vehicle and a user with the features of claim 1. Advantageous embodiments and further developments, as well as a vehicle for carrying out the method, are described in the dependent claims.

[0009] A generic method for interaction between a vehicle and a user, wherein the user performs an interaction action to interact with the vehicle, wherein the vehicle uses detection means to detect the interaction action and, depending on an analysis of the interaction action, generates a recommendation for action for the user and outputs this to the user via output means and / or implements it directly, is further developed according to the invention by the following method steps: A: Feeding the input data generated by the detection means to a first machine learning model, wherein the first machine learning model is trained to generate an interaction text describing the interaction action based on the input data, and generating the interaction text;

[0010] B: Processing the interaction text by a first large language model, wherein the first large language model is trained to generate action recommendations based on the interaction text and to generate at least one action recommendation;

[0011] C: Processing the action recommendations generated by the first large language model by a second large language model, wherein the second large language model is trained to generate a probability of success for the respective action recommendation based on the action recommendations, and generating the probability of success for all processed action recommendations;

[0012] D: Comparing the success probabilities with a defined success probability threshold; and

[0013] E1: Issuance and / or execution of the action recommendation with the highest probability of success, if the probability of success is greater than the probability of success threshold; or

[0014] E2: Repeat steps A to D, adjusting a random factor for the first machine learning model, until step E1 is reached if the probability of success is less than or equal to the probability of success threshold.

[0015] The method according to the invention involves the automatic evaluation of user behavior during the execution of interaction actions using a cascade of various artificial intelligence-based algorithms and the determination of suitable recommendations for action by this cascade. Through the interaction of different AI models in a cascade, an algorithm tailored to the respective application can be used, which allows for particularly high reliability in generating suitable recommendations for action. Furthermore, the AI ​​models can be specifically trained, thereby increasing reliability with increasing operating time. The assistance provided by the respective recommendations for action makes it easier for the user to operate the vehicle or its underlying functions, thus making the vehicle safer to operate.In particular, vehicle functions that are otherwise unused or not fully utilized are made available to the user, thereby increasing the likelihood that the user will operate the vehicle with a higher degree of automation. This improves not only safety but also the user experience and comfort. Workshop or dealer visits resulting from a lack of user knowledge can be reduced or even completely avoided.

[0016] Relevant recommendations for action are generated as needed and distributed via appropriate output channels. This eliminates the need for the user to consult manuals, an app, or a program to explain vehicle functions and learn how to operate them correctly.

[0017] The vehicle can be equipped with a wide variety of sensing devices. This allows interactions to be captured via multimodal input channels. For example, the vehicle interior can be captured using one or more cameras. These can include cameras already installed in the vehicle, such as a camera integrated into the instrument cluster or the rearview mirror for driver monitoring. Such a camera can be capable of capturing infrared images. This allows for visual recording of the vehicle interior, even under adverse lighting conditions. Furthermore, the acoustic environment in the vehicle interior can be recorded using one or more microphones. This allows for the acoustic capture of interactions. Additionally, buttons or mechanical actuators can be integrated into operable devices, such as a mechanism for locking the opening position of a glove compartment, or similar devices.This allows the system to record the operation of physical devices in the vehicle interior, particularly taking into account the degree of activation. Furthermore, user actions performed via standard human-machine interfaces can be recorded and tracked. This makes it possible to determine which buttons, knobs, rotary push-buttons, sliders, or similar devices the user operates, and how they are operated, including the speed, force, and other relevant factors. Operational actions entered via a graphical user interface, such as a touchscreen, can also be recorded. Camera systems and / or depth sensors can also be used to capture operating gestures. For this purpose, the pose of the user's limbs within the vehicle interior can be tracked using motion sensors, for example.

[0018] The data capture devices allow for the recording of all possible interactions between the user and the vehicle. This includes, for example, switching vehicle functions on and off, configuring vehicle functions, manipulating physical devices such as opening the glove compartment, adjusting the interior mirror, and the like.

[0019] The interaction can also be captured implicitly or indirectly using the data acquisition system. For example, the user can communicate with other vehicle occupants or with people outside the vehicle via telephone or video conference. In such a conversation or chat, the user can comment on specific vehicle functions and their operation. These comments can be positive, negative, or neutral. For example, the user can express criticism, difficulties in using the system, misunderstandings, or suggestions for improvement. This information can be aggregated over time and evaluated at a suitable point in time. In an advantageous embodiment of the method according to the invention, data sources external to the vehicle can also be included.If the user has given their permission, messages linked to their user account in an online forum relating to the use and operation of corresponding vehicle functions can also be evaluated.

[0020] The data acquisition devices generate data, referred to below as input data, which is processed by an in-vehicle computing unit. Accordingly, local instances of the first machine learning model, the first major language model, and the second major language model are executed on the computing unit in the vehicle. Multiple physical computer systems can collaborate to implement the computing unit. The computing unit can therefore be a single unit, a network of several control units from vehicle subsystems, or a central on-board computer. As a result, the vehicle generates action recommendations related to the respective interaction. These action recommendations can be presented as text and displayed on a screen in the vehicle. The text can also be read aloud by computer-generated speech.Output devices can include, for example, displays and / or loudspeakers. On a display device, recommendations for action can be enhanced with images, pictograms, photos, videos, rendered animations, and the like. A wide variety of display devices can be used, such as screens, head-up displays, augmented reality glasses, virtual reality glasses, projectors, and the like. Control elements relevant to operation, such as a rotary knob, can be visually highlighted in the vehicle interior, for example, by direct illumination with a moving light, a lighting element assigned to the respective control element, and so on. Haptic feedback can also be provided to the user, for example, through bumps, vibrations, or similar sensations.Haptic feedback is provided via elements in the vehicle interior that are touched by the user. For example, if the wrong control is used, an initial haptic feedback signal may be given, and a second haptic feedback signal may be given when the correct control is used.

[0021] The first machine learning model is trained to generate corresponding interaction texts for the input data provided by the various data acquisition tools. To do this, interactions are recognized and classified in the input data and then assigned appropriate labels by the first machine learning model. This process of assigning appropriate labels to the interactions can also be referred to as "labeling." The various interactions are then linked into sequences, resulting in a coherent text. Such a text can be processed by large language models.

[0022] The first large language model is trained to generate action recommendations based on the processed interaction texts. Large language models (LLMs) are particularly well-suited for artificially generating content. A large language model is trained using a so-called generative pretrained transformer (GPT). Since the first and second large language models are executed locally in the vehicle, they have fewer parameters compared to large language models running on a high-performance computing cluster. This allows for timely data generation on the vehicle's hardware. The action recommendations generated by the first large language model can, for example, be in text form.However, the recommended actions can also be implemented as program code or control signals for downstream control units. This allows for direct execution of the recommendations, meaning the vehicle itself can directly implement each recommendation. These recommendations can also include images, videos, pictograms, and the like. Sounds can also be generated by the first major speech model. Therefore, recommendations can include or be based on media content.

[0023] The action recommendations generated by the first major language model are then processed by the second major language model. The first major language model preferentially generates at least two or more action recommendations, which can then be compared to determine the most suitable one. The second major language model is trained to generate a probability of success for each action recommendation. For all action recommendations fed into the second major language model, respective probabilities of success are calculated. These probabilities are then compared, and the action recommendation with the highest probability of success is either displayed to the user or automatically implemented. Each action recommendation is first compared to the success probability threshold.Only those recommendations for action that exceed the success probability threshold, in other words, that can contribute to an improvement in user interaction, are issued to the user or implemented directly. The respective success probability threshold can be defined by the developers of the inventive method on a situation-specific basis.

[0024] If, however, no profitable recommendation for action can be generated, steps A to D are repeated until a suitable recommendation can be output. In particular, the same input data set is processed again when step A is repeated. By adjusting a random factor in the first machine learning model, it is ensured that different results are produced for the same input data. Thus, a different interaction text is generated, which in turn leads to different behavior. For example, the temperature value of the artificial neural network underlying the large language model can be adjusted, or parameters for top-k sampling, top-p sampling, or similar can be changed.

[0025] Large language models are prone to hallucinations. This means that results can be generated that are inconsistent with the training data used, i.e., nonsensical or even "invented." With regard to the method according to the invention, the first large language model could therefore issue a recommendation for action that is not feasible. Since the recommendations of the first large language model are checked for their probability of success by the second large language model, the output of such erratic recommendations to the user can be prevented. A correspondingly low probability of success value is thus determined for such a hallucination-based recommendation.

[0026] The initial training of the AI ​​models used in the inventive method, i.e., the first machine learning model as well as the first and second large language models, can be carried out in the usual way. For example, the respective AI model is trained by means of supervised learning.

[0027] An advantageous further development of the method according to the invention provides that the interaction text is enriched with a problem parameter during or after its generation, representing the extent to which the interaction action was problematic for the user; after step A, the problem parameter is compared with a defined problem parameter threshold; and steps B to E are only executed if the problem parameter is greater than the problem parameter threshold.

[0028] Using the problem parameter, the output of action recommendations to the user can be limited to cases where such recommendations are not obvious to the user, or action recommendations are only generated if the underlying interaction was particularly problematic. This prevents the user from being distracted or disturbed by unnecessary action recommendations. The problem parameter can be generated by the first machine learning model during the generation of the interaction text. In this case, the problem parameter is generated during the process of generating the interaction text. However, it is also conceivable that the input data or the interaction text generated by the first machine learning model is processed by a further algorithm that generates the aforementioned problem parameter.In particular, a second machine learning model can be used for this purpose, which is trained accordingly.

[0029] The execution of the procedure is preferably terminated if: after a specified number of executions of step E2, step E1 is still not reached; upon re-execution of steps A to D, the probability of success decreases from iteration to iteration; and / or after re-execution of step A based on current input data, the problem parameter is less than or equal to the problem parameter threshold.

[0030] Depending on the situation, it may be impossible to generate a suitable recommendation for the user that would improve their interaction behavior. In such a case, the process is terminated to avoid issuing unnecessary or misleading recommendations.

[0031] If the procedure is terminated based on the probability of success, various rules can be defined to determine the extent to which the probability of success is reduced. For example, the minimum number of iterations to be performed can be fixed, such as at least five. A measure of the decrease in the probability of success from iteration to iteration can also be specified, such as a 10% reduction per iteration or the asymptotic approach to a reduced value. Alternatively, the termination can be triggered by falling below a critical threshold for the probability of success, such as falling below 50%.

[0032] It is also conceivable to collect current input data, i.e., to track user behavior over time. This allows, taking the problem parameter into account, the system to recognize that the user has already found a solution for the appropriate operation of the vehicle function. In this case, issuing further recommendations for action is no longer necessary.

[0033] A further advantageous embodiment of the method according to the invention provides that the input data is temporarily stored in a ring buffer within the vehicle. This enables efficient data processing, as irrelevant data is deleted after a certain period of time or after a certain amount of data has been written. Furthermore, it ensures data privacy. The input data is removed from the ring buffer at a specific time, preventing personal data from being accessed and misused for other purposes after that point.

[0034] For example, a sufficient amount of input data can be written to the ring buffer to track the user's interaction behavior for the past minute. This also allows for the identification of comparatively long-lasting interactions. If there is no need to generate corresponding recommendations for action during this period, this now irrelevant data is overwritten by newly generated input data.

[0035] According to a further advantageous embodiment of the method according to the invention, it is further provided that a set consisting of a corresponding interaction text, a recommendation for action suitable for step E1, and a subsequent interaction text generated for the subsequent interaction action immediately following the interaction action underlying the interaction text is transmitted to a central computing unit external to the vehicle for further analysis, wherein a case evaluation parameter is generated in the further analysis, and the case evaluation parameter is supplied as a reward value to a local instance of the first machine learning model, first large language model, and / or the second large language model executed on the central computing unit for further training, wherein in the training the first machine learning model, the first large language model, and the second large language model are adapted in order to maximize the reward value.By continuously retraining the first machine learning model, the first large language model, and the second large language model, the reliability of the inventive method for generating action recommendations can be continuously increased. This retraining takes place on the central, vehicle-external computing unit, for example, in the form of a cloud server or server cluster. The AI ​​models trained on the central computing unit can then be distributed to the vehicles in a fleet. By taking into account the data generated in the individual vehicles of the fleet, comprehensive and therefore particularly effective training is possible. Due to the large amount of data, the probability of generating suitable action recommendations for a wide variety of scenarios is increased.The case evaluation parameter can be set manually on the central computing facility, for example by a developer.

[0036] The subsequent interaction action describes the user's reaction to the action recommendation or the user's behavior after the action recommendation has been issued. For example, it can be determined that the user correctly implements the action recommendation, thereby operating the vehicle function as intended. However, it could also happen that the user implements the action recommendation incorrectly or that the action recommendation does not lead to the desired result. This subsequent interaction action can be described by a corresponding subsequent interaction text, thus enabling the evaluation of the success of the execution of the method according to the invention.

[0037] According to an advantageous embodiment of the method according to the invention, the sentence is processed by a third machine learning model, which is trained to generate the case evaluation parameter based on the interaction text, the action recommendation, and the subsequent interaction text. This allows the continuous retraining of the aforementioned machine learning models to be automated. Thus, manual case evaluation is no longer necessary. Preferably, the third machine learning model is also a large language model.

[0038] According to a further advantageous embodiment of the method according to the invention, it is also provided that the user enters the case evaluation parameter directly in the vehicle. This allows the user to evaluate the interaction with the vehicle as well as the action recommendation generated by the vehicle. This information can then be taken into account for future generations of action recommendations.

[0039] A further advantageous embodiment of the method according to the invention provides that the user is uniquely identified and user-specific recommendations for action are generated. Different vehicle users can be identified in a proven manner. For example, the user can log in to the vehicle's infotainment system with a personal username and password. Alternatively, a characteristic identification feature can be identified, such as a code carried in a key fob, a transponder, or a smartphone app. A user can also be identified using biometric features such as an iris scan, a facial scan, voice matching, or the like.

[0040] Generating user-specific recommendations is possible, for example, through tailored training of the described AI models for each individual user. This allows for the provision of differently trained AI models for different users. Consequently, the resulting recommendations can be evaluated differently by different users, leading to further training of the respective AI models for each user. User-specifically trained AI models can then be transferred across vehicles. If a specific user operates different vehicles, for instance, the transfer of a user-adapted AI model to each other vehicle is enabled. Local instances of the first machine learning model, the first large language model, and the second large language model can be transferred between vehicles via the central computing unit.According to a further advantageous embodiment of the method according to the invention, it is further provided that at least one interaction-action-specific problem parameter threshold is specified and / or at least one adaptive problem parameter threshold is specified. By providing different problem parameter thresholds, the various user requirements when operating different vehicle functions can be met. Thus, a limit can be specified for different vehicle functions above which action recommendations are issued. For two functions that are equally difficult for the user to operate, it can therefore be achieved that an action recommendation is issued for the first vehicle function, but not for the second. In particular, the problem parameter thresholds can be specified by the user.

[0041] Problem parameter thresholds can also be adaptive. This means that the value of the problem parameter threshold can be changed. For example, if it is detected that the user repeatedly cancels the output of a corresponding action recommendation, the corresponding problem parameter threshold can be raised to prevent the action recommendation from being displayed again in this case, or to postpone it to a situation in which the user is identified as having an even greater difficulty. The problem parameter threshold can also be lowered, for example, if no action recommendation is displayed and it is further observed that the user exhibits erratic operating behavior.

[0042] According to the invention, a vehicle comprising a sensor, an output device, and a processing unit is configured to carry out the method described above. The vehicle can be any road vehicle, such as a car, truck, van, bus, or the like. It could also be a rail vehicle, watercraft, or aircraft. Suitable sensor devices include cameras, microphones, radar sensors, touch-sensitive control surfaces, motion sensors, and the like. Suitable output devices include, for example, display devices, loudspeakers, actuators, and the like. The processing unit serves to process input data generated by the sensor devices and to control the output devices.Accordingly, the first machine learning model, the first large language model, and the second large language model are executed by the computing unit. The computing unit thus has at least read access to a computer-readable storage medium containing a computer program product that includes machine-interpretable instructions, the execution of which by a processor of the computing unit causes the processing unit to provide the method according to the invention.

[0043] The corresponding machine-interpretable instructions required for further training of the respective AI models can be executed by the central computing unit.

[0044] Further advantageous embodiments of the inventive method for the interaction between a vehicle and a user as well as of the inventive vehicle also result from the exemplary embodiments which are described in more detail below with reference to the figures.

[0045] This shows:

[0046] Fig. 1 is a schematic side view of a vehicle according to the invention; and Fig. 2 is a flowchart of a method according to the invention for interaction between a user and the vehicle.

[0047] Figure 1 shows a side view of a vehicle 1 according to the invention. In the embodiment shown in Figure 1, a user 2 attempts to open the glove compartment as an interaction action and shakes an operating lever of the glove compartment for this purpose.

[0048] Vehicle 1 includes various data acquisition devices 3, such as a camera, a microphone, a touch-sensitive display, and the like. These devices 3 enable the user 2's interaction behavior to be captured and recorded. Input data 6, generated by the data acquisition devices 3 and shown in Figure 2, is processed on a computing unit 10 within vehicle 1 to generate action recommendations 4 for user 2. These action recommendations 4 can be output via output devices 5, such as the aforementioned touch-sensitive display, a loudspeaker, or the like, and / or implemented directly by vehicle 1. Furthermore, vehicle 1 has a communication connection, provided by a telecommunications unit 11, to an external central computing unit 9. This central computing unit 9 is, for example, a cloud server or server cluster.Data generated by the inventive method, insights gained therefrom, and the AI ​​models used for processing can be transferred to the central computing unit 9 or are available there in a local version. This allows the respective AI models to be further trained on the central computing unit 9 and subsequently distributed to the vehicles 1 of a vehicle fleet.

[0049] The process according to the invention is explained with reference to Figure 2. In step 201, the behavior of the user 2 is recorded using the detection means 3, and corresponding input data 6 is generated. For example, the shaking of the glove compartment is detected in camera images, as is a corresponding rattling noise recorded by the microphones. In step A, the input data 6 are fed to a first machine learning model ML1. The first machine learning model ML1 is trained to generate an interaction text 7 describing the interaction action of the user 2 with the vehicle 1, based on the input data 6.

[0050] In step B, this interaction text 7 is fed into a first large language model, LLM1. The first large language model, LLM1, is trained to generate action recommendations 4 based on interaction text 7.

[0051] Subsequently, in step C, the generated action recommendations 4 are processed by a second large language model, LLM2. This second large language model, LLM2, is trained to generate a probability of success 8 for each underlying action recommendation 4, based on those recommendations.

[0052] In step D, the generated success probabilities 8 are compared with a defined success probability threshold. If the success probability 8 is greater than the success probability threshold, step E1 is executed. In this step, the action recommendation 4 with the highest success probability 8 is output and / or executed. If, however, the success probability 8 is less than or equal to the success probability threshold, step E2 is executed. This step involves repeating steps A to D, adjusting a random factor for the first machine learning model ML1, until step E1 is reached. If this is not possible, the method according to the invention can be terminated.

[0053] For example, a camera captures that user 2 is shaking the glove compartment, but it doesn't open. From this interaction, the first machine learning model, ML1, generates the following linguistic sequence: "User tries to open glove compartment using the control lever, glove compartment doesn't open; user tries to open glove compartment using the control lever, glove compartment doesn't open; user tries to open glove compartment using the control lever and shakes the control lever, glove compartment doesn't open." User 2 is continuously observed, so interaction text 7 contains several similar descriptions of the user's action.

[0054] Advantageously, the interaction text 7 can be enriched with a problem parameter during or after its generation, representing the extent to which the interaction was or is problematic for user 2. After step A, the problem parameter can be compared with a defined problem parameter threshold. The problem parameter threshold can also be interaction-specific and / or adaptive. In this case, a particularly high value for the problem parameter is determined, so that the problem parameter exceeds the problem parameter threshold. This results in action recommendations 4 being generated in vehicle 1. As action recommendations 4, the first large language model LLM1 can, for example, generate the following recommendations, which are evaluated by the second large language model LLM2 as follows:

[0055] LLM1:

[0056] - The vehicle displays a message on the screen indicating that the glove compartment is stuck.

[0057] - Driver gets annoyed and says something like, "Oh, come on." LLM2:

[0058] - Probability of problem resolution by following the recommended course of action is negative; repeat dialogue simulation.

[0059] LLM1 :

[0060] - The vehicle activates a spotlight function to illuminate the correct button for opening the glove compartment. An acoustic signal is emitted.

[0061] - The driver uses the illuminated button, the glove compartment opens.

[0062] LLM2:

[0063] - Probability of problem solving by following the recommended action is negative, illumination in the vehicle is too bright, spotlight cannot be visually detected, acoustic signal can be misinterpreted as a warning signal and cause confusion, renewed dialogue simulation.

[0064] LLM1 :

[0065] - The vehicle displays a pop-up dialog on the screen with the content: "Do you want to open the glove compartment? Here are the options -> [LINK to user manual chapter on opening the glove compartment]",

[0066] - Driver confirms the link and selects the voice interaction option: "Hey Mercedes, open the glove compartment".

[0067] - The glove compartment is opened.

[0068] LLM2:

[0069] - Probability of problem solving by following the recommended course of action is positive; start of the generated interaction.

[0070] Since the glove compartment has now opened, the user interaction can be successfully completed. Consequently, the problem parameter now falls below the problem parameter threshold, so no further action recommendations need to be issued.

[0071] The data generated during the interaction can preferably be transmitted in an anonymized manner to the central computing facility 9 for further analysis and, in particular, for further training of the first machine learning model ML1, the first large language model LLM1, and the second large language model LLM2.

Claims

Mercedes-Benz Group AG Patent claims 1. A method for interaction between a vehicle (1) and a user (2), wherein the user (2) performs an interaction action to interact with the vehicle (1), wherein the vehicle (1) records the interaction action using detection means (3) and, depending on an analysis of the interaction action, generates a recommendation for action (4) for the user (2) and outputs this to the user (2) via output means (5) and / or implements it directly, characterized by the following process steps: A: Feeding the input data (6) generated by the acquisition means (3) to a first machine learning model (ML1), wherein the first machine learning model (ML1) is trained to generate an interaction text (7) describing the interaction action based on the input data (6) and generating the interaction text (7); B: Processing the interaction text (7) by a first large language model (LLM1), wherein the first large language model (LLM1) is trained to generate action recommendations (4) based on the interaction text (7) and generating at least one action recommendation (4); C: Processing the action recommendation (4) generated by the first large language model (LLM1) by a second large language model (LLM2), wherein the second large language model (LLM2) is trained to generate a probability of success (8) for the respective action recommendation (4) based on action recommendations (4) and generating the probability of success (8) for all processed action recommendations (4); D: Comparing the success probabilities (8) with a specified success probability threshold; and E1: Issuance and / or execution of the action recommendation (4) with the highest probability of success (8), if the probability of success (8) is greater than the probability of success threshold; or E2: Repeat steps A to D, adjusting a random factor for the first machine learning model (ML1), until step E1 is reached if the success probability (8) is less than or equal to the success probability threshold.

2. Method according to claim 1, characterized in that the interaction text (7) is enriched with a problem parameter during or after its generation, representing the extent to which the interaction action was problematic for the user (2); after step A, the problem parameter is compared with a defined problem parameter threshold; and steps B to E are only executed if the problem parameter is greater than the problem parameter threshold.

3. Method according to claim 2, characterized in that the execution of the method is terminated if: after a specified number of executions of step E2, step E1 is still not reached; when steps A to D are re-executed, the probability of success (8) decreases from iteration to iteration; and / or after re-executing step A based on current input data, the problem parameter is less than or equal to the problem parameter threshold.

4. Method according to one of claims 1 to 3, characterized in that the input data (6) is temporarily stored in a ring buffer in the vehicle (1).

5. Method according to one of claims 1 to 4, characterized in that a set consists of a corresponding interaction text (7), an action recommendation suitable for step E1 (4) and a text directly related to the The interaction text (7) underlying the interaction action is transmitted to a vehicle-external central computing unit (9) for further analysis, in which a case evaluation parameter is generated, and the case evaluation parameter is supplied as a reward value to a local instance of the first machine learning model (ML1), first large language model (LLM1) and / or the second large language model (LLM2) executed on the central computing unit (9) for further training, in which the first machine learning model (ML1), the first large language model (LLM1) and the second large language model (LLM2) are adapted in order to maximize the reward value.

6. Method according to claim 5, characterized in that the sentence is processed by a third machine learning model, wherein the third machine learning model is trained to generate the case evaluation parameter based on the interaction text (7), the action recommendation (4) and the subsequent interaction text.

7. Method according to claim 5 or 6, characterized in that the user (2) enters the case evaluation parameter himself in the vehicle (1).

8. Method according to any one of claims 1 to 7, characterized in that Users (2) are uniquely identified and user-specific recommendations for action (4) are generated.

9. Method according to one of claims 2 to 8, characterized in that at least one interaction-action-specific problem parameter threshold is specified and / or at least one adaptive problem parameter threshold is specified.

10. Vehicle (1), comprising recording means (3), output means (5) and a Computing unit (10), characterized in that the recording means (3), output means (5) and the computation unit (10) are connected to The execution of a method according to one of claims 1 to 9 is set up.

Citation Information

Patent Citations

  • Device and method for interactive information output and / or assistance for the user of a motor vehicle

    DE102006049965A1

  • Information e.g. vehicle data, device for use in motor vehicle, has display device, and processing device attached to storage that stores vehicle data, where vehicle data are displayed in partially animated manner via display device

    DE102006052897A1

  • Intelligent User Manual System for Vehicles

    US20210023945A1