Active multi-mode vehicle-mounted housekeeper

By interacting with multiple context sources through the in-vehicle management system, it enables switching between audio and non-audio channels. Combined with a touchscreen interface and display, it solves the problems of in-vehicle assistant systems being unable to proactively initiate communication and low efficiency in voice interaction. It provides proactive information services and multi-turn dialogue capabilities, thereby improving the efficiency of driver information services.

CN122003349APending Publication Date: 2026-05-08CERENCE OPERATING CO
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CERENCE OPERATING CO
Filing Date
2024-10-15
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing in-vehicle assistant systems typically cannot initiate communication proactively, and voice communication is inefficient in complex environments, making it difficult to meet the diverse information needs of drivers.

Method used

The system employs an in-vehicle concierge system that interacts with multiple context sources through the infotainment system. It uses a mode selector to switch between audio and non-audio channels, and combines a touchscreen interface and display to provide proactive information services. It also generates prompts and interaction mode selections through dynamic context to enable multi-turn dialogues and task execution.

Benefits of technology

It improves the interaction efficiency of the in-vehicle assistant system in complex environments, can proactively provide information services, reduce driver distraction, and adapt to diverse information needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122003349A_ABST
    Figure CN122003349A_ABST
Patent Text Reader

Abstract

An apparatus for providing an information service to a user in a vehicle includes an on-board steward, a prompter, and an interaction mode. An in-vehicle steward is hosted by an infotainment system in a vehicle and participates in interaction with a user based on context. A prompter generates a prompt for a model in response to an instruction from the on-board steward, the instruction being based on the contextual information, and the prompt being selected to cause the model to generate content that guides the user to select. The interaction mode delivers content to the user and provides information from the user to the on-board steward.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims the priority date of U.S. Provisional Application 63 / 543,868, which is October 12, 2023. It also claims the priority date of U.S. Provisional Application 63 / 685,291, which is August 21, 2024, and the priority date of U.S. Provisional Application 63 / 703,238, which is October 4, 2024, and the priority dates of U.S. Provisional Applications 63 / 706,114 and 63 / 706,118, which are all incorporated herein by reference. Background Technology

[0003] This invention relates to in-vehicle electronic equipment, and more particularly to electronic equipment for transmitting messages to a suitable subset of a group of vehicles.

[0004] In modern vehicles, infotainment systems with managed in-vehicle assistants are not uncommon. These assistants provide various information services to the vehicle's occupants.

[0005] The in-vehicle assistant responds to communications from occupants, but typically does not initiate such communications. Furthermore, the in-vehicle assistant usually communicates via voice. This is particularly convenient for drivers who expect to observe the vehicle's surroundings rather than read messages on a screen. This reduces driver distraction. Summary of the Invention

[0006] In one aspect, the invention features an apparatus for providing information services to a user in a vehicle. This apparatus includes an infotainment system installed in the vehicle, a context source communicating with the infotainment system, an in-vehicle assistant hosted by the infotainment system, a prompter that generates model prompts in response to instructions from the in-vehicle assistant, the instructions being context-based, and selecting prompts to cause the model to generate content guiding the user to make choices; the model being a language model; and an interaction mode for delivering content to the user and providing information from the user to the in-vehicle assistant. The in-vehicle assistant participates in the interaction with the user based on the context from the context source.

[0007] The embodiments also include those with a mode selector. In these embodiments, the mode selector selects an inbound channel and an outbound channel in response to a selection signal from the vehicle management system, the selection signal being context-based; the mode selector switches between first and second modes in response to context, each mode including an inbound channel and an outbound channel, wherein in the first mode, both the inbound and outbound channels are audio channels, and wherein in the second mode, at most one of the inbound and outbound channels is an audio channel; the mode selector switches between a first mode and a second mode in response to context, each mode including an inbound channel and an outbound channel, wherein in the first mode, both the inbound and outbound channels are audio channels, and wherein in the second mode, both the inbound and outbound channels are non-audio channels.

[0008] Furthermore, in embodiments including a mode selector, the mode selector switches between a first mode and a second mode in response to context, wherein each mode includes an inbound channel and an outbound channel, wherein in the second mode, the inbound channel includes a touchscreen interface with selections identified by content from the model, and the mode selector switches between the first mode and the second mode in response to context, wherein each mode includes an inbound channel and an outbound channel, wherein in the first mode, the outbound channel is selected to provide voice information to the user and also to provide text to the user via a display, the text being generated by the model.

[0009] Other embodiments include a mode selector that responds to context by switching between a first interaction mode and an interaction mode selected from a group consisting of a second interaction mode and a third interaction mode, wherein the first interaction mode, the second interaction mode, and the third interaction mode are selected from a group consisting of an audio mode, a non-audio mode, a mixed mode, a privacy mode, and a mute mode.

[0010] In some embodiments, the context includes information indicating a message for the user. In such an embodiment, the in-vehicle assistant causes the prompter model to generate a summary of the message to be delivered to the user.

[0011] In these embodiments, the context includes information indicating a first message and a second message to the user, as well as preference information indicating that the first message has a higher priority than the second message. In these embodiments, the in-vehicle administrator delivers the first message before the second message based on the preference information.

[0012] Other embodiments include embodiments in which the context includes one or more of the following: occupancy information, message information, location information, traffic information, driver information, vehicle information, preference information, and environmental information.

[0013] Furthermore, in this embodiment, the vehicle management system is configured to change the interaction mode based on the presence of an additional passenger in the vehicle, who is a passenger other than the user.

[0014] In other embodiments, the interaction mode is the current interaction mode selected from a plurality of interaction modes, including a first interaction mode and a second interaction mode. In such an embodiment, the current interaction mode is the first interaction mode, and the vehicle management system is configured to change the current interaction mode to the second interaction mode in response to a change in context.

[0015] The embodiments also include those in which the context includes driver information and vehicle information. In some of these embodiments, in response to the context, the vehicle manager identifies the location of facilities for modifying the driver's state and the vehicle's state.

[0016] In other embodiments, the context includes driver information indicating the driver's state. In these embodiments, based on the driver information, the in-vehicle assistant proposes interactions to facilitate changes in the driver's state, including cognitive stimuli.

[0017] These and other features of the invention will be apparent from the following detailed description and accompanying drawings, wherein: Attached Figure Description

[0018] Figure 1 A vehicle with an infotainment system is shown, which has a message interceptor that blocks incoming messages;

[0019] Figure 2 It shows Figure 1 Details of the message interceptor;

[0020] Figure 3 and Figure 4 The touch interface and display are shown in two different states;

[0021] Figure 5 It shows Figure 1 The hierarchy within the in-vehicle assistant;

[0022] Figure 6 It shows Figure 5 The output of the top-level model shown is after interaction with a domain-specific authorized party.

[0023] Figure 7 It shows Figure 5 The output of the top-level model shown is after interaction with a domain-specific licensee.

[0024] Figure 8 It shows something similar to Figure 5 The embodiment shown is modified to process task chains;

[0025] Figure 9 It shows the use of Figure 1 The user interface of the in-vehicle concierge;

[0026] Figure 10 It shows Figure 9 Alternative embodiments of the user interface; and

[0027] Figure 11 It shows Figure 10 Details of the multi-level architecture shown. Detailed Implementation

[0028] Figure 1 The vehicle 10 is shown with a passenger compartment 12, in which a user 14 and possibly one or more other vehicle occupants are seated in seats 16.

[0029] Each seat 16 has an associated speaker 18, microphone 20, camera 22, mass sensor 24, and touch interface 26. The microphone 20, camera 22, mass sensor 24, and touch interface 26 are components of sensor group 28.

[0030] The vehicle 10 also includes an infotainment system 30 that receives input from a group of sensors 28 and provides audible output via a speaker 18 and / or visual output via at least one display 32.

[0031] Now for reference Figure 2 The infotainment system 30 includes an in-vehicle assistant 34 that participates in interactions with the user 14. These interactions include dialogues with the user 14. During these dialogues, the in-vehicle assistant 34 receives context 36 from an aggregator 38 and uses this context 36 as the basis for controlling a prompter 40, which generates prompts 42 for a model 44. As used herein, "model" refers to a "large language model".

[0032] In response to such a prompt 42, model 44 outputs the content 46 to be conveyed to user 14 using one of several interaction modes 48. Mode selector 50 selects from interaction modes 48 based on mode selection switch 52 received from vehicle manager 34. The settings on mode selection switch 52 depend in part on the context 36 obtained from aggregator 38.

[0033] Aggregator 38 aggregates several different kinds of context 36. These include: occupancy information 54, which is derived from signals from sensor group 28, and in particular from camera 22 and mass sensor 24, and indicates how many occupants are in vehicle 10, where they are sitting, and in some cases, who they are; message information 56, which includes messages 58 to be relayed to user 14, including information such as who sent message 58, who else received message 58, and the content of message 58, all of which provide useful context 36 for use by vehicle manager 34; location information 60, such as location information obtained from a geolocation system; and traffic information 62, which is typically generated by remote... The traffic server provides: driver status information 64, which is obtained from observations made by cameras and microphones; vehicle information 66, such as fuel supply, tire pressure, and other information derived from sensors inside the vehicle 10; preference information 68, some of which is stored as user preferences stored locally or remotely by individual users, and some of which is obtained by observing user activities and behavioral patterns, including observing user activities through which users express or imply user intentions; and environmental information 70, which includes temperature, precipitation, atmospheric conditions, ambient lighting, and similar information about the vehicle's operating environment, as well as information about nearby points of interest.

[0034] Each interaction mode 48 includes an inbound channel 72 and an outbound channel 74. The in-vehicle assistant uses the inbound channel 72 to receive information from the occupant 10 and uses the outbound channel 74 to provide information to the user 14. As shown, the inbound channel 72 passes through the mode selector 50 toward the in-vehicle assistant 34.

[0035] Interaction mode 48 is different from each other based on the modes used for inbound channel 72 and outbound channel 74.

[0036] In the illustrated embodiment, there are three types of interaction modes 48: audio mode 76, non-audio mode 78, and mixed mode 80.

[0037] In audio mode 76, both inbound channel 72 and outbound channel 74 are audio channels. This is the useful default interaction mode 48.

[0038] In non-audio mode 78, neither the inbound channel 72 nor the outbound channel 74 are audio channels. This mode provides greater privacy because other occupants in vehicle 10 will find it more difficult to eavesdrop on interactions.

[0039] In hybrid mode 80, one of the inbound channel 72 and the outbound channel 74 is audio, while the other is not. An example of a non-audio inbound channel 72 is a channel with a touch sensor 26 at its end. An example of a non-audio outbound channel 74 is a channel with a display 32 at its end. An example of hybrid mode 80 is a mode in which the outbound channel 74 is suppressed. This is referred to as "mute mode".

[0040] Based on context 36, the vehicle management system 34 can determine that privacy is important, and therefore use one or two channels in a non-audio mode.

[0041] Already introduced Figure 2 It is helpful to consider some examples of how the structures shown operate in specific use cases.

[0042] In one example, the vehicle manager 34 receives context 36, which includes occupancy information 34 and message information 56. Message information 56 indicates to user 14 that several messages 58 are pending.

[0043] The in-vehicle management system 34 automatically draws attention to the waiting message 58 and suggests actions, such as:

[0044] "Good morning, Mr. Phelps. You have a voice message from Mrs. Phelps, a text message from the agency, and two text messages from unknown parties. Some of these may be fraudulent. Would you like to hear Mrs. Phelps's text message first?"

[0045] In this example, the vehicle management system 34 has already used occupancy information 54 to identify user 14. After doing so, it then accesses preference information 68 to determine how to prioritize messages from the identified user 14. This allows the vehicle management system 34 to engage in customized interactions with that user 14.

[0046] In this situation, the in-car assistant 34 has already discovered from preference information 68 that user 14 habitually listens to messages from Mrs. Phelps before any other messages. Unsurprisingly, user 14 replies to in-car assistant 34 in a manner consistent with his observed habit:

[0047] "Yes, please tell me what Sandra wants to say."

[0048] Like many voice messages, this message is embellished with filler words, repetitions, pauses, and similar features characteristic of real-time voice. Having recognized these features, the vehicle assistant 34 provides both message 58 and instructions summarizing message 58 to the prompter 40 before delivery via the outbound channel 74.

[0049] In response, prompter 40 generates a suitable prompt 42 and provides it to model 44 along with the message 58 to be summarized. Model 44 then generates content 46, which summarizes message 58 in a form that retains the content but delivers it more fluently. An example of such a summary has the following form:

[0050] "Sandra said the fridge was empty and suggested dinner at a French restaurant. Should I make a reservation?"

[0051] As is evident again from the foregoing, the vehicle manager 34 is configured to do more than simply wait for instructions on what to do next. In this case, based on message information 56 including the message content, the vehicle manager 34 has suggested actions to be taken.

[0052] After receiving the user's consent to the action, the vehicle management system 34 retrieves location information 60 and environmental information 70 from the context aggregator 38. Together, these reveal the presence of nearby restaurants. Therefore, the vehicle management system 34 proposes an action plan:

[0053] "I've found two French restaurants: 'Chez Fantine' and 'Le Soleil Bleu'. 'Chez Fantine' is the closest to your current location. Would you like me to take you there?"

[0054] The in-car assistant 34, which has been programmed to be proactive, once again provided specific suggestions for action, which user 14 readily accepted:

[0055] "Yes, that's a great idea. I've forgotten where it is."

[0056] The in-car management system 34 disclosed in this article does not passively and strictly follow user instructions, but proactively provides further measures:

[0057] "It would be my pleasure. Turn left onto Macondo Avenue two miles ahead. If you'd like, I can also book a table for you there and let Sandra know where to meet you."

[0058] After identifying environmental information 70 indicating that the weather would be fine that evening and that the selected restaurant had outdoor seating, the in-car assistant 34 added:

[0059] "By the way, since the weather is so nice, would you like me to reserve an outdoor seat for you?"

[0060] Due to the nature of his work at the agency and the suspicious nature of the two unidentified messages, User 14 did not wish to be exposed to passersby. Therefore, he rejected the following options:

[0061] "No, thank you, I prefer to be indoors."

[0062] After being prompted with the word "prefer", the in-car concierge 34 updates the preference information 68 to indicate that Mr. Phelps prefers indoor dining.

[0063] A few minutes later, after making a reservation at "La Maison Fantine" and after also sharing the reservation information with Mrs. Phelps, the car concierge 34 used traffic information 62 and location information 60 and said:

[0064] "Reservation complete, and Sandra has been called. Turn left onto Macondo Street half a mile ahead. Traffic is clear, and you should arrive at the restaurant in about ten minutes."

[0065] As vehicle 10 approaches, the vehicle assistant 34 recognizes its imminent arrival and proactively offers another suggestion:

[0066] "We'll be there in five minutes. Would you like me to order your usual beer? That way it can be served right away when we arrive."

[0067] During the multi-turn dialogue in the aforementioned interactions, the in-vehicle assistant 34 has utilized many different types of context 36, used model 44 to provide appropriate summaries of messages 58, and consistently guided the user 14 along a path to effectively complete the interaction. This differs from conventional digital assistants, which are inherently more passive.

[0068] A notable feature of the aforementioned interaction is its use of audio mode 76. While audio mode 76 is generally convenient in vehicle 10, there are situations where it is advantageous to pause voice interaction and use different interaction modes 48 for all or part of the interaction.

[0069] For example, traffic conditions or ambient noise may change to the point that speaking becomes too difficult. In this case, the in-vehicle assistant 34 sends a mode selection signal 52, which causes the mode selector 50 to switch to a hybrid mode 80 characterized by an inbound channel 72 terminating at the touch sensor 26 and an outbound channel 74 terminating at the display 32. Alternatively, the user 14 can initiate a change in the interaction mode 48 via verbal command. In other examples, the in-vehicle assistant 34 determines the need for privacy based on occupancy information 54 and excludes the use of the audio mode 76.

[0070] Since the changes in interaction mode 48 are somewhat unpredictable, it is particularly useful for the vehicle management system 34 to display its recent communications on display 32 and provide a menu of response options on touch interface 26 for the user to choose from. However, since the topics of recent communications are unpredictable, the responses to be displayed on touch interface 26 cannot be pre-programmed. They must be generated dynamically.

[0071] To address this challenge, while providing communication to user 14, the vehicle management system 34 also enables the prompter 40 to compose a prompt 42. This prompt causes the model 44 to generate one or more action-oriented statements to be displayed on the touch interface 26 for the user to select a topic for. The limited space of the touch interface 26 and the user's limited attention limit the length of these action-oriented statements. As a result, the prompt 42 causes the model 44 to compose action-oriented statements that convey the options presented by the vehicle management system 34 in a concise manner.

[0072] The in-vehicle assistant 34 also enables model 44 to predict the most likely course of the conversation, thereby allowing it to provide prompts to guide or instruct user 14 in their interactions. In some cases, these prompts, generated based on contextual logic, have the benefit of providing user 14 with information about options that user 14 may not be aware of.

[0073] In some cases, the vehicle assistant 34 initiates interactions that do not involve communication with another person. In one illustrated example, the vehicle assistant 34 receives vehicle information 66 indicating low fuel and location information 60 indicating no service stations within the next eighty miles as context 36. The vehicle assistant 34 also receives driver information 64 indicating the driver is experiencing an increased level of drowsiness. In response to this context 36, the vehicle assistant 34 identifies from vehicle information 66 that vehicle 10 is using diesel and user 14 needs stimulants, and uses environmental information 70 to learn about upcoming rest stops that offer diesel and various stimulant-containing beverages. After doing so, the vehicle assistant 34 makes the following suggestions:

[0074] “You look a bit tired and low on fuel. The truck service area at the Dead Man’s Valley exit has diesel and coffee. Why don’t you stop for a bit? The next service area is 80 miles away.”

[0075] Then, after resuming on the road, the vehicle assistant 34 continues to monitor the context 36, including driver information 64. When user drowsiness is detected, and considering that a moderate increase in cognitive load helps to stay alert, the vehicle assistant 34 attempts to guide user 14 through a short question-and-answer session. Therefore, the vehicle assistant 34 causes model 44 to generate questions for appropriate responses and issues the following invitation:

[0076] "How about we play a quiz game to pass the time during this long drive?"

[0077] After receiving the driver's consent, Model 44 began providing the questions it generated. When the first question was ready, the in-vehicle assistant 34 said,

[0078] "Okay, let's start the Q&A. Who scored the winning goal in the 2014 World Cup?"

[0079] In response to the driver's answer "Michael Jordan," the car assistant replied with 34 comments:

[0080] "Close! But Michael Jordan played basketball. The correct answer is 'Mario Goetz.' Next question: Who is the star player on the opposing team?"

[0081] Figure 3 The display 32 and touch interface 26 are shown after model 44 generates text 82 summarizing the choice occupant 14 needs to make, and three buttons 84 that invite the user to touch touch interface 26 to make a choice. In a typical embodiment, model 44 also generates button text 86 for the buttons 84 based on context and is subject to the constraint that they must fit the buttons 84. Touch interface 26 also displays some dynamically generated prompts 88 that alert occupant 14 to other capabilities of vehicle assistant 34. In some embodiments, dynamically generated prompts 88 also serve as buttons 82 to convey the user's intent to vehicle assistant 34.

[0082] Figure 4 This shows what happens after the voice has been received from crew member 14. Figure 3 The display 32. In this case, the display 32 displays echo text 90, which shows the occupant's words, as best understood by the vehicle butler 34.

[0083] Therefore, the in-vehicle assistant 34 described in this paper proactively provides information services to users by utilizing different types of context 36. This is done by both providing action-oriented statements to prompt users and prompting the model to provide appropriate generated content, which dynamically changes in response to changes in context.

[0084] The multi-turn dialogues described here often require the use of specialized information sources to perform specialized tasks. Figure 5 An implementation of an architecture that facilitates the ease with which the vehicle butler 34 can perform these specialized tasks is shown.

[0085] For example, in the first instance, the in-vehicle assistant 34 mentions a recently arrived text message. In order to retrieve the message and then read it aloud, the in-vehicle assistant 34 will need to know how to manipulate the message. This may require the in-vehicle assistant 34 to interface with the software that handles message transmission.

[0086] In another example described herein, the in-car assistant 34 already offers the option to make restaurant reservations. However, the in-car assistant 34 may not necessarily know how to do this. Furthermore, programming the in-car assistant 34 to do this is impractical, especially considering the diversity of existing systems and the fact that they change over time. Instead, the in-car assistant 34: identifies external applications 92A, 92B that know how to make reservations, interacts with those external applications 92A, 92B, and relays the results to user 14, in this case, Mr. Phelps.

[0087] Typical external applications 92A and 92B interact with other entities through their application programming interfaces (APIs) 94A and 94B. Therefore, in order to interact with external applications 92A and 92B, the vehicle management system 34 uses the APIs 94A and 94B of those external applications. Each external application 92A and 92B in the external application set 36 will have its own API 94A and 94B. This means that the vehicle management system 34 must somehow know how to interact with many different APIs 94A and 94B.

[0088] In cases where the external application set 36 has only a few external applications 92A, 92B, it is practical to build an in-vehicle management system 34 that knows the relevant application programming interfaces 94A, 94B. However, as the number of external applications 92A, 92B increases, the task of ensuring that the in-vehicle management system 34 can effectively use them also increases. Furthermore, the application programming interfaces 94A, 94B are not necessarily static. They are prone to change over time when developers of external applications 94A, 94B add or remove features or when they modify existing features. Therefore, the technical problem to be solved is to alleviate the burden of ensuring that the in-vehicle management system 34 can interact with the constantly changing application programming interfaces 94A, 94B.

[0089] Figure 5 The diagram illustrates the hierarchy 96 upon which the vehicle management system 34 relies for interaction with application programming interfaces 94A and 94B. Hierarchy 96 comprises a first level 98A and a second level 98B. The first level includes a top-level agent 100, which may be referred to as the “backbone” or “orchestrator” of hierarchy 96. The second level 98B includes domain-specific authorized parties 102A and 102B, which may be referred to as “domain-specific actions” or “domain-specific tools.”

[0090] In the illustrated embodiment, the first level 98A is exactly the top level of level 96. However, Figure 5 The architecture shown is inherently recursive. Nothing prevents each domain-specific licensee 102A, 102B from acting as the first-level 98A and having subdomain-specific licensees under it.

[0091] Each domain-specific authorized party 102A, 102B is configured to respond to prompts related to a specific domain. Examples of domains include: navigation domain, music domain, general knowledge domain, and vehicle control domain. Typically, a text unit can have more than one meaning. To determine which of the several meanings of the text should be applied, additional information is needed. This additional information identifies the "domain." As an example, when accompanied by information to apply the "music" domain, the term "laying tracks" would be interpreted as meaning the action of recording music. Conversely, when accompanied by information to apply the "railway" domain, the same text would be interpreted as meaning the action of laying pairs of tracks for train use. Dividing requests into domains is often useful so that a given set of words can be assigned to the appropriate meaning.

[0092] The top-level agent 100 includes the top-level hint generator 104 and the top-level model 106.

[0093] In a preferred embodiment, the top-level model 106 is a large language model, hereinafter referred to as the "model". In this case, the model receives text input and provides an output consistent with its training. In some embodiments, the output of the top-level model 106 is natural language text. However, this is not necessarily the case. For example, in some embodiments, the output of the top-level model 106 includes structured text, such as structured text forming API commands. This output can then be used to perform functions in response to text input.

[0094] Top-level suggestion builder 104 receives top-level query 108 from user 14. Top-level suggestion builder 104 uses top-level query 108, context information 36, and domain information 110A to construct top-level suggestion 112, and then provides it to top-level model 106.

[0095] Top-level hint 112 does not only respond to top-level query 108. Top-level hint 112 prompts top-level model 106 to respond with top-level model output 114, which includes: information identifying multiple domains related to processing top-level query 108, inference steps depending on identifying those domains, and an action plan for responding to top-level query 108.

[0096] The top-level agent 100 also includes a top-level query generator 126 and a top-level receiver 128.

[0097] The top-level query builder 126 uses the first output 112 to construct domain-specific queries 130A, 130B and provides them to the corresponding domain-specific authorized parties 102A, 102B identified in the action section 118 of the top-level model output 114. Each domain-specific query 130A, 130B uses information obtained from the action input section 118 of the same top-level model output 114.

[0098] For each domain-specific authorized party 102A, 102B, the domain-specific queries 130A, 130B are no different from the top-level query 108 from user 14. This means that the top-level query builder 126 does not need to know the details of the interaction with external applications 92A, 92B. This knowledge is contained in the relevant domain-specific authorized parties 102A, 102B. The domain-specific queries 130A, 130B only need to be specific enough to trigger the correct application of the knowledge contained in the domain-specific authorized parties 102A, 102B.

[0099] As a result of the aforementioned architecture, by utilizing its combined layer 96, it is possible to bypass training the top-level model 106 to output all kinds of API calls. Instead, the top-level model 106 only needs to be trained to identify which domain-specific licensees 102A, 102B will know how to generate specific kinds of API calls.

[0100] Each domain-specific authorized party 102A, 102B provides its corresponding domain-specific response 132A, 132B back to the top-level agent 100, and specifically, to the top-level receiver 128. The top-level receiver 128 weaves the domain-specific responses 132A, 132B from the different domain-specific authorized parties 102A, 102B into a coherent top-level response 134, and then provides it back to the user 14 in response to the top-level query 108.

[0101] Therefore, the top-level agent 100 performs a form of classification in which it receives the top-level query 108, identifies various experts, namely the domain-specific authorized parties 102A and 102B required to process the top-level query 108, and then provides the corresponding domain-specific queries 130A and 130B to these domain-specific authorized parties 102A and 102B.

[0102] Having discussed the structure and operation of the top-level agent 100, it is now useful to consider the structure and operation of the representative domain-specific licensees 102A and 102B. The structure and operation of the representative domain-specific licensees 102A and 102B represent all domain-specific licensees 102A and 102B.

[0103] like Figure 5 As shown, the domain-specific licensees 102A and 102B have structures very similar to those of the top-level agent 100. The domain-specific licensees 102A and 102B are characterized by domain-specific prompt builders 136A and 136B and domain-specific models 138A and 138B. Like the top-level model 106, the domain-specific models 138A and 138B are large language models.

[0104] Domain-specific licensees 102A and 102B have been trained to use application programming interfaces 94A and 94B associated with selected external applications 92A and 92B relevant to their respective domains. Therefore, domain-specific licensees 102A and 102B have been trained to adapt to a certain number of application programming interfaces 94A and 94B. However, because domain-specific licensees 102A and 102B are specific to a single structural domain, this number is not very high. Most importantly, it is not high enough to degrade the overall performance of the agent.

[0105] Domain-specific suggestion builders 136A and 136B use domain-specific queries 130A and 130B and domain information 110B to construct a domain-specific suggestion 140A, which is then provided to a domain-specific model 138A. The domain-specific model 138A provides application-specific outputs 142A and 142B to API builders 144A and 144B, which then construct application-specific API calls 146A and 146B and provide them to the relevant external applications 92A and 92B.

[0106] In response to application-specific API calls 146A and 146B, external applications 92A and 92B provide application-specific responses 142A and 142B to application-specific receivers 148A and 148B of the domain-specific authorized party. The application-specific receivers 148A and 148B then transform the application-specific responses 142A and 142B into domain-specific responses 132A and 132B that ultimately go to the top-level receiver 128, as already discussed.

[0107] The proxy layer 96 eliminates the need for the top-level proxy 100 to know the details of the various application programming interfaces 94A and 94B. Instead, the top-level proxy 100 only needs to be able to identify which domains are relevant to the top-level query 108 and specify what each domain-specific authorized party 102A and 102B needs.

[0108] Figure 6 An example of top-level model output 114 is shown, generated by top-level hint 112 constructed by top-level hint builder 104 based on occupant's top-level query 108. Clearly, in a single step, top-level model 106 responds by providing a thinking section 116 that analyzes top-level query 108 into separate and distinct first and second domain-specific tasks; providing an action section 118 that identifies a first domain-specific authorized party 102A to perform the first domain-specific task; and providing an action input section 120 that is currently empty. It has been found that by forcing top-level model 106 to provide a thinking section 116 that includes reasoning steps to reach action section 118, top-level model 106 becomes constrained to more reliably generate valid top-level hints 112.

[0109] The action input section 120 will be populated with information used to initiate interaction with the first domain-specific authorized party 102A through subsequent iterations to facilitate the return of useful information from the first domain-specific authorized party 102A. Examples of action input sections 120 include executable commands and questions to be responded to. If the input is context-dependent and contains anaphoric relations, anaphoric pronouns, implicit contextual references, and / or omissions, the input is interpreted based on the dialogue history to resolve context dependencies before being used as input for the domain-specific authorized party 102A. Preferably, the action input section 120 (which forms the basis for input for the domain-specific authorized party 102A) can always be understood as a self-contained question or command.

[0110] Figure 7 A second top-level model output 114 is shown, generated by the top-level agent 100 interacting with the first domain-specific licensee 102A in the manner indicated in the action input section 118. This second top-level model output 114 includes an observation section 121 containing information provided by the first domain-specific licensee 102A. The action section 118 identifies the appropriate second domain-specific licensee 102B, and the action input section 120 provides the basis for a domain-specific query 130B to be provided to the second domain-specific licensee 102B.

[0111] In some cases, processing a top-level query 108 requires the consecutive use of two or more domain-specific authorized parties 102A, 102B. Such a top-level query 108 defines a task chain. For example, a top-level query 108 of the form "Please send the current price of the Nol element to Septimus Selden" requires the first task to be performed by the first domain-specific authorized party 102A, which knows how to determine the price of the Nol element, and the second task to be performed by the second domain-specific authorized party 102B, which knows how to send the message.

[0112] These tasks are domain-specific tasks corresponding to different domains. The first task involves retrieving product price information. The second task involves sending a message to a specific person. The second domain-specific task depends on the result of the first domain-specific task. After all, it's impossible to send the current price of a nore element before knowing its actual current price.

[0113] In this scenario, the vehicle management system 34 performs repeated iterations. In each iteration, the vehicle management system 34 identifies the appropriate domain-specific authorized parties 102A and 102B, consults the domain-specific authorized parties 102A and 102B, and saves the results of the consultation for use as context in subsequent iterations.

[0114] After all iterations are complete, the results of the iterations are packaged into a top-level response (134) and then provided to the user. Figure 8 The illustrated embodiment performs an iterative process in which context information collected from earlier iterations informs subsequent iterations.

[0115] The processing required to generate top-level response 134 extends across a time interval. This time interval extends between the start and finish times. Between the start and finish times, top-level response 134 is early, and processing is in progress.

[0116] Figure 8 The illustrated embodiment supports the processing of a top-level query 108 that requires the use of multiple domain-specific licensees 102A, 102B (or it is conceivable to reuse a single domain-specific licensee 102A), wherein the output of the processing of the first domain-specific licensee 102A is used to form the input of one or more second domain-specific licensees 102B.

[0117] The advantage of the implementation described herein is that the first and second domain-specific licensees 102A and 102B remain substantially unchanged. The main difference lies in minor perturbations to the structure and operation of the top-level agent 100.

[0118] Top-level agent 100, specifically top-level hint generator 104, receives top-level query 108 defining the task chain. Top-level hint builder 104 ultimately produces a series of top-level hints 112, each hint initiating an iteration. Each such top-level hint 112 is provided to top-level model 106. As a result, top-level model 106 produces a series of top-level model outputs 114, one for each iteration. In each iteration, top-level query builder 126 receives one of the top-level model outputs 114. It uses this to form second-level queries 130A, 130B, to be passed to one of the domain-specific authorized parties 102A, 102B.

[0119] In addition to the top-level query 108, the top-level hint builder 146 also receives context information 36, domain information 110A, and interaction context 150 as input.

[0120] When the domain-specific authorized parties 102A and 102B complete their processing, they provide their domain-specific responses 132A and 132B to the context updater 152. After processing the domain-specific responses 132A and 132B into a form suitable for inclusion in the interaction context, the context updater 152 adds them to the interaction context 150. As a result, the interaction context 150 is an accumulation of information about the results of previous iterations performed by the top-level agent 100.

[0121] During the processing of the persistent state of top-level query 108, top-level hint builder 104 provides a series of top-level hints 112 to top-level model 106. Each such top-level hint 112 triggers an iteration. During each such iteration, top-level model 106 provides top-level model output 114 to query builder 126, which then provides domain-specific queries 130A and 130B to domain-specific authorized parties 102A and 102B. This, in turn, produces domain-specific results 132A and 132B.

[0122] The domain-specific results 132A and 132B from each iteration are provided to the context updater 152. The context updater 152 then updates the interaction context 110, which is then used by the top-level hint builder 146 to construct subsequent top-level hints 152.

[0123] The iteration continues until, at some point, the top-level model output 114 indicates that the top-level response 134 is no longer earlier. In this iteration, which is actually the last iteration, the top-level model output 114 is suitable to be used as the top-level response 134. Therefore, the top-level query generator 126 simply passes it to the user 14 as the top-level response 134 to the top-level query 108.

[0124] During the first iteration, the interaction context 150 is essentially empty. As processing continues, the interaction context 150 accumulates information that is then used to drive subsequent iterations. The top-level hint generator 104 uses the accumulated body of the interaction context 150 to generate a new top-level hint 112 that clarifies the nature of the entire task chain defined by the top-level query 108 and collapses any updates related to the tasks that have been executed into this new top-level hint 112. The top-level hint generator 104 then closes the new top-level hint 112 using a request to execute the next task in the task chain.

[0125] Therefore, in the context of this example, once the top-level prompt builder 104 has identified that the first task in the task chain (i.e., determining the price of nore) has been completed, it will provide a new top-level prompt 112 to the top-level model 106, which is in the form of: "When the crew requests to send the price of nore to Septimus Selden and finds that the price is $58 per milligram, create a prompt to perform the next step in the crew request...".

[0126] In this scenario, top-level model 106 creates top-level model output 114, which prompts top-level query builder 126 to provide a domain-specific query 130B, enabling a domain-specific authorized party 102B to compose a message for Septimus Selden, including the fact that noreium is currently sold for twenty cents per gram. Upon completing this task, the domain-specific authorized party 102B provides the relevant information to context updater 152, which then updates the interaction context 150. Based on this updated interaction context 150, top-level prompt builder 104 identifies that the next iteration is the final iteration and outputs a top-level prompt 112 that causes a top-level response 134 to be provided to user 14.

[0127] Figure 9 It shows Figure 5 and Figure 8 The user interface 154 is shown in detail. User interface 154 translates the top-level query 108 into a form suitable for use by the vehicle management system 34, and also translates the top-level response 134 from the vehicle management system 34 into a form suitable for the user 14. User interface 154 has a text-to-speech converter 156 that translates the top-level response 134 into an audio signal from the speaker 18, and a speech-to-text converter 158 that translates the top-level query 108 received at the microphone 20 into text for use by the vehicle management system 34.

[0128] To allow the mode selector 50 to select within the interactive mode 48, it is useful to modify the user interface 154, such as... Figure 10 As shown, it also includes a touch-to-text converter 160 and a text-to-image converter 162. This provides the mode selector 50 with four ways to operate the user interface 154.

[0129] In response to user 14 touching a specific button 84, touch text converter 160 retrieves the text associated with that button from the table and provides that text to vehicle assistant 34. For vehicle assistant 34, this is entirely equivalent to receiving the same text in voice form via microphone 20.

[0130] Therefore, in Figure 3 In this context, if user 14 touches the dynamically generated prompt 88, the touch text converter 160 will provide the in-vehicle concierge 34 with a top-level query 108, such as "Please recommend a French restaurant within a 500 picosecond radius of the intersection of Third and Fourth Streets." Alternatively, if user 14 presses the "OK" button 86, the corresponding top-level query 108 might take the form of: "Please reserve two seats at 'Le Feu' tonight at 7 PM." In either case, the in-vehicle concierge 34 will use the already combined... Figure 5 and Figure 8The user interface 154 operates in a manner described above. In this way, the user interface 154 provides the user 14 with the ability to speak the top-level query 108 or simply use the top-level query 108 that has already been proposed by the vehicle manager 34.

[0131] Text-to-image converter 162 operates in a similar manner. However, text-to-image translator 162 does not simply type text onto display 32 in the manner of an old-fashioned teletypewriter. Text-to-image converter 162 uses the graphical capabilities of display 32 to display information in top-level response 134. For example, in response to a top-level query 108 in the form of “Where are we?”, text-to-image converter 162 will not simply type the words “You are currently on the Westville Highway, heading north towards the cloverleaf interchange of the Wakanda Connector.” Alternatively, text-to-image converter 162 could display map 164, where movement points 166 indicate vehicle locations, such as… Figure 4 As shown. In some embodiments, displayable objects such as map 164 are provided by a domain-specific licensee and passed as objects embedded in top-level response 134. Such objects can be reproduced by text-to-image converter 162, but are ignored by text-to-speech converter 156. It should be noted that such objects will have names expressed in the form of some string of alphanumeric characters, i.e., “text,” and will therefore be appropriate operands that text-to-image converter 162 can operate on.

[0132] As already pointed out, the Car Butler 34 predicts the possible course of the conversation. Figure 3 The effect of this prediction is shown. Based on the text shown in the dynamically generated prompt 88, the in-car concierge has clearly recognized that offering the action of reserving a table at a restaurant creates a significant likelihood that the user will inquire about other restaurants. Since the in-car concierge 34 also has access to preference information 68, it recognizes that specifically inquiring about restaurant suggestions would be more useful than simply providing information about restaurants in general.

[0133] Figure 11 The in-vehicle assistant 34 is shown, in which the prediction model 168 receives updated context from the context updater 152 to serve as the basis for predicting what the user 14 might want to ask next, given one or more previous top-level queries 108. The prediction model 168 provides this information to the query builder 126, which sends the corresponding natural language query back to the user interface 154 as a top-level response 134. The pattern selector 50 then determines whether the top-level response 134 is placed on a dynamically generated prompt 88 (such as...). Figure 3 (As shown), or spoken to user 14 via speaker 18.

[0134] The invention and its preferred embodiments have been described, and the novel and patent-protected aspects we claim are shown in the appended claims.

Claims

1. An apparatus for providing information services to a user in a vehicle, the apparatus comprising an infotainment system installed in the vehicle, a context source communicating with the infotainment system, an in-vehicle assistant hosted by the infotainment system, a prompter, and an interaction mode, wherein the in-vehicle assistant engages in interaction with the user based on context from the context source, the prompter generates model prompts in response to instructions from the in-vehicle assistant, the instructions being based on the context and the prompts being selected to cause the model to generate content guiding the user to make selections, the model being a language model, and the interaction mode being used to deliver the content to the user and provide information from the user to the in-vehicle assistant.

2. The apparatus of claim 1, further comprising a mode selector, the mode selector selecting the inbound channel and the outbound channel in response to a selection signal from the vehicle manager, the selection signal being based on the context.

3. The apparatus of claim 1, further comprising a mode selector that switches between a first mode and a second mode in response to the context, wherein, Each of the modes includes an inbound channel and an outbound channel, wherein in the first mode, both the inbound channel and the outbound channel are audio channels, and wherein in the second mode, at most one of the inbound channel and the outbound channel is an audio channel.

4. The apparatus of claim 1, further comprising a mode selector that switches between a first mode and a second mode in response to the context, wherein, Each of the modes includes an inbound channel and an outbound channel, wherein in the first mode, both the inbound channel and the outbound channel are audio channels, and wherein in the second mode, both the inbound mode and the outbound mode are non-audio channels and audio channels.

5. The apparatus of claim 1, further comprising a mode selector that switches between a first mode and a second mode in response to the context, wherein, Each of the modes includes an inbound channel and an outbound channel, wherein, in the second mode, the inbound channel includes a touchscreen interface having selections recognized by content from the model.

6. The apparatus of claim 1, further comprising a mode selector, the mode selector switching between a first mode and a second mode in response to the context, wherein, Each of the modes includes an inbound channel and an outbound channel, wherein, in the first mode, the outbound channel is selected to provide voice information to the user and also to provide text to the user via a display, the text having been generated by the model.

7. The apparatus of claim 1, further comprising a mode selector that, in response to the context, switches between a first interaction mode and an interaction mode selected from the group consisting of a second interaction mode and a third interaction mode, wherein the first interaction mode, the second interaction mode, and the third interaction mode are selected from the group consisting of an audio mode, a non-audio mode, a mixed mode, a privacy mode, and a mute mode.

8. The apparatus according to claim 1, wherein, The context includes information indicating a message for the user, and wherein the in-vehicle manager causes the prompter to prompt the model to generate a summary of the message for delivery to the user.

9. The apparatus according to claim 1, wherein, The context includes information indicating a first message and a second message for the user, as well as preference information indicating that the first message has a higher priority than the second message, wherein the in-vehicle manager delivers the first message before the second message based on the preference information.

10. The apparatus according to claim 1, wherein, The context includes occupancy information, message information, location information, traffic information, driver information, vehicle information, preference information, and environmental information.

11. The apparatus according to claim 1, wherein, The in-vehicle assistant is configured to change the interaction mode based on the presence of an additional passenger in the vehicle, the additional passenger being a passenger other than the user.

12. The apparatus according to claim 1, wherein, The interaction mode is the current interaction mode selected from a plurality of interaction modes, including a first interaction mode and a second interaction mode, wherein the current interaction mode is the first interaction mode, and wherein the vehicle management system is configured to change the current interaction mode to the second interaction mode in response to a change in the context.

13. The apparatus according to claim 1, wherein, The context includes driver information and vehicle information, wherein, in response to the context, the vehicle manager identifies the location of facilities for modifying the driver's state and the vehicle's state.

14. The apparatus according to claim 1, wherein, The context includes driver information indicating the driver's state, and wherein, based on the driver information, the in-vehicle assistant proposes an interaction to facilitate a change in the driver's state, the interaction including cognitive stimuli.