Predictive multimodal interface for automotive assistant

WO2026080833A3PCT designated stage Publication Date: 2026-06-04CERENCE OPERATING CO

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CERENCE OPERATING CO
Filing Date
2025-10-10
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing automotive assistants often provide responses that are not directly relevant to the occupant's immediate needs, wasting the occupant's time and attention, especially when the occupant is busy driving.

Method used

The automotive assistant employs a bidirectional interaction method that includes generating leading and trailing responses. Leading responses directly answer occupant queries, while trailing responses predict and provide supplemental information based on interaction context, using a query predictor to anticipate the occupant's likely next query.

Benefits of technology

Enhances the relevance of responses by providing timely and contextually appropriate information, reducing distractions for the occupant by anticipating their needs and improving interaction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025050478_04062026_PF_FP_ABST
    Figure US2025050478_04062026_PF_FP_ABST
Patent Text Reader

Abstract

A method executed by an automotive assistant includes, at an occupant-interface, receiving a query, providing it to the automotive assistant, and receiving a leading response back from the assistant. The leading response answers the query. This is followed by receiving a trailing response before any further queries. The trailing response solicits another query, which is then provided to the automotive assistant. The assistant generates the leading response based on interaction context that includes an earlier query and response. Predictor output that is based on updated interaction context predicts a second query, which forms the basis for the trailing response. The leading and trailing responses are then sent to the occupant interface.
Need to check novelty before this filing date? Find Prior Art

Description

-0016-US-PRVPREDICTIVE MULTIMODAL INTERFACE FOR AUTOMOTIVE ASSISTANTCROSS-REFERENCE TO RELATED APPLICATIONS[00011 This application claims priority to U.S. provisional application Serial No.63 / 706,118 filed October 11, 2024, the disclosure of which is hereby incorporated in its entirety by reference herein.TECHNICAL FIELD

[0002] A modern vehicle typically includes an infotainment system on which an automotive assistant carries out various tasks on behalf of an occupant. The automotive assistant receives a query and provides a response to that query.BACKGROUND100031 The response of an automotive assistant is targeted to the query. After all, an occupant who is busy driving the car has little interest in hearing about unrelated subject matter.

[0004] For example, upon receiving a request for navigation instructions, the automotive assistant will launch a GPS application and provide it with the necessary information to begin realtime navigation.SUMMARY

[0005] In one aspect, the disclosure features a method that is by an automotive assistant that engages in a bidirectional interaction with an occupant of a vehicle. Such a method includes steps carried out at the occupant interface and further steps carried out at the automotive assistant.

[0006] The steps at the occupant interface include receiving a first occupant-query from the occupant and providing it to the automotive assistant. This is followed by receiving, from the automotive assistant, a first leading-response that answers the first occupant-query and, prior to receiving any further queries from the occupant, receiving a trailing response from the automotive-0016-US-PRV assistant, and providing it to the occupant. This trailing response solicits a second occupant-query from the occupant. The leading and trailing responses are generated at the automotive assistant by steps set forth below.

[0007] Assuming that the occupant responds to the solicitation, there is a further step of providing the second occupant-query to the automotive assistant and receiving, from the automotive assistant, a second leading-response, which answers the second occupant-query.[0008| The steps undertaken at the automotive assistant include generating the first leading-response based on the interaction context and updating context information. This context information includes interaction context that includes at least one query and response.

[0009] Further steps undertaken at the automotive assistant include generating a predictor output that predicts the second occupant-query based on the updated interaction context, generating the trailing response based on the predictor output, and transmitting both the first leading-response and the trailing response to the occupant-interface.

[0001] These and other features of the invention will be apparent from the following detailed description and the accompanying figures, in which:BRIEF DESCRIPTION OF THE DRAWINGS|0011] FIG. 1 shows a vehicle having an automotive assistant;

[0012] FIG. 2 shows the architecture of the automotive assistant shown in FIG. 1;

[0013] FIG. 3 shows details of the predictor shown in FIG. 2.DETAILED DESCRIPTION

[0014] As required, detailed embodiments of the present invention are disclosed herein; however, it is to be understood that the disclosed embodiments are merely exemplary of the invention that may be embodied in various and alternative forms. The figures are not necessarily-0016-US-PRV to scale; some features may be exaggerated or minimized to show details of particular components. Therefore, specific structural and functional details disclosed herein are not to be interpreted as limiting, but merely as a representative basis forteaching one skilled in the art to variously employ the present invention.

[0015] FIG. 1 shows a vehicle 10 in which an occupant 12 communicates with an automotive assistant 14 via an occupant-interface 16. This may be in the form of acoustic utterances including spoken words, phrases, or commands from vehicle occupants. The occupant-interface 16 receives information from the occupant 12 and transforms it into an occupant query 18 that it sends to the automotive assistant 14. The information from the occupant 12 reaches the occupant- interface 16 in the form of speech via a microphone 20 or in the form of a haptic input, either from a touchscreen display 22 or a physical button in the vehicle 10.

[0016] The automotive assistant 14 processes the occupant query 18 and generates therefrom a query response 24, which it returns to the occupant-interface 16. The information in the query response 24 includes one or more of a speech component, which is delivered to the occupant 12 via a loudspeaker 26, and a graphic component, which is delivered to the occupant 12 via the touchscreen display 22, and a combination thereof.

[0017] In the course of formulating the query response 24, the automotive assistant communicates with a remote server 28, via a network interface 30, an application set 32 that includes external applications 34.1, 34.2, or some combination thereof.

[0018] Communication with the application set 32 includes the automotive assistant 14 generating an application-query 36 that is customized to a particular one of the external applications 34.1 and receiving a corresponding application-response 38 from that external application 34.1.

[0019] An interaction between the occupant 12 and the automotive assistant 14 involves communication towards the occupant 12 and away from the occupant 12. For ease of exposition, those communications that travel towards the occupant 12 shall be referred to herein as “responses” whereas those that travel away from the occupant shall be referred to herein as “queries.”-0016-US-PRV

[0020] As shown in FIG. 1, there are two kinds of response 24: a leading response 24.1 and a trailing response 24.2.

[0021] A leading response 24.1 answers the original occupant query 18. A trailing response 24.2 carries supplemental information that goes beyond what is needed to answer the occupant query 18. To see the distinction, it is useful to consider a use case that illustrates an example of each.100221 Suppose the occupant 12 asks, “Is there a baseball diamond near me?” In such a case, the automotive assistant 14 provides a leading response 24.1, which possibly includes a graphic component and an audio component. The graphic component would be a map showing nearby baseball diamonds with a notable pin or flag on the nearest one. This would be displayed on the touchscreen display 22. The audio component would come from having the occupant-interface 16 causing the loudspeaker 26 to announce, “The nearest baseball diamond is next to “Jack’s Abattoir” on Porterhouse Road.”

[0023] As is apparent, the leading response 24.1 simply answers the original occupant query 18. The leading response 24.1 therefore always occurs in response to an occupant query 18.

[0024] The trailing response 24.2 is quite different.

[0025] A trailing response 24.2 arises as a result of the automotive assistant’s attempt to predict a query that has two properties: it has not been made by the occupant 12 and it is likely to be made by the occupant 12 given what has already occurred. Upon making such a prediction, the automotive assistant 14 invites the occupant 12 to consummate those queries. This invitation is the trailing response 24.2.

[0026] The automotive assistant 14 makes predictions based at least in part on a dialog history that comprises queries 18 already made and query responses 24 to those queries 18. As a result, a trailing response 24.2 occurs at the automotive assistant’s own initiative. It does not have to actually “answer” a question posed by an occupant 12. However, the trailing response 24.2 cannot be made without the recent antecedent of an occupant query 18.-0016-US-PRV

[0027] In response to the foregoing occupant query 18, in which the occupant 12 sought out a nearby baseball diamond, the automotive assistant 14 predicts that the likelihood of a request for navigation instructions receives a boost as a result of the recent request for the location of a diamond. Based on the vehicle’s driving history, the automotive assistant recognizes that the vehicle 10 has not traveled to this part of the city very often. This further boosts the likelihood of a request for navigation instructions. With this being the case, the automotive assistant 14 promptly follows with a trailing response 24.2 of the form, “I see that you have not driven in this neighborhood often. Shall I fetch directions to the baseball diamond?”10028] The vehicle 10 may include one or more processors configured to perform certain instructions, commands and other routines as described herein. Internal vehicle networks 126 may also be included, such as a vehicle controller area network (CAN), an Ethernet network, and a media oriented system transfer (MOST), etc. The internal vehicle networks may allow the processor to communicate with other vehicle systems, such as a vehicle modem, a GPS module and / or Global System for Mobile Communication (GSM) module configured to provide current vehicle location and heading information, and various vehicle electronic control units (ECUs) configured to corporate with the processor.

[0029] The processor may execute instructions for certain vehicle applications, including navigation, infotainment, climate control, etc. Instructions for the respective vehicle systems may be maintained in a non-volatile manner using a variety of types of computer-readable storage medium. The computer-readable storage medium (also referred to herein as memory, or storage) includes any non-transitory medium (e.g., a tangible medium) that participates in providing instructions or other data that may be read by the processor. Computer-executable instructions may be compiled or interpreted from computer programs created using a variety of programming languages and / or technologies, including, without limitation, and either alone or in combination, Java, C, C++, C#, Objective C, Fortran, Pascal, Java Script, Python, Perl, and PL / structured query language (SQL).

[0030] FIG. 2 shows an automotive assistant 14 that generates leading responses 24.1 and trailing responses 24.2 of the type discussed in connection with FIG. 1.-0016-US-PRV

[0031] The illustrated automotive assistant 14 has a hierarchy 40 having an upper level 42 and a lower level 44. The upper level 42 sends upper-level queries 46 to domain-specific delegees 48 in the lower level 44, of which only one is shown. Note that upper-level queries 46 travel away from the user 12, which is why they are “queries.” In response, the domain- specific delegee 48 that was selected to receive the upper-level query 46 provides a domain- specific response 50 back to the upper level 42. A domain-specific response 50 travels towards the user 12, which makes it a “response” instead of a “query.”

[0032] The upper level 42 features an upper-level prompt builder 52 that receives the occupant query 18 from the occupant-interface 16. The upper-level prompt builder 52 also has access to context information, which includes domain information 54, interaction context 56, user preferences, contents derived from internal and external sensors, including vehicle state, location, traffic, weather, occupancy, and time and date.

[0033] The interaction context 56 records interactions between the occupant 20 and the automotive assistant 14. A context updater 58 observes the domain-specific responses 50 along with the upper-level queries 46 from the upper-level query builder 52. As a result, the interaction context 56 includes previous queries 18 and previous query responses 24 that can be drawn upon by both the upper-level prompt builder 52.

[0034] Based on its inputs, the upper-level prompt builder 52 generates an upper-level prompt 62 that has been tailored to cause an upper-level language model 64 to output an upper-level model-output 66. The upper-level model-output 66 serves two purposes: it identifies a domainspecific delegee 48 that is appropriate for processing the occupant query 18 and it provides information to be used by an upper-level query builder 68 in formulating the upper-level query 46.[0035| Processing continues with the upper-level query 46 being provided to whichever domainspecific delegee 48 is appropriate for the domain. The domain-specific delegee 48 makes an application-call 36 to the appropriate external application 34.1 and receives an application response 38 therefrom.-0016-US-PRV

[0036] The domain-specific del egee 48, having received the application response 38, generates its domain-specific response 50. The information in this domain-specific response 50 is distributed to the context updater 58 and to an upper-level response-manager 70. The context updater 58 uses this information to update the interaction context 56.

[0037] The upper-level response-manager 70 packages this information into a leading response 24.1. This leading response 24.1 comprises one or more of a speech component and a graphic component, both of which have been formatted into a form for use by the occupant- interface 16. In some cases, the upper-level response manager 70 simply relays what it has been provided from the lower level 44. This is equivalent to simply bypassing the upper level 42 since it achieves substantially the same result in substantially the same way.

[0038] The information packaged by the upper-level response-manager 70 need not come from only one domain-specific delegee 48. It is quite possible for the upper-level response- manager 70 to collect information from two or more domain-specific delegees 48 and to assemble a response 24, which is either a leading response 24.1 or a trailing response 24.2, that includes information incorporated from several domain-specific delegees 48.|0039| The foregoing process describes how the automotive assistant 14 creates a leading response 24.1. The process for creating a trailing response 24.2 is somewhat different. This process relies on a query predictor 72.

[0040] The query predictor 72, like the upper-level prompt builder 52, receives interaction context 56. Since the interaction context 56 amounts to a history of interactions, it inherently includes occupant queries 18 and the various query responses 24 thereto. The job of the query predictor 72 is to answer the question, “Given the attached interaction context, what supplemental information would the occupant 12 most likely want to see next?”

[0041] As was the case for the information packaged into the leading response 24.1, the supplemental information packaged in the trailing response 24.2 is multimodal in nature.

[0042] Some of it is presentable on the touchscreen display 22 and some of it deliverable via the loudspeaker 26. However, in many cases, the nature of the supplemental information makes it-0016-US-PRV impracticable to deliver via the loudspeaker 26 for much the same reason that restaurant customers who listen to a waiter recite a litany of daily specials will lapse into inattention by about the third one. As such, it is particularly useful for the query predictor 72 to orchestrate the presentation of such supplemental information into a form that can be visually displayed on the touchscreen display 22. This has dual advantages: first, the occupant 12 will have random access rather than sequential access to the information and second, the occupant 12 will be able to make a selection by touching either a configurable button or an active region of the touchscreen display 22 itself.

[0043] As shown in FIG. 3, the query predictor 72 includes a predictor language-model 74 that is much like the upper-level language-model 64. This is most useful for supplemental information that is primarily textual.

[0044] The query predictor 72 also includes a graphical model 76 that produces supplemental information for delivery via a transient tile 78 on the touchscreen display 22. In one example, the query predictor 72 causes the transient tile 78 to appear on the touchscreen display 22 as a set of options together with an invitation to touch a particular actuator to select a corresponding one of the options.

[0045] In some examples, the touchscreen display 22 presents the set of options is a onedimensional list, with perhaps bullets or numbers providing visual cues to separate elements of the list. For example, as a follow up to the leading response 24.1 that announced the existence of the baseball diamond on Porterhouse Road, the trailing response 24.2 manifests as a transient tile 78 on the touchscreen display 22 showing a list that includes the baseball diamond on Porterhouse Road together with other baseball diamonds that are within a predetermined distance from the vehicle, together with an invitation to select one of the elements on that list to receive directions to that particular baseball diamond.

[0046] Another example takes advantage of the inherently two-dimensional nature of the touchscreen display 22. In such an example, the options in the set of options are overlaid on a map of the surrounding area, thus providing the occupant 12 with more topological information than could be had using the loudspeaker 26. In this example, the occupant 12 is invited to touch a location of a particular option on the map to obtain supplemental information, such as directions,-0016-US-PRV to reach the baseball diamond corresponding to that option. In such cases, touching the location generates a second occupant-query 18 using pre-defined text. This second occupant-query 18 is then provided to the automotive assistant 14 for processing in much the same way as the original occupant-query 18. As a result, the pre- defined text is processed as if it were spoken by the occupant 12.

[0047] Embodiments include those in which the touchscreen display 22 renders one or more of individual buttons to be pressed to select an option from the list or on the map and also renders text or graphics that are configured to communicate, to the occupant 12, what will occur upon pressing a particular one of those buttons. Among the possibilities of what will occur upon pressing a particular button is that of providing pre-defined text for use as a second occupant query 18 for the automotive assistant 14 to process.

[0048] FIG. 3 shows a suitable implementation of a query predictor 72 that provides textual predictions and graphical predictions.

[0049] For textual prediction, the predictor language-model 74 receives a first prediction- prompt 76 from a predictor-prompt builder 80. In response to the first prediction-prompt 76, the predictor model 74 provides a prediction-model output 82 that serves as the basis for the textual or audio component of the trailing response 24.2.

[0050] For graphics prediction, the graphic model 76 receives a second prediction- prompt 84 from the predictor-prompt builder 80. In response to the second prediction-prompt 82, the graphic model 76 provides a graphic-model output 86 that serves as the basis for the graphic component of the trailing response 24.2.|0051[ A response packager 88 uses the prediction-model output 82 and the graphic- model output 86 to provide an output to the upper-level response manager 70 for use in building the trailing response 24.2. The upper-level response manager 70 then does so in much the same way that it builds the leading response 24.1.

[0052] While exemplary embodiments are described above, it is not intended that these embodiments describe all possible forms of the invention. Rather, the words used in the-0016-US-PRV specification are words of description rather than limitation, and it is understood that various changes may be made without departing from the spirit and scope of the invention. Additionally, the features of various implementing embodiments may be combined to form further embodiments of the invention.[00531 Having described the invention and a preferred embodiment thereof, what is claimed as new and secured by letters patent is:

Claims

-0016-US-PRVWHAT IS CLAIMED IS:

1. A method executed by an automotive assistant that engages in a bidirectional interaction with an occupant of a vehicle, the method comprising: at an occupant-interface: receiving a first occupant-query from an occupant, providing a first occupant-query to an automotive assistant, receiving, from the automotive assistant, a first leading-response, wherein a first leading-response answers the first occupant-query, prior to receiving further queries from the occupant, receiving a trailing response from the automotive assistant and providing the trailing response to the occupant, wherein the trailing response solicits a second occupant-query from the occupant, providing a second occupant-query to the automotive assistant, and receiving, from the automotive assistant, a second leading-response, wherein the second leading-response answers the second occupant-query, at the automotive assistant, generating the first leading-response based on the interaction context, updating context information, the context information including interaction context that includes at least one query and response, generating a predictor output that predicts the second occupant-query based on the updated interaction context, generating the trailing response based on the predictor output, and transmitting both the first leading-response and the trailing response to the occupant-interface.

2. The method of claim 1, wherein the interaction context includes a history of interactions between the occupant and the vehicle.

3. The method of claim 2, wherein the history of interactions includes previous queries and previous query responses.-0016-US-PRV4. The method of claim 1, wherein the trailing response is provided to the occupant via visual cues on a display screen.

5. The method of claim 1, wherein the trailing response is provided to the occupant via a speaker.

6. The method of claim 1, wherein the predictor output includes a language model.

7. The method of claim 1, wherein the first query response includes an audible output.

8. The method of claim 1, wherein the first query response includes an audible output and the trailing response includes a visual output.

9. A method executed by an automotive assistant that engages in a bidirectional interaction with an occupant of a vehicle, the method comprising: receiving a first occupant-query from a vehicle occupant, receiving an interaction context, generating a first leading-response based on the first occupant-query, generating a predictor output that predicts a second occupant-query based on the interaction context, generating a trailing response and providing the trailing response to the occupant, wherein the trailing response is based on the predictor output and solicits a second occupant-query from the occupant.

10. The method of claim 9, wherein the interaction context includes a history of interactions between the occupant and the vehicle.

11. The method of claim 10, wherein the history of interactions includes previous queries and previous query responses.-0016-US-PRV12. The method of claim 9, wherein the trailing response is provided to the occupant via visual cues on a display screen.

13. The method of claim 9, wherein the trailing response is provided to the occupant via a speaker.

14. The method of claim 9, wherein the predictor output includes a language model.

15. The method of claim 9, wherein the first query response includes an audible output.

16. The method of claim 9, wherein the first query response includes an audible output and the trailing response includes a visual output.

17. A method executed by an automotive assistant that engages in a bidirectional interaction with an occupant of a vehicle, the method comprising: receiving a first occupant-query from a vehicle occupant, receiving an interaction context, generating a first leading-response based on the first occupant-query, generating a predictor output that predicts a second occupant-query based on the interaction context, generating a trailing response and providing the trailing response to the occupant, wherein the trailing response is based on the predictor output and solicits a second occupant-query from the occupant, receiving the second occupant-query from a vehicle occupant, and generating a second leading-response, wherein the second leading-response answers the second occupant-query.-0016-US-PRV18. The method of claim 9, wherein the interaction context includes a history of interactions between the occupant and the vehicle include previous queries and previous query responses.

19. The method of claim 17, wherein the trailing response is provided to the occupant via at least one of visual cues on a display screen and audible cues at a speaker.

20. The method of claim 17. wherein the first query response includes an audible output and the trailing response includes a visual output.