Determination of Location and Route through Natural Conversation

By enabling users to interactively refine route queries through natural conversation, the navigation application addresses the inefficiencies of conventional systems, reducing cognitive overload and computational demands while enhancing user safety.

JP2025516248AInactive Publication Date: 2025-05-27GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024563932
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-05-02
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Conventional navigation applications struggle to provide users with tailored and efficient route proposals, especially when the range of available routes is vast or the user is flexible about the destination, leading to cognitive overload and inefficient use of computational resources.

Method used

The technology enables users to select a destination and route through an interactive conversation with a navigation application, using spoken or input natural language to refine queries and narrow down route options, thereby reducing the need for users to browse through extensive lists of routes.

Benefits of technology

This approach reduces user time and cognitive overload by providing a more natural and efficient way to select routes, while also reducing computational requirements and enhancing safety by allowing users to navigate without visually selecting routes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025516248000001_ABST
    Figure 2025516248000001_ABST
Patent Text Reader

Abstract

A computing device may implement a method for determining a location and a route through natural conversation. The method may include receiving, from a user, voice input including a search query to initiate a navigation session, and generating a set of navigation search results corresponding to the search query. The set of navigation search results includes a plurality of destinations or a plurality of routes corresponding to one or more destinations. The method further includes providing an audio request to the user to narrow down the set of navigation search results, and receiving, in response to the audio request, subsequent voice input from the user including a refined search query. The method further includes providing one or more refined navigation search results corresponding to the refined search query, including a subset of the plurality of destinations or the plurality of routes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to route determination, and more particularly to determining locations and routes through natural conversations.

Background Art

[0002] The description of the background art provided herein is for the purpose of generally presenting the context of the present disclosure. The work of the inventors named herein is not, in the context of this background art section, admitted as prior art to the present disclosure, either explicitly or implicitly, along with aspects of this description that may not be eligible as prior art for some other reason at the time of filing.

[0003] Generally speaking, conventional navigation applications that provide a route to / from a destination are widely prevalent in modern culture. These conventional navigation applications can provide step-by-step and turn-by-turn navigation to reach a pre-programmed destination by driving and / or some other means of transportation (e.g., walking, public transportation, etc.). In particular, in conventional navigation applications, the user specifies a starting point and a destination, and then the user is presented with a set of route proposals based on various means of transportation. Thus, these route proposals are typically provided to the user as a result of a single interaction between the user and the application, where the user enters a query and is later presented with a list of route proposals.

[0004] However, in situations where the range of available routes is wide or the user is flexible regarding the accuracy of the destination, this conventional single interaction method may be insufficient. For example, a user may wish to navigate a hiking course without assuming a specific hiking course, and as a result, there may be a very large number of possible route configurations to reach multiple different hiking courses. If each of these route options were displayed, they could overwhelm the user or simply take too much time to browse, so the user may select an undesirable route or not select a route at all. Furthermore, to display each of the numerous routes, a large amount of computational resources are correspondingly required to determine and provide each of those routes on the client device.

[0005] Therefore, generally, conventional navigation applications are unable to provide users with reachable, particularly user-tailored route proposals, and there is a need for a navigation application that can determine locations and routes through natural conversation to avoid these problems associated with conventional navigation. Summary of the Invention

[0006] By using the technology of the present disclosure, a user's computing device may be enabled to select a destination and a route through an interactive conversation with a navigation application. Specifically, a user can initiate a route query (e.g., a navigation session) through a spoken or input natural language request, which can lead to follow-up questions from the navigation application and / or further refinements from the user. A navigation application utilizing the present invention can then narrow down the route / destination provided to the user based on the user's responses to the follow-up questions. In this way, the present invention enables the user to narrow down the selection of the user's destination or route in a more natural way than the prior art through a two-way conversation with their device. As a result, the present invention can reduce the user's time and cognitive overload by eliminating the need to browse through a long list of different route proposals and manually compare them. In this way, the present invention solves the technical problem of efficiently determining a route to a destination. This is further enabled by the fact that the routes provided to the user for selection are a narrowed-down list of all possible routes. That is, since the number of routes provided is reduced, it means that the computational requirements necessary to provide routes to the user are reduced compared to the prior art. In this way, the present invention provides a more computationally efficient means for determining a route to a destination. An additional technical advantage provided by the present invention is the advantage of a safer means for providing routes to the user for selection. The disclosed technology that enables a user to narrow down and select a set of routes to a destination by voice input is less distracting for the user compared to the prior art of looking at the routes displayed on a screen and selecting one of those routes by touch input. With the disclosed technology, a vehicle driver can select a route without taking their eyes off the road or taking their hands off the control of the vehicle.Furthermore, a user who is a vehicle driver can safely narrow down or update a route while already moving along that route using voice input and a conversational interface. In this way, the disclosed technology provides a safer means for selecting and narrowing down a route to a destination. Additionally, since the present invention prompts the user to explicitly state their preferences as part of the conversation flow, it can also provide route suggestions that better meet the user's needs and preferences than the prior art. However, embodiments of the present invention are not particularly limited to achieving effects based on the user's preferences. Some of the disclosures of the present invention are independent of the user's preferences.

[0007] The present invention can function in either a speech-based interface or a touch-based interface setting. However, for ease of explanation, the conversation flow between the user and the navigation application (and corresponding processing components) described herein can generally be in the context of a speech-based configuration. Nevertheless, with respect to a touch-based interface, questions for clarification can be presented to the user in the user interface, and the user can answer those questions for clarification via free-form text input or via UI elements (such as a drop-down menu). The embodiments disclosed herein, which are described in the context of a speech-based interface, can also be applied in the context of a touch-based interface. All embodiments disclosed herein where the input or output is described in the context of a voice-based interface can be adapted to apply in the context of a touch-based interface.

[0008] In the first example, the techniques of the present disclosure can resolve locations when the user is flexible regarding the destination. The user may be traveling in Switzerland and can start a navigation session by saying "Navigate to a nearby hiking spot." If there are multiple hiking spots that meet the "nearby" constraint, the navigation system may respond to the user with an audio request such as "What is the maximum acceptable time?" to narrow down the set of routes. The user may respond to the audio request by stating "Within 30 minutes by car." However, when the number of available options is still relatively large, the navigation application may generate subsequent voice requests such as "Some of the top-rated options require taking a cable car from the parking lot. Are you willing to do so? The total travel time is expected to be less than 30 minutes." The user can respond with "Yes, that's fine," and the navigation application can respond with several options for hikes within 30 minutes of travel time that are highly rated but include both driving and taking a cable car. Thereafter, the user can further narrow down the route options returned in the follow-up statement or accept one of the suggestions provided.

[0009] In a second example, the techniques of the present disclosure can be configured to generate filtered natural language route suggestions. For example, a user may arrive at an airport in Italy and want to navigate to a hotel. The user can ask a navigation application, "Show me the route to Tenuta il Cigno." The navigation application can respond with several different route proposals along with the top candidate by stating, "The route I recommend is the shortest route, but it involves driving on a single-lane road for 10 miles." Instead of accepting the proposal, the user may say, "The short journey is certainly appreciated, but is it possible to reduce the time spent on the single-lane road?" The navigation application may then propose an alternative route that is longer but only requires driving on a single-lane road for 2 miles to reach the user's destination. The user may accept this alternative route, view the route, or start the navigation session.

[0010] In a third example, the techniques of the present disclosure can provide clarification of the conversation during a navigation session. Similar to the above example, a user may be navigating from an airport to a hotel at a vacation destination. While the user is in transit, the user may encounter the possibility of a detour that takes a similar amount of time but has different characteristics. As the user approaches this detour, the navigation system can prompt the user by stating, "There is an alternative route on the left with a similar ETA. The distance is shorter, but there is temporary roadwork and there may be a slight delay." In response, the user may say, "Okay, got it, let's take that route," or "No, I don't want to deviate from the current route," and the navigation application can continue with the original route or switch to the alternative route as needed.

[0011] In this way, aspects of the present disclosure provide a technical solution to the problem of non-optimal route suggestions by automatically filtering route options based on a conversation between a user and a navigation application. Aspects of the present disclosure also provide a technical effect for the problem of narrowing down safer routes based on a conversational interaction between a user and a navigation application. In particular, in a conversational interaction, less cognitive input from the user is required. Thus, since the user does not need to physically view and physically select a route on the device's display, the user's attention is diverted less. Instead, the user can verbally narrow down and select a route while driving a vehicle or operating in another way. As described above, conventional systems automatically provide a list of route options in response to a single query raised by the user. As a result, in conventional systems, the search / judgment criteria applied to generate a list of route options by a single query of the user are severely restricted, and thus the ability to narrow down the list of routes presented to the user is severely restricted. Consequently, conventional systems may irritate the user by providing an overwhelming amount of possible routes, many of which are not optimized for the user's particular situation. In contrast, the technology of the present disclosure eliminates such frustrating interactions with the navigation application by conversing with the user until the application has enough information to determine a narrowed set of optimal route suggestions, each of which is tailored to the user's particular situation. Also, the narrowed set of optimal route suggestions requires fewer computational resources to process and provide to the user for selection, since the set of narrowed routes contains fewer routes than the original set contains. Thus, a more computationally efficient technology is disclosed compared to the prior art.

[0012] An exemplary embodiment of the technology of the present disclosure is a method in a computing device for determining a location and a route through natural conversation. The method includes receiving, from a user, an audio input including a search query to start a navigation session; generating, by one or more processors, a set of navigation search results responsive to the search query, the set of navigation search results including a plurality of destinations or a plurality of routes corresponding to one or more destinations; providing, by one or more processors, an audio request to the user to narrow down the set of navigation search results; receiving, in response to the audio request, a subsequent audio input from the user including a narrowed search query; and providing, by one or more processors, one or more narrowed navigation search results responsive to the narrowed search query, the one or more narrowed navigation search results including a subset of the plurality of destinations or the plurality of routes.

[0013] Another exemplary embodiment is a computing device for determining a location and a route through natural conversation. The computing device includes a user interface, one or more processors, and a computer-readable memory, optionally non-transitory, coupled to the one or more processors and storing instructions that, when executed by the one or more processors, cause the computing device to receive, from a user, voice input including a search query to initiate a navigation session; generate a set of navigation search results responsive to the search query, the set of navigation search results including a plurality of destinations or a plurality of routes corresponding to one or more destinations; provide an audio request to the user to narrow down the set of navigation search results; receive, in response to the audio request, subsequent voice input from the user including a narrowed search query; and provide one or more narrowed navigation search results responsive to the narrowed search query, the one or more narrowed navigation search results including a subset of the plurality of destinations or the plurality of routes.

[0014] Yet another exemplary embodiment is a computer-readable medium, optionally non-transitory, storing instructions for determining a location and a route through natural conversation, the instructions, when executed by one or more processors, causing the one or more processors to receive, from a user, an audio input including a search query to initiate a navigation session; generate a set of navigation search results responsive to the search query, the set of navigation search results including a plurality of destinations or a plurality of routes corresponding to one or more destinations; provide an audio request to the user to narrow down the set of navigation search results; receive, from the user, a subsequent audio input including a narrowed search query responsive to the audio request; and provide one or more narrowed navigation search results responsive to the narrowed search query, the one or more narrowed navigation search results including a subset of the plurality of destinations or a subset of the plurality of routes.

[0015] Another exemplary embodiment is a method in a computing device for determining a location and a route through natural conversation. The method includes receiving an input from a user to initiate a navigation session, generating, responsive to the user input, one or more destinations or one or more routes, and providing a request to the user to narrow down the response to the user input. Responsive to the request, the method includes receiving a subsequent input from the user and providing, responsive to the subsequent user input, one or more updated destinations or one or more updated routes. BRIEF DESCRIPTION OF THE DRAWINGS

[0016]

Figure 1A

Figure 1B

Figure 2A

Figure 2B

Figure 2C

Figure 3A

Figure 3B

Figure 4

DETAILED DESCRIPTION OF THE INVENTION

[0017] SUMMARY As described above, navigation applications typically receive user input and automatically generate a number of route options that the user can select. However, in such situations, it may sometimes be better to follow up with the user with a clear question or statement, thereby enabling the user to narrow down the range of possible routes so that the user can reduce the set of route selections. The techniques of the present disclosure achieve this clarification by supporting a conversational route configuration that (i) detects situations where a follow-up question (referred to herein as an "audio request") is beneficial and (ii) provides the user with an opportunity to clarify their preferences in order to identify an optimal route. It will be appreciated that the techniques of the present disclosure may also achieve clarification and suggestions of optimal routes in a manner that does not depend on the user's preferences. For example, objectively safer, faster, or shorter route suggestions may be provided based on the conversational route configuration.

[0018] Generally speaking, a user's computing device can generate a set of filtered navigation search results based on a series of inputs received from the user as part of a conversational dialogue with the user computing device. More specifically, the user computing device can receive voice input from the user that includes a search query to initiate a navigation session. A navigation session generally corresponds to a set of navigation instructions intended to guide the user from a current location or a specified location to a destination, and such navigation instructions may be rendered on a user interface for display to the user or communicated audibly through an audio output component of the user computing device. Then, the user computing device can generate a set of navigation search results corresponding to the search query, and the set of navigation search results may include multiple destinations or multiple routes corresponding to one or more destinations.

[0019] At this point, the user computing device may determine that it can / should narrow down the set of navigation search results before providing the search results to the user (e.g., via a navigation application). For example, the user computing device may determine that the number of route options included in the set of navigation search results is too large (e.g., exceeds a route presentation threshold), is likely to confuse the user, and / or is likely to overwhelm the user as is, or that it would be too computationally costly to provide the set of search results to the user. Additionally, or alternatively, the user computing device may determine that the optimal route included in the set of navigation instructions is characterized by potentially dangerous and / or otherwise unusual driving conditions and that it is necessary to inform the user of these things before or during the navigation session. In any case, if the user computing device determines that it is necessary to prompt the user for an audio request, the user computing device can provide the user with an audio request to narrow down the set of navigation search results.

[0020] In response thereto, and in response to an audio request, the user computing device may receive subsequent voice input from the user that includes a refined search query. This refined search query may include keywords or other phrases that directly correspond to keywords or phrases included as part of the audio request, enabling the user computing device to narrow a set of navigation search results based on the user's subsequent voice input. For example, the audio request provided to the user by the user computing device may prompt the user to specify a maximum desired travel time to a destination. In response, the user may state, "I don't want to be on the road for more than 30 minutes." The user computing device may receive this subsequent voice input from the user, interpret that 30 minutes is the maximum desired travel time, and filter the set of navigation search results by excluding routes with an estimated travel time exceeding 30 minutes. Thereafter, the user computing device may provide one or more refined navigation search results corresponding to the refined search query, including a subset of multiple destinations or multiple routes.

[0021] In this way, aspects of the present disclosure provide a technical solution to the problem of suboptimal route suggestions by automatically filtering route options based on conversations between a user and a navigation application. Conventional systems automatically provide a list of route options in response to a single query raised by the user, and as a result, are severely limited in the search / judgment criteria applied to generate the list of route options and the ability to narrow down the list of routes provided to the user. Such conventional systems generally frustrate the user by providing an overwhelming amount of possible routes, many of which are not optimized for the user's specific situation. In contrast, the techniques of the present disclosure eliminate such frustrating interactions with the navigation application by conversing with the user until the application has enough information to determine a narrowed set of navigation search results, each tailored to the user's specific situation. The techniques of the present disclosure provide a technical solution to the problem of optimizing computational resources when providing route suggestions by narrowing down possible routes through conversations with the user.

[0022] Furthermore, the present technology improves the overall user experience when using a navigation application, and more broadly, when receiving navigation instructions to a desired destination. In some examples, the present technology automatically determines a set of filtered navigation search results that are tailored / curated to the user's preferences, such that they are determined through an intuitive and non-distracting conversation between the user and their computing device. This improves the user's satisfaction with their travel plans, reduces the user's distraction while traveling to the desired destination, and reduces user confusion and frustration resulting from sub-optimal and / or otherwise irrelevant / inappropriate navigation recommendations from conventional navigation applications, thereby providing a safer, more user-friendly and relevant experience. Accordingly, the present technology enables a safer, more user-specific, and more enjoyable navigation session to the desired destination.

[0023] Exemplary Hardware and Software Components Referring initially to FIG. 1A, an exemplary communication system 100 that can implement techniques for determining locations and routes through natural conversation includes a user computing device 102. The user computing device 102 may be, for example, a portable device such as a smartphone or a tablet computer. The user computing device 102 may also be a wearable device such as a laptop computer, a desktop computer, a personal digital assistant (PDA), a smartwatch, or smart glasses. In some embodiments, the user computing device 102 may be removably attached to a vehicle, embedded in a vehicle, and / or capable of interacting with a vehicle's head unit to provide navigation instructions.

[0024] The user computing device 102 may include one or more processors 104 and a memory 106 that stores machine-readable instructions executable on the processor(s) 104. The processor(s) 104 may include one or more general-purpose processors (e.g., a CPU) and / or a dedicated processing unit (e.g., a graphics processing unit (GPU)). The memory 106 may optionally be a non-transitory memory and may include one or several suitable memory modules such as random access memory (RAM), read-only memory (ROM), flash memory, and other types of persistent memory. The memory 106 can store instructions for implementing a navigation application 108 that can provide navigation guidance (e.g., by displaying guidance via the user computing device 102 or issuing audio instructions), display an interactive digital map, request and receive routing data to provide driving, walking, or other navigation guidance, and provide various location-based content such as traffic, points of interest (POI), and weather information.

[0025] Furthermore, memory 102 may include a language processing module 109a configured to implement and / or support the techniques of the present disclosure for determining locations and routes through natural conversations. That is, language processing module 109a may include an automatic speech recognition (ASR) engine 109a1 configured to transcribe voice input from the user into a set of text. Further, language processing module 109a may include a text-to-speech (TTS) engine 109a2 configured to convert the text into an audio output such as an audio request, a navigation command, and / or other output for the user. In some scenarios, language processing module 109a may include a natural language processing (NLP) model 109a3 configured to output a text transcription, intent interpretation, and / or audio output related to voice input received from a user of user computing device 102. As described herein, ASR engine 109a1 and / or TTS engine 109a2 may be included as part of NLP model 109a3 to transcribe user voice input into a text set, convert text output into audio output, and / or perform any other suitable functions described herein as part of a conversation between user computing device 102 and the user.

[0026] Generally, the language processing module 109a may include computer-executable instructions for training and operating the NLP model 109a3. Generally, the language processing module 109a can train one or more NLP models 109a3 by establishing a network architecture, or topology, and adding layers associated with one or more activation functions (e.g., rectified linear units, softmax, etc.), loss functions, and / or optimization functions. Such training can generally be performed using symbolic methods, machine learning (ML) models, and / or any other suitable training methods. More generally, the language processing module 109a can train the NLP model 109a3 to perform two techniques, namely syntactic analysis and semantic analysis, such that the user computing device 102, and / or any other suitable device (e.g., the vehicle computing device 151), can understand words spoken by the user and / or words generated by a text-to-speech synthesis program (e.g., the TTS engine 109a2) executed by the processor 104.

[0027] Syntactic analysis generally involves using basic grammar rules to analyze text and identify the overall sentence structure, how specific words within the sentence are structured, and how the words within the sentence relate to each other. Syntactic analysis can include one or more subtasks such as tokenization, part-of-speech (PoS) tagging, parsing, lemmatization and stemming, stop word removal, and / or any other suitable subtasks or combinations thereof. For example, the NLP model 109a3 can use syntactic analysis to generate a text transcription from an audio input from the user. Additionally or alternatively, the NLP model 109a3 can receive a text transcription as a set of text from the ASR engine 109a1 for performing semantic analysis on the set of text.

[0028] Semantic analysis generally involves analyzing text in order to understand its meaning and / or capture it in other ways. Specifically, the NLP model 109a3 to which semantic analysis is applied can study the meaning of each individual word included in the text transcription in a process known as lexical semantics. Using these individual meanings, the NLP model 109a3 can then examine the various combinations of words included in the sentences of the text transcription to determine one or more contextual meanings of the words. Semantic analysis can include one or more subtasks, such as semantic ambiguity resolution, relation extraction, sentiment analysis, and / or other appropriate subtasks or combinations thereof. For example, using semantic analysis, the NLP model 109a3 can generate one or more interpretations of intent based on the text transcription from syntactic analysis.

[0029] In these aspects, the language processing module 109a can include an artificial intelligence (AI) trained conversation algorithm (e.g., a natural language processing (NLP) model 109a3) configured to interact with a user accessing the navigation app 108. The user may be directly connected to the navigation app 108 to provide oral input / responses (e.g., voice input), and / or the user request may include text input / responses that the TTS engine 109a2 (and / or other appropriate engine / model / algorithm) converts to audio input / responses for the NLP model 109a3 to interpret. When the user accesses the navigation app 108, the input / responses spoken by the user and / or generated by the TTS engine 109a2 (or other appropriate algorithm) can be analyzed by the NLP model 109a3 to generate a text transcription and an interpretation of intent.

[0030] The language processing module 109a can train one or more NLP models 109a3 to apply these NLP techniques and / or other NLP techniques using multiple training voice inputs from multiple users. As a result, the NLP model 109a3 can be configured to output text transcription and corresponding intent interpretation based on syntactic and semantic analysis of the user's voice input.

[0031] In certain aspects, one or more types of machine learning (ML) can be used by the language processing module 109a to train the NLP model(s) 109a3. The ML may be used by an ML module 109b that can store an ML model 109b1. The ML model 109b1 may be configured to receive a set of texts corresponding to user input and output an intent and a destination based on the set of texts. The NLP model(s) 109a3 may be, and / or may include, one or more types of ML models such as the ML model 109b1. More specifically, in these aspects, the NLP model 109a3 is a machine learning model (e.g., a large language model (LLM)) trained by the ML module 109b using one or more training datasets of texts to output one or more training intents and one or more training destinations, as further described herein, or can include it. For example, an artificial neural network, a recurrent neural network, a deep learning neural network, a Bayesian model, and / or any other suitable ML model 109b1 can be used to train and / or otherwise implement the NLP model(s) 109a3. In these aspects, the training may be performed by repeatedly training the NLP model(s) 109a3 using labeled training samples (e.g., training user input).

[0032] When the NLP model(s) 109a3 is an artificial neural network, the training of the NLP model(s) 109a3 can generate weights as by-products or parameters that can be initialized with random values. The weights can be changed using any of several gradient descent algorithms as the network is repeatedly trained to reduce the loss and converge the values output by the network to the expected values, or "learned" values. In an embodiment, a regression neural network lacking an activation function can be selected, and the input data can be normalized by mean centering to determine the loss and quantify the accuracy of the output. Such normalization can use a mean squared error loss function and mean absolute error. The artificial neural network model can be verified and cross-validated using standard techniques such as holdout, K-fold, etc. In an embodiment, multiple artificial neural networks can be trained and operated separately and / or trained separately and operated in conjunction.

[0033] In an embodiment, one or more NLP models 109a3 may include an artificial neural network having an input layer, one or more hidden layers, and an output layer. Each of the layers of the artificial neural network may include any number of neurons. The multiple layers may be linearly connected to each other such that neurons can pass outputs from one neuron to the next, or the neurons may be networked such that they transmit inputs and outputs non-linearly. In general, it should be understood that many configurations and / or connections of an artificial neural network are possible. For example, the input layer may correspond to input parameters given as a complete sentence, or input parameters separated according to word or character (e.g., fixed-width) limitations. The input layer may, in some embodiments, correspond to a large number of input parameters (e.g., one million inputs) and may be analyzed sequentially or in parallel. Further, the various neurons and / or neuron connections within the artificial neural network may be initialized with any number of weights and / or other training parameters. Each of the neurons in the hidden layer can analyze one or more of the input parameters from the input layer and / or one or more of the outputs from one or more of the previous hidden layers to generate a determination or other output. The output layer may include one or more outputs, each of which indicates a prediction. In some embodiments and / or scenarios, the output layer includes only a single output.

[0034] FIG. 1A shows the navigation application 108 as a stand-alone application, but note that the functionality of the navigation application 108 can also be provided in the form of an online service accessible via a web browser running on the user computing device 102, a plug-in or extension of another software application running on the user computing device 102, etc. The navigation application 108 can generally be provided in different versions for different operating systems. For example, a manufacturer of the user computing device 102 can provide a software development kit (SDK) that includes the navigation application 108 for the Android (trademark) platform, another SDK for the iOS (trademark) platform, and so on.

[0035] Memory 106 can also store an operating system (OS) 110, which can be any type of suitable mobile or general-purpose operating system. User computing device 102 can further include a global positioning system (GPS) 112 or other suitable positioning module, a network module 114, a user interface 116 for displaying map data and directions, and an input / output (I / O) module 118. Network module 114 can include one or more communication interfaces such as hardware, software, and / or firmware of an interface to enable communication via a cellular network, a Wi-Fi network, or other suitable network such as network 144 described later. I / O module 118 can include I / O devices that can receive input from the surrounding environment and / or the user and provide output to the surrounding environment and / or the user. I / O module 118 can include a touch screen, a display, a keyboard, a mouse, buttons, keys, a microphone, a speaker, etc. In various embodiments, user computing device 102 can include fewer components than illustrated in FIG. 1A, or conversely, can include additional components.

[0036] The user computing device 102 may communicate with the external server 120 and / or the vehicle computing device 150 via the network 144. The network 144 may include one or more of an Ethernet-based network, a private network, a cellular network, a local area network (LAN), and / or a wide area network (WAN) such as the Internet. The navigation application 108 may send map data, navigation directions, and other geographic content from the map database 156 to the vehicle computing device 150 for display on the cluster display unit 151. Additionally or alternatively, the navigation application 108 may access maps, navigation, and geographic location content stored locally on the user computing device 102 and may periodically access the map database 156 during navigation to update local data or to access real-time information such as real-time traffic data. Further, the user computing device 102 may be directly connected to the vehicle computing device 150 via any suitable direct communication link 140 such as a wired connection (e.g., a USB connection).

[0037] In certain embodiments, the network 144 may include any communication link suitable for short-range communication and may comply with communication protocols such as Bluetooth (trademark) (e.g., BLE), Wi-Fi (e.g., Wi-Fi Direct), NFC, ultrasonic signals, etc. Additionally or alternatively, the network 144 may be, for example, Wi-Fi, a cellular communication link (e.g., compliant with 3G, 4G, or 5G standards). In some scenarios, the network 144 may also include a wired connection.

[0038] External server 120 may be a remotely located server that includes the processing power and executable instructions necessary to perform some or all of the actions described herein with respect to user computing device 102. For example, external server 120 may include a language processing module 120a similar to language processing module 109a included as part of user computing device 102, and module 120a may include one or more of ASR engine 109a1, TTS engine 109a2, and / or NLP model 109a3. External server 120 may also include a navigation application 120b and an ML module 120c similar to navigation application 108 and ML module 109b included as part of user computing device 102.

[0039] Vehicle computing device 150 includes one or more processors 152 and a memory 153 that stores computer-readable instructions executable by processors 152. Memory 153 can store a language processing module 153a, a navigation application 153b, and an ML module 153c that are respectively similar to language processing module 153a, navigation application 108, and ML module 109b. Navigation application 153b can support functions similar to navigation application 108 from the vehicle side and can facilitate the rendering of information displays as described herein. For example, in a particular aspect, user computing device 102 may provide vehicle computing device 150 with an accepted route received by the user and corresponding navigation instructions that are to be provided to the user as part of the accepted route. Navigation application 153b can then proceed to render the navigation instructions within cluster unit display 151 and / or generate an audio output that verbally provides the navigation instructions to the user via language processing module 153a.

[0040] In any case, the user computing device 102 may be communicatively coupled to various databases such as a map database 156, a traffic database 157, and a point of interest (POI) database 159. From these databases, the user computing device 102 can obtain navigation-related data. The map database 156 may include map data such as map tiles, visual maps, road shape data, road type data, speed limit data, etc. The traffic database 157 can store not only real-time traffic information but also past traffic information. The POI database 159 can store descriptions, locations, images, and other information about landmarks or points of interest. FIG. 1A shows the databases 156, 157, and 159, but the user computing device 102, the vehicle computing device 150, and / or the external server 120 may be communicatively coupled to additional or, conversely, fewer databases. For example, the user computing device 102 and / or the vehicle computing device 150 may be communicatively coupled to a database that stores weather data.

[0041] Referring to FIG. 1B, the user computing device 102 may transmit information for rendering / displaying navigation instructions within the vehicle environment 170. The user computing device 102 may be disposed within the vehicle 172 or may be a smartphone. However, although FIG. 1B shows the user computing device 102 as a smartphone, this is for ease of explanation, and the user computing device 102 may be any suitable type of device and may include any suitable type of portable or non-portable computing device.

[0042] In any case, vehicle 172 may include a head unit 174, which may include user computing device 102 in some embodiments and / or may house it in other ways. Even if head unit 174 does not include user computing device 102, device 102 can communicate with head unit 174 (e.g., via a wireless or wired connection) to send navigation information such as maps or voice commands and / or information displays to head unit 174 so that head unit 174 can display or emit light. Further, vehicle 172 includes a cluster display unit 151 that can display information transmitted from user computing device 102. In certain embodiments, the user can interact with user computing device 102 by interacting with head unit controls. Additionally, vehicle 172 may provide a communication link 140, which may include, for example, a wired connection to vehicle 172 (e.g., via a USB connection), through which user computing device 102 can send navigation information and corresponding navigation commands for rendering within cluster display unit 151, display 176, and / or for rendering as audio output via speaker 184.

[0043] Accordingly, the head unit 174 may include a display 176 for outputting navigation information such as a digital map. Of course, the cluster display unit 151 may also display such navigation information including a digital map. Such a map rendered within the cluster display unit 151 can provide the driver of the vehicle 172 with navigation instructions at a more optimal position, so that as a result, the driver may not be forced to look away from the active road too much during driving in order to safely navigate to the intended destination. Nevertheless, the display 176 in some embodiments may include a software keyboard for entering text input that can include names or addresses such as a destination, a departure place, etc.

[0044] The hardware input controls 178 and 180 at the head unit 174 and the steering wheel respectively can each be used for entering alphanumeric characters or performing other functions for requesting navigation guidance. For example, the hardware input controls 178, 180 may be, or may include, a rotary control (e.g., a rotary knob), a track pad, a touch screen, and / or any other suitable input control. The head unit 14 may also include audio input and output components such as a microphone 24 and a speaker 26, for example. By way of example, the user computing device 102 may be communicatively connected to the head unit 174 (e.g., via Bluetooth (trademark), WiFi, a cellular communication protocol, a wired connection, etc.) or may be included in the head unit 174. The user computing device 102 can present map information via the cluster display unit 151, issue voice commands for navigation via the speaker 184, and receive input from the user via the head unit 174 (e.g., by the user interacting with the input controls 178 and 180, the display 176, or the microphone 182).

[0045] Examples of conversations / analysis executed to determine filtered navigation search results Regarding the technology of the present disclosure for determining routes and locations through natural conversations, it will be described below with reference to the conversation flow and processing workflow shown in FIGS. 2A-2C. Throughout the description of FIGS. 2A-2C, the actions described as being executed by the user computing device 102 may, in some embodiments, be executed by the external server 120, the vehicle computing device 150, and / or may be executed in parallel by the user computing device 102, the navigation server 120, and / or the vehicle computing device 150. For example, the user computing device 102, the navigation server 120, and / or the vehicle computing device 150 may utilize the language processing modules 109a, 120a, 153a, and / or the machine learning modules 109b, 120c, 153c to determine routes and locations through natural conversations with the user.

[0046] In particular, FIG. 2A shows an exemplary conversation 200 between user 202 and user computing device 102 of FIG. 1A for determining a location and route through natural conversation. User 202 can converse with user computing device 102 by voice to prompt user 202 to provide an explanation in order to determine a narrowed set of navigation search results that enable user 202 to move to the desired destination. That is, user 202 can provide user input to user computing device 102 (transmission to user computing device 102 shown as 204a). The user input can generally include the desired destination of user 202, as well as additional criteria related to routing to the desired destination of user 202 that user 202 includes. For example, user 202 can state "Navigate to ABC Hotel", and user 202 may further state "I don't want to drive for more than 25 minutes". Thus, the user input includes a destination (ABC Hotel) and an additional criterion (travel time of 25 minutes or less).

[0047] Using this user input, user computing device 102 can generate an initial set of navigation search results that meet one or both of the criteria of user 202. For example, the first set of navigation search results may include multiple routes to one or more ABC Hotels and / or multiple routes to different hotels / accommodations within a distance of 25 minutes or less. In any case, user computing device 102 may determine that the number of candidate routes is too large to provide to user 202 and / or, in other cases, may determine that the set of navigation search results should be filtered to provide user 202 with a set of narrowed navigation search results.

[0048] In that case, the user computing device 102 can generate an audio request that is output to the user 202 via a speaker 206 that can be integrated as part of the user computing device 102 (e.g., part of the I / O module 118) (transmission to the user 202 shown as 204b). The audio request can prompt the user 202 to provide additional criteria and / or details corresponding to the destination and / or route desired by the user 202 so that the user computing device 102 can narrow down the set of navigation search results (e.g., via the machine learning module 109b). Continuing with the above example, the audio request transmitted to the user 202 via the speaker 206 can state "What is the address of the ABC Hotel you are using?", and the audio request can further state "Some of the routes involve driving on toll roads. Is that okay?". In this way, the user computing device 102 can request additional information from the user 202 to filter (e.g., exclude) routes that do not meet and / or cannot otherwise meet the additional criteria that may be provided by the user 202 in response to the audio request.

[0049] However, the audio request can provide the user with various clarification options that can be repeated. For example, user 202 may be traveling in Switzerland and can provide user input by saying "Navigate to a nearby hiking spot." User computing device 102 may generate a number of route options for the set of navigation search results that include several different options for reaching the hiking course from a car parking lot. Device 102 may respond to user 202 with an audio request stating, "Some of the top-rated options require taking a cable car from the parking lot. Do you want to do that? The total travel time will probably be less than 30 minutes." In this way, user computing device 102 can provide user 202 with an audio request that can quickly exclude many route options based on the response of user 202 indicating whether taking a cable car from the parking lot is acceptable or not.

[0050] As another example, user computing device 102 may provide an audio request that includes a suggestion to assist user 202 in determining the optimal route among a set of navigation search results. In this example, user 202 may arrive at an airport and may want to navigate to a hotel by asking user computing device 102 "Show me the way to ABC Hotel." User computing device 102 can respond with several different route proposals along with the top candidate by stating "The route I recommend is the shortest route, but it involves driving on a single-lane road for 10 miles." If user 202 has no problem driving on a single-lane road for 10 miles, user 202 may accept the proposed route and thereby end the route search. However, if user 202 rejects the proposed route, user computing device 102 may exclude not only the proposed route but also all routes that include driving on a single-lane road for at least 10 miles. Thus, by suggesting proposed routes based on specified criteria, user computing device 102 may be able to narrow down the set of navigation search results without directly prompting (and potentially distracting) user 202.

[0051] As yet another example, an audio request may be configured to provide clarification of the conversation during a navigation session. Similar to the above example, user 202 may be navigating from an airport to a hotel, and while user 202 is in transit, user 202 may encounter the possibility of a detour that takes a similar amount of time but has different characteristics. As the detour is approached, user computing device 102 may prompt user 202 with an audio request such as "There is an alternative route with a similar ETA to Exit 213A. The distance is shorter, but there are temporary structures and there may be a 3-minute delay. Do you want to take the alternative route?" In response, user 202 may accept or reject the alternative route, and user computing device 102 may continue with the original route or switch to the alternative route as needed. Thus, in this example, the set of navigation search results may include the original route and the alternative route, and the audio request may prompt user 202 to filter the set of navigation search results by determining which of the two routes user 202 prefers. In this way, the audio request provided by user computing device 102 can actively / continuously search for and / or filter the set of navigation search results before / during the navigation session to ensure that user 202 receives an optimal routing experience to the destination.

[0052] Furthermore, it should be noted that the user computing device 102 can generally allow several seconds (e.g., 5 to 10 seconds) for the user 202 to respond after transmitting a voice request via the speaker 206, in order to give the user 202 sufficient time to consider an appropriate response without continuously listening to the interior of the vehicle. By default, the user computing device 102 may not activate the microphone and / or other listening devices (e.g., included as part of the I / O module 118) while the navigation app 108 is running and / or while processing information received via the microphone by, for example, the processor 104, the language processing module 109a, the machine learning module 109b, and / or the OS 110, or in accordance with them. Thus, the user computing device 102 may not actively listen to the interior of the vehicle during a navigation session and / or at any other time, except when the user computing device 102 provides an audio request to the user 202, in which case the user computing device 102 may expect an oral response from the user 202 within a few seconds of transmission.

[0053] In any case, user 202 can receive an audio request and, in response, can provide subsequent user input (transmission to user computing device 102 shown as 204c). The subsequent user input can generally include additional route / destination criteria based on the requested information included as part of the audio request provided by user computing device 102. Continuing with the previous example, user 202 can provide subsequent user input of "ABC Hotel is located at 123 Main Street, Chicago, Illinois" and "No, I want to avoid toll roads" in response to the audio requests "What is the address of ABC Hotel where you are staying?" and "Some of the routes include driving on toll roads. Is that okay?" Thus, in this example, user 202 provides additional location information related to the desired destination and routing information to exclude toll roads that can be used by user computing device 102 to narrow down the set of navigation search results. Accordingly, user computing device 102 may receive the subsequent user input and proceed to generate a narrowed set of navigation search results. User computing device 102 can provide this narrowed set of navigation search results to user 202 as an audio output (e.g., by speaker 206), as a visual output on a display screen (e.g., cluster display unit 151, display 176), and / or as a combination of audio output / visual output.

[0054] As described with reference to FIG. 2A, to better understand the processing performed by user computing device 102, FIG. 2B illustrates a user input analysis sequence 210 for outputting a set of audio requests and navigation search results. User input analysis sequence 210 generally includes the user computing device 102 analyzing / operating on user input during two separate periods 212, 214 to generate two separate outputs. That is, during the first period 212, the user computing device 102 receives user input and proceeds to generate a text transcription of the user input using the language processing module 109a. Thereafter, during the second period 214, the user computing device analyzes the text transcription of the user input to output a set of audio requests and / or navigation search results using the language processing module 109a and / or the machine learning module 109b.

[0055] More specifically, during the first period 212, the user computing device 102 receives user input through an input device (e.g., a microphone as part of the I / O module 118). Next, the user computing device 102 utilizes the processor 104 to execute instructions included as part of the language processing module 109a to transcribe the user input into a set of text. The user computing device 102 can cause the processor 104 to execute instructions including, for example, an ASR engine (e.g., ASR engine 109a1) to transcribe the voice-based input of the user input received by the I / O module 118 into a text transcription of the user input. Of course, as described above, the execution of the ASR engine to transcribe the user input into a text transcription (as well as any other actions described with reference to FIGS. 2B and 2C) can be performed by the user computing device 102, the external server 120, the vehicle computing device 150, and / or other suitable components or combinations thereof.

[0056] This transcription of the user input can then be analyzed during a second period 214 by a processor 104 that executes instructions including, for example, a language processing module 109a and / or a machine learning module 109b to output, for example, a set of audio requests and / or navigation search results. Specifically, the instructions including the language processing module 109a and / or the machine learning module 109b can cause the processor 104 to interpret the text transcription to determine the user's intent along with values corresponding to a destination and / or other constraints. For example, the user's intent can include a movement to a desired destination, and the value of the destination can correspond to a specific location (e.g., Chicago, Illinois) or a general location (e.g., a nearby hiking trail), and the other constraints can include other details corresponding to the user's intent (e.g., "by car," "within a distance of less than 10 miles" to move, etc.).

[0057] When the destination is generally described (e.g., "a nearby restaurant"), to determine the value of the destination from the user input, the user computing device 102 may first parse and extract this destination information from the user input. Next, the user computing device 102 can access a database (e.g., a map database 156, a POI database 159) or other appropriate repository to search for corresponding locations by fixing the search to the user's current location and / or viewport. Next, the user computing device 102 may identify candidate destinations and routes to each candidate destination based on the similarity between the locations in the repository and the destination determined from the user input, thereby creating an initial set of navigation search results.

[0058] However, before determining whether to generate an audio request, user computing device 102 can prune this initial set of navigation search results by excluding candidate destinations and routes that do not match other details corresponding to the user's intent and / or are not appropriately addressed in other ways. For example, if a candidate destination is further away than a distance specified by the user as a maximum distance in the user input, the candidate destination can be excluded from the initial set of navigation search results. Further, each destination / route can receive a score corresponding to, for example, the overall similarity of the destination / route to values extracted from the user input.

[0059] Once user computing device 102 determines and filters / prunes the initial set of navigation search results to generate a set of navigation search results, device 102 can proceed to determine whether to provide an audio output to the user. User computing device 102 can make this determination based on several criteria, such as (i) the total number of routes / destinations that would be provided to the user as part of the set of navigation search results, (ii) the device type and / or surface type (e.g., smartphone, tablet, wearable device, etc.) that the user is using to receive navigation instructions, (iii) the entry point and / or input type (e.g., voice-based input, touch-based input) that the user used to enter the user input, (iv) whether the scores corresponding to the destinations / routes included in the set of navigation results are high enough (e.g., relative to a score threshold), and / or any other suitable criteria or combinations thereof.

[0060] For example, the user computing device 102 can determine that the total number of routes included as part of a set of navigation search results is 20 and the route presentation threshold is 15. In some examples, the route presentation threshold is set based on a determination of the computational cost involved in providing the set of results. For example, in this example, providing a set of 16 results exceeds the threshold, and providing this set of results requires more computational resources compared to a set of results less than the threshold amount. As a result, the user computing device 102 compares the total number of routes with the route presentation threshold and determines that the total number of routes does not meet the route presentation threshold and that an audio request should be generated. Thus, if any of the above criteria are applied by the user computing device 102 and any of the applied criteria do not meet their respective thresholds (e.g., route presentation threshold, score threshold) and / or have their respective values (e.g., device type, input type) that require an audio request, the device 102 can generate an audio request.

[0061] In response to determining that an audio request should be generated, the user computing device 102 can proceed to generate an audio request, for example, using the language processing module 109a. The user computing device 102 can generally proceed to generate an audio request by considering which audio request reduces the number of destinations / routes included in the set of navigation search results the most. That is, the user computing device 102 can analyze the attributes corresponding to each destination / route to determine which attribute is the most common among the destinations / routes included in the set of navigation search results, and can generate an audio request based on one or more of these most common attributes.

[0062] As an example, a set of navigation search results may include 20 route options to a specific destination, and each route option may have a distance to travel to reach the specific destination that is substantially different from all other route options. Thus, the user computing device 102 can generate an audio request that prompts the user to provide a distance requirement in order to most efficiently narrow down the set of navigation search results by excluding routes that do not meet the user's distance requirement.

[0063] As another example, a set of navigation search results may include eight route options to a specific destination, and each route option may be substantially different from all other route options in terms of the type of road (e.g., highway, country road, scenic route, urban area) that the user may drive on to reach the specific destination. Thus, the user computing device 102 can generate an audio request that prompts the user to provide a preference for road type in order to most efficiently narrow down the set of navigation search results by excluding routes that do not meet the user's road type preference.

[0064] The user computing device 102 can generate the text of the audio request by utilizing the language processing module 109a and, in certain aspects, a large language model (LLM) (e.g., Language Model for Dialog Applications (LaMDA)) (not shown) included as part of the language processing module 109a. Such an LLM can be conditioned / trained to generate audio request text based on specific most common attributes of the set of navigation search results, and / or the LLM can be trained to receive a natural language representation of a candidate route / destination as input and output a set of texts representing audio requests based on the most common attributes.

[0065] In any case, once the user computing device 102 has fully generated the text of the audio request, the device 102 may proceed to synthesize the text into speech for the audio output of the request to the user. Specifically, the user computing device 102 can send the text of the audio output to a TTS engine (e.g., TTS engine 109a2) to aurally output the audio request via a speaker (e.g., speaker 206) so that the user can listen to and interpret the audio output. Additionally or alternatively, the user computing device 102 can also visually prompt the user by displaying the text of the audio request on a display screen (e.g., cluster display unit 151, display 176), so that the user can interact with the display screen (e.g., click, tap, swipe, etc.) and / or verbally respond to the audio request.

[0066] When the user receives an audio request from the user computing device 102, the user can provide subsequent user input. This user computing device 102 may proceed to receive this subsequent user input and narrow down a set of navigation search results as shown in Figure 2C. More specifically, Figure 2C shows a subsequent user input analysis sequence 220 for outputting a set of narrowed-down navigation search results. The subsequent user input analysis sequence 220 generally includes the user computing device 102 analyzing / operating on subsequent user input during two separate periods 222, 224 to generate two separate outputs. That is, during the first period 222, the user computing device 102 proceeds to receive subsequent user input and generate a text transcription of the subsequent user input using the language processing module 109a. Thereafter, during the second period 224, the user computing device analyzes the text transcription of the subsequent user input to output a set of narrowed-down navigation search results using the language processing module 109a and / or the machine learning module 109b.

[0067] More specifically, during the first period 222, the user computing device 102 receives user input through an input device (e.g., a microphone as part of the I / O module 118). Next, the user computing device 102 utilizes the processor 104 to execute instructions included as part of the language processing module 109a to transcribe the subsequent user input into a set of text. The user computing device 102 can cause the processor 104 to execute instructions including, for example, an ASR engine (e.g., the ASR engine 109a1) to transcribe the subsequent user input from the audio-based input received by the I / O module 118 into a text transcription of the subsequent user input.

[0068] This transcription of subsequent user input can then be analyzed during a second period 224 by a processor 104 that executes instructions including, for example, a language processing module 109a and / or a machine learning module 109b to output a set of filtered navigation search results. Specifically, the instructions including the language processing module 109a and / or the machine learning module 109b can cause the processor 104 to interpret the text transcription of the subsequent user input to determine a subsequent user intention along with values corresponding to a filtered destination and / or other constraints. For example, the subsequent user intention can include determining whether the subsequent user input is related to an audio request, the value of the filtered destination can correspond to a specific location (e.g., Chicago, Illinois) or a general location (e.g., a nearby hiking trail), and the other constraints can include other details corresponding to the subsequent user intention (e.g., traveling "by car", traveling "less than 10 miles", etc.).

[0069] When the user computing device 102 receives subsequent user input and determines a subsequent user intention and a value of a filtered destination and / or other constraints, the device 102 can narrow down / filter a set of navigation search results by excluding candidate destinations and routes that do not match and / or otherwise do not appropriately correspond to the subsequent user intention, the value of the filtered destination, and / or other details corresponding to the other constraints. Further, each destination / route included in the set of navigation search results can receive a score (e.g., from the machine learning module 109b) corresponding to the overall similarity of the destination / route to values extracted from the subsequent user input. As an example, if a candidate route receives a score of 35 due to relative dissimilarity to values extracted from the subsequent user input and the score threshold for remaining part of the set of navigation results is 75, the candidate route can be excluded from the set of navigation search results.

[0070] Generally speaking, user computing device 102 can repeat any suitable number of the actions described herein with reference to FIGS. 2B and 2C to provide the user with a set of filtered navigation search results. For example, after receiving subsequent user input, user computing device 102 may determine that subsequent audio output needs to be provided to the user. Thus, in this example, user computing device 102 may proceed to generate a subsequent audio request to the user as described above with reference to FIG. 2B. User computing device 102 can then receive further user input in response to the subsequent audio request and proceed to further filter the set of navigation search results until the criteria used by device 102 to determine whether to generate an audio request are met.

[0071] In any case, when user computing device 102 determines that all criteria corresponding to generating an audio request are met, device 102 can determine that the set of navigation search results is a set of filtered navigation search results suitable for providing to the user. Thus, user computing device 102 can proceed to provide the user with the set of filtered navigation search results as an audio output and / or visual display. The filtered set of navigation search results, when provided to the user, may include any suitable information corresponding to each route, such as total travel distance, total travel time, number of road changes / branches, and / or other suitable information or combinations thereof. Further, all information included as part of each route of the filtered set of navigation search results can be provided to the user as an audio output (e.g., via speaker 206) and / or as a visual display on the display screen of any suitable device (e.g., I / O module 118, cluster display unit 151, display 176).

[0072] Of course, the user may decide to further narrow down the set of navigation search results and independently provide an input to the user computing device 102 to that effect (e.g., without being prompted by the user computing device 102). In certain embodiments, the user may provide a user input that includes a specific trigger phrase or word, whereupon the user computing device 102 receives user input for a certain duration following the user input that includes the trigger phrase / word. The user may initiate the input collection of the user computing device 102 in this or a similar manner, and the device 102 can proceed with receiving and interpreting the user input in the same manner as described above with reference to FIGS. 2A-2C. For example, the user may independently say, "I prefer a shorter distance route and intend to drive on single-lane roads." The user computing device 102 can receive this user input and proceed to narrow down the set of navigation search results to provide to the user as described above.

[0073] Examples of conversations / analysis performed to provide navigation instructions When user computing device 102 successfully generates a filtered set of navigation search results, the user may examine the results to determine an optimal route to a desired destination. To illustrate actions performed by user computing device 102 as part of a route approval process, FIGS. 3A and 3B show an exemplary route acceptance and route adjustment sequence that provides input for selecting / adjusting a route included as part of a set of filtered navigation search results selected by the user. As described above, each route included as part of a set of filtered navigation search results includes a turn-by-turn sequence of directions to a destination as part of a navigation session. As described herein, the user can provide input regarding acceptance and / or adjustment of a route included in a set of filtered navigation search results such that the turn-by-turn sequence of directions provided by user computing device 102 during a navigation session is also changed in response to user input regarding the currently accepted route.

[0074] More specifically, FIG. 3A shows an exemplary transition 300 between the user 202 providing a route acceptance input and the user computing device 102 displaying navigation instructions corresponding to the accepted route. The user 202 may provide a route acceptance input to the user computing device 102 indicating acceptance of a route included as part of a set of filtered navigation search results. Next, the user computing device 102 can receive the route acceptance input and proceed to initiate a navigation session that includes turn-by-turn navigation instructions corresponding to the accepted route. Thus, the user computing device 102 may not only provide verbal turn-by-turn instructions to the user 202, but may also proceed to render turn-by-turn instructions on the display screen 302 of the device 102 for the user 202 to view.

[0075] During a navigation session, the user computing device 102 can display a map via the display screen 302 depicting, among other things, the position of the user computing device 102, the direction of travel of the user computing device 102, the estimated arrival time, the estimated distance to the destination, the estimated travel time to the destination, the current navigation direction, one or more future navigation directions of a set of navigation instructions corresponding to the accepted route, one or more user-selectable options for changing the display or adjusting the navigation direction, and the like. The user computing device 102 can also issue audio instructions corresponding to the set of navigation instructions.

[0076] As an example, the user computing device 102 can provide the user 202 with a set of filtered navigation search results that include three candidate routes to the desired destination of the user 202. The user 202 may provide a route acceptance input indicating that the user 202 wishes to take a first candidate route included as part of the set of filtered navigation search results. The user computing device 102 can receive this route acceptance input from the user 202, provide a first navigation instruction included as part of the first candidate route (referred to herein as the "accepted route" in this example), and proceed to render on the display screen 302 a map that includes a visual representation of the first navigation instruction. As the user 202 moves along the accepted route, the user computing device 102 can provide the user 202 with successive navigation instructions (e.g., first, second, third) both orally and visually as the user 202 approaches each waypoint along the accepted route to enable the user 202 to follow the accepted route. When the user 202 reaches the destination, which is the end of the accepted route, the user computing device 102 can deactivate the navigation session.

[0077] However, in certain situations, user 202 may wish to change from the accepted route to an alternative route and / or may be forced to do so. FIG. 3B shows an exemplary route update sequence 320 for updating the navigation instructions provided to user 202 by prompting user 202 with the option to switch to an alternative route. Specifically, user computing device 102 may actively participate in the navigation session initiated by user 202 when user computing device 102 determines that the alternative route may be a more optimal route than the accepted route. User computing device 102 may make such a determination based on, for example, updated traffic information along the accepted route (e.g., from traffic database 157) and / or any other suitable information.

[0078] Based on this determination, user computing device 102 may generate an alternative route output (a transmission to user 202 shown by 322a) that provides user 202 with the option to adjust the current navigation session to follow the alternative route. For example, user computing device 102 may verbally provide the alternative route output through speaker 206 and / or may visually indicate the alternative route output through prompt 324. As shown in FIG. 3B, the alternative route output may state, "There is an alternative route that shortens the travel time by 10 minutes. Do you want to switch to the alternative route?" This statement may be presented verbally to user 202 via speaker 206 and may also be presented visually via display screen 322.

[0079] If user 202 determines to provide an oral user input (transmission to user computing device 102 as indicated by 322b), user 202 can verbally respond to the alternative route output within a short period (e.g., 5 - 10 seconds) after the output is provided to user 202 for the user computing device 102 to receive the oral user input. The user computing device 102 may receive the oral user input and proceed to process / analyze the oral user input in the same manner as the analysis described herein with reference to FIGS. 2A - 2C. Specifically, if user 202 determines to accept the alternative route, the user computing device 102 can start an updated navigation session and provide alternative turn-by-turn navigation instructions based on the alternative route. Alternatively, if user 202 determines to reject the alternative route, the user computing device 102 may continue to provide turn-by-turn navigation instructions corresponding to the accepted route and may not start an updated navigation session.

[0080] Furthermore, the visual rendering of the alternative route output may include interactive buttons 324a, 324b that enable user 202 to physically interact with the display screen 322 to accept or reject switching to the alternative route. When the user receives the prompt 324, the user can interact with the prompt 324 by pressing, clicking, tapping, swiping, etc. on one of the interactive buttons 324a, 324b. If the user selects the "Yes" interactive button 324a, the user computing device 102 can instruct the navigation application 108 to generate and render turn-by-turn navigation guidance as part of an updated navigation session corresponding to the alternative route. If the user selects the "No" interactive button 324b, the user computing device 102 may continue to generate and render turn-by-turn navigation instructions corresponding to the accepted route and may not generate / render an updated navigation session.

[0081] Exemplary Logic for Determining Location and Route by Natural Conversation FIG. 4 is a flow diagram of an exemplary method 400 for determining a location and a route through natural conversation, which can be implemented on a computing device such as the user computing device 102 of FIG. 1. For ease of explanation only, it should be understood that the "user computing device" described herein with reference to FIG. 4 may correspond to the user computing device 102. Further, throughout the description of FIG. 4, the actions described as being performed by the user computing device 102 may, in some embodiments, be performed by the external server 120, the vehicle computing device 150, and / or may be performed in parallel by the user computing device 102, the navigation server 120, and / or the vehicle computing device 150. For example, the user computing device 102, the navigation server 120, and / or the vehicle computing device 150 may utilize the language processing modules 109a, 120a, 153a, and / or the machine learning modules 109b, 120c, 153c to determine a route and a location through natural conversation with the user.

[0082] Referring to FIG. 4, the method 400 can be implemented by a user computing device (e.g., the user computing device 102). The method 400 can be stored in a computer-readable memory and implemented as a set of instructions executable by one or more processors of a user computing device (e.g., the processor(s) 104).

[0083] At block 402, method 400 includes receiving, from a user, voice input that includes a search query to initiate a navigation session (block 402). Method 400 can further include an optional step of transcribing the voice input into a set of text (block 404). In certain embodiments, method 400 can further include, by one or more processors, parsing the set of text to determine a destination value and extracting the destination value from the set of text. Further, in these embodiments, method 400 can include searching for the destination value in a destination database (e.g., map database 156, external server 120) and identifying a plurality of destinations based on the result of searching the destination database. Thus, in these embodiments, method 400 can further include generating one or more routes to each of the plurality of destinations.

[0084] Method 400 also includes generating, in response to the search query, a set of navigation search results (block 406). The set of navigation search results can include a plurality of destinations or a plurality of routes corresponding to the plurality of destinations. In some embodiments, generating a set of navigation search results in response to the search query further includes transcribing the voice input into a set of text and applying a machine learning (ML) model to the set of text to output user intent and destinations. In these embodiments, the ML model can be trained using one or more training datasets of text to output one or more training intents and one or more training destinations.

[0085] In certain aspects, generating a set of navigation search results responsive to a search query further includes generating one or more candidate routes to each of a plurality of destinations based on a respective set of attributes for each candidate route of the one or more candidate routes. In these aspects, each respective set of attributes may include one or more of (i) a means of transportation, (ii) the number of changes, (iii) the total travel distance, (iv) the total travel time, (v) the total travel distance on each road included, or (vi) the total travel time on each road included.

[0086] Method 400 further includes providing an audio request to the user to narrow down the set of navigation search results (block 408). In some aspects, method 400 may further include determining whether to provide an audio request to the user based on at least one of (i) the total number of routes included in the plurality of routes, (ii) the device type of the device used by the user to provide voice input, (iii) the type of input provided by the user, or (iv) the second number of routes included in the plurality of routes that meet a quality threshold.

[0087] In certain aspects, method 400 may include verbally communicating an audio request for the user's consideration by a text-to-speech (TTS) engine (e.g., TTS engine 109a2). Further, in some aspects, providing an audio request to the user to narrow down the set of navigation search results further includes determining the primary attributes of the plurality of routes that result in the greatest reduction of the plurality of routes and generating an audio request to the user based on the primary attributes. In certain aspects, providing an audio request to the user to narrow down the set of navigation search results may further include generating an audio request based on the attributes of the plurality of routes by running a large language model (LLM).

[0088] Method 400 further includes receiving, in response to the audio request, subsequent voice input from the user that includes a refined search query (block 410). In certain aspects, method 400 may further include recognizing the voice input and the subsequent voice input based on a trigger phrase included as part of both the voice input and the subsequent voice input.

[0089] Method 400 may further include an optional step of filtering a set of navigation search results based on subsequent user input (block 412). That is, in certain aspects, method 400 includes transcribing the voice input into a set of text, (a) providing the user with an audio request to narrow a set of navigation search results, (b) receiving, in response to the audio request, subsequent voice input from the user that includes a refined search query, and (c) filtering the set of navigation search results to generate one or more refined navigation search results by excluding routes from among a plurality of routes based on the subsequent voice input. In certain aspects, filtering the set of navigation search results to generate one or more refined navigation search results further includes excluding routes within a set of routes having respective relevance scores that do not meet a relevance threshold based on a natural language transcription of the subsequent voice input by executing a machine learning (ML) model.

[0090] Furthermore, in these aspects, the natural language transcription may not be parsed, and the ML model may be configured to receive the natural language transcription and the route as inputs to output a relevance score for each route. Alternatively, the ML model may be trained using the transcription string of the audio input and the training routes to output a relevance score corresponding to each training route. The relevance score may generally indicate how relevant a particular route is based on the transcription string of the user input. In this way, instead of parsing the user input to extract explicit attributes, the ML model can operate in a more "end-to-end" manner by determining the relevance score of each route based on the user's input. For example, the ML model may receive as inputs the natural language transcription of a subsequent user input such as "It would be better if there were no single-lane roads" and two routes from a set of routes. The first route may include a navigation instruction that instructs the user to move along a series of single-lane roads, and the second route may include a navigation instruction that instructs the user not to move along single-lane roads.

[0091] Continuing with the above example, the ML model may output relevance scores for two routes that may indicate either relevance as an indicator of route executability or route non - executability. That is, since the first route includes a series of single - lane roads, the ML model may output a relatively high (e.g., 9 out of 10) relevance score for the first route, and since the second route does not include single - lane roads, the ML model may output a relatively low (e.g., 1 out of 10) relevance score for the second route. Thus, the relevance score may indicate route non - executability because the first route has a high relevance score based on including a series of single - lane roads (which the user does not desire), while the second route has a low relevance score based on not including single - lane roads (which the user prefers). Alternatively, since the first route includes a series of single - lane roads, the ML model may output a relatively low (e.g., 1 out of 10) relevance score for the first route, and since the second route does not include single - lane roads, the ML model may output a relatively high (e.g., 9 out of 10) relevance score for the second route. Thus, the relevance score may indicate route executability because the first route has a low relevance score based on including a series of single - lane roads (which the user does not desire), while the second route has a high relevance score based on not including single - lane roads (which the user prefers).

[0092] Method 400 may also include an optional step (block 414) of determining whether to provide the user with subsequent audio requests based on one or more filtered navigation search results. Specifically, optionally, user computing device 102 may determine whether the set of navigation search results meets a route presentation threshold (block 416). If user computing device 102 determines that the set of navigation search results does not meet the route presentation threshold (the no branch of block 416), method 400 returns to block 408, and user computing device 102 can provide the user with subsequent audio requests. However, if user computing device 102 determines that the set of navigation search results meets the route presentation threshold (the yes branch of block 416), method 400 can proceed to block 418. It should be understood that method 400 may include repeatedly executing each of blocks 408-416 (and / or any other block of method 400) any suitable number of times until one or more filtered navigation search results meet the route presentation threshold.

[0093] In any case, method 400 further includes providing one or more filtered navigation search results in response to a filtered query that includes a subset of multiple destinations or multiple routes (block 418). In some aspects, method 400 may further include providing one or more filtered navigation search results for the user to view in the user interface.

[0094] In certain aspects, method 400 may include receiving, from a user, a verbal route acceptance input indicating an accepted route from one or more filtered navigation search results. In these aspects, method 400 may further include, at a user interface, displaying the accepted route for viewing by the user and initiating a navigation session along the accepted route by providing verbal navigation instructions corresponding to the accepted route as the user moves along the accepted route.

[0095] In some aspects, providing one or more filtered navigation search results in response to a filtered query may further include generating, by executing a large language model (LLM), a text summary for each route of a subset of multiple routes. Further, in these aspects, method 400 may include, at a user interface, providing the subset of multiple routes and each respective text summary for viewing by the user.

[0096] In certain aspects, method 400 may further include receiving, from a user, a selection of the accepted route to initiate a navigation session along the accepted route. Further, in these aspects, method 400 may include, during the navigation session, determining that an alternative route improves at least one of (i) the user's arrival time, (ii) the user's travel distance, or (iii) the user's time on a particular road. Method 400 may also include, during the navigation session, prompting the user with an option to switch from the selected route to the alternative route via either a verbal prompt or a text prompt.

[0097] Aspects of the present disclosure 1. A method in a computing device for determining a location and a route through natural conversation, the method comprising: receiving, from a user, an audio input including a search query to start a navigation session; generating, by one or more processors, a set of navigation search results corresponding to the search query, the set of navigation search results including a plurality of destinations and / or a plurality of routes corresponding to one or more destinations; providing, by the one or more processors, an audio request to the user to narrow down the set of navigation search results; receiving, in response to the audio request, a subsequent audio input from the user including a narrowed search query; and providing, by the one or more processors, one or more narrowed navigation search results corresponding to the narrowed search query and including a subset of the plurality of destinations and / or the plurality of routes.

[0098] 2. Transcribing, by an automatic speech recognition (ASR) engine, the audio input into a set of text; (a) providing, by the one or more processors, the audio request to the user to narrow down the set of navigation search results; (b) receiving, in response to the audio request, the subsequent audio input from the user including the narrowed search query; (c) filtering, by the one or more processors, the set of navigation search results to generate the one or more narrowed navigation search results by excluding routes among the plurality of routes based on the subsequent audio input; (d) determining, by the one or more processors, whether to provide a subsequent audio request to the user based on the one or more narrowed navigation search results; and (e) repeating (a) to (e) until the one or more narrowed navigation search results meet a threshold. The method according to aspect 1 further includes this.

[0099] 3. Filtering the set of navigation search results to generate the one or more filtered navigation search results includes the one or more processors executing a machine learning (ML) model to exclude, based on a natural language transcription of the subsequent voice input, each route in a set of routes having a relevance score that does not meet a relevance threshold, wherein the natural language transcription is not parsed and the ML model is configured to receive the natural language transcription and the route as inputs to output a relevance score for each route, the excluding, the method of aspect 2 further comprising.

[0100] 4. The method according to any one of aspects 1 to 3, further comprising providing, on a user interface, the one or more filtered navigation search results for viewing by the user.

[0101] 5. Determining, by the one or more processors, whether to provide the audio request to the user based on at least one of (i) a total number of routes included in the plurality of routes, (ii) a device type of a device used by the user to provide the voice input, (iii) an input type provided by the user, or (iv) a second number of routes included in the plurality of routes that meet a quality threshold, the method according to any one of aspects 1 to 4 further comprising.

[0102] 6. The method according to any one of aspects 1 to 5, further comprising orally communicating the audio request for consideration by the user by a text-to-speech (TTS) engine.

[0103] 7. Further comprising receiving, from the user, an oral route acceptance input indicating an accepted route from the one or more filtered navigation search results; displaying, on the user interface, the accepted route for viewing by the user; and starting the navigation session along the accepted route by providing, by the one or more processors, an oral navigation instruction corresponding to the accepted route when the user moves along the accepted route, the method according to any one of Aspects 1 to 6.

[0104] 8. Generating the set of navigation search results in response to the search query comprises transcribing the voice input into a set of text and applying, by the one or more processors, a machine learning (ML) model to the set of text to output user intent and a destination, wherein the ML model is trained using one or more training datasets of text to output one or more training intents and one or more training destinations, the method according to any one of Aspects 1 to 7.

[0105] 9. Further comprising transcribing the voice input into a set of text; parsing, by the one or more processors, the set of text to determine a value of a destination; extracting, by the one or more processors, the value of the destination from the set of text; and searching, by the one or more processors, for the value of the destination in a destination database, the method according to any one of Aspects 1 to 8.

[0106] 10. Further comprising identifying, by the one or more processors, the plurality of destinations based on a result of searching the destination database; and generating, by the one or more processors, one or more routes to each destination of the plurality of destinations, the method according to Aspect 9.

[0107] 11. Generating the set of navigation search results in response to the search query comprises, by the one or more processors, generating, for each destination of the plurality of destinations, the one or more candidate routes to the respective destination based on respective sets of attributes for each candidate route of the one or more candidate routes, wherein each respective set of attributes includes one or more of (i) means of transportation, (ii) number of changes, (iii) total travel distance, (iv) total travel time, (v) total travel distance on each road included, or (vi) total travel time on each road included, the method according to any of aspects 1-10, further comprising the generating.

[0108] 12. Providing the audio request to the user to narrow down the set of navigation search results further comprises, by the one or more processors, determining the primary attributes of the plurality of routes that would result in the greatest reduction of the plurality of routes, and generating, by the one or more processors, the audio request to the user based on the primary attributes, the method according to any of aspects 1-11.

[0109] 13. Providing the audio request to the user to narrow down the set of navigation search results further comprises generating, by the one or more processors executing a large language model (LLM), the audio request based on the attributes of the plurality of routes, the method according to any of aspects 1-12.

[0110] 14. Providing one or more filtered navigation search results in response to the filtered search query includes the one or more processors executing a large language model (LLM) to generate a text summary for each route of the subset of the plurality of routes, and providing, at the user interface, the subset of the plurality of routes and each respective text summary for viewing by the user. The method according to any one of aspects 1-13 further includes this.

[0111] 15. Receiving a selection of the accepted route from the user to start the navigation session moving along the accepted route, and during the navigation session, determining that an alternative route improves at least one of (i) the user's arrival time, (ii) the user's travel distance, or (iii) the user's time on a particular road, and during the navigation session, prompting the user with an option to switch from the selected route to the alternative route via either an oral prompt or a text prompt. The method according to any one of aspects 1-14 further includes this.

[0112] 16. Further including recognizing the voice input and the subsequent voice input by the one or more processors based on a trigger phrase included as part of both the voice input and the subsequent voice input. The method according to any one of aspects 1-15 further includes this.

[0113] 17. A computing device for determining a location and a route through natural conversation, comprising: a user interface; one or more processors; a computer-readable memory, optionally non-transitory, coupled to the one or more processors and storing instructions, wherein when the instructions are executed by the one or more processors, the computing device is caused to: receive, from a user, an audio input including a search query to start a navigation session; generate a set of navigation search results corresponding to the search query, the set of navigation search results including a plurality of destinations or a plurality of routes corresponding to one or more destinations; provide an audio request to the user to narrow down the set of navigation search results; receive, in response to the audio request, a subsequent audio input from the user including a refined search query; and provide one or more refined navigation search results corresponding to the refined search query, the one or more refined navigation search results including a subset of the plurality of destinations or the plurality of routes.

[0114] 18. When the command is executed by the one or more processors, the computing device is caused to: transcribe the voice input into a set of text by an automatic speech recognition (ASR) engine; (a) provide the audio request to the user to narrow down the set of navigation search results; (b) receive, in response to the audio request, subsequent voice input from the user that includes the narrowed search query; (c) filter the set of navigation search results to generate one or more narrowed navigation search results by excluding a route among the plurality of routes based on the subsequent voice input; (d) determine whether to provide a subsequent audio request to the user based on the one or more narrowed navigation search results; and (e) repeatedly execute (a) to (e) until the one or more narrowed navigation search results meet a threshold. The computing device according to aspect 17.

[0115] 19. A computer-readable medium, optionally non-transitory, storing instructions for determining a location and a route through natural conversation, the instructions, when executed by one or more processors, causing the one or more processors to: receive voice input from a user that includes a search query to start a navigation session; generate a set of navigation search results corresponding to the search query, the set of navigation search results including a plurality of destinations or a plurality of routes corresponding to one or more destinations; provide an audio request to the user to narrow down the set of navigation search results; receive, in response to the audio request, subsequent voice input from the user that includes the narrowed search query; and provide one or more narrowed navigation search results corresponding to the narrowed search query that include a subset of the plurality of destinations or the plurality of routes. The computer-readable medium.

[0116] 20. When the command is executed by the one or more processors, the one or more processors are caused to transcribe the voice input into a set of text by an automatic speech recognition (ASR) engine, and (a) provide the audio request to the user to narrow down the set of navigation search results; (b) receive, in response to the audio request, subsequent voice input from the user that includes the narrowed search query; (c) filter the set of navigation search results to generate one or more narrowed navigation search results by excluding routes among the plurality of routes based on the subsequent voice input; (d) determine whether to provide a subsequent audio request to the user based on the one or more narrowed navigation search results; and (e) repeatedly execute (a) to (e) until the one or more narrowed navigation search results meet a threshold. The computer-readable medium according to aspect 19, which further causes the above to be performed.

[0117] 21. A computing device for determining a location and a route through a natural conversation, comprising a user interface, one or more processors, and a non-transitory computer-readable memory coupled to the one or more processors and storing instructions, wherein when the instructions are executed by the one or more processors, the computing device is caused to perform any of the methods disclosed herein. The computing device.

[0118] 22. A tangible non-transitory computer-readable medium storing instructions for determining a location and a route through a natural conversation, wherein when the instructions are executed by one or more processors, the one or more processors are caused to perform any of the methods disclosed herein. The non-transitory computer-readable medium.

[0119] 23. A method in a computing device for determining a location and a route through natural conversation, the method comprising: receiving an input from a user to initiate a navigation session; generating, by one or more processors, one or more destinations or one or more routes in response to the user input; providing, by the one or more processors, a request to the user for narrowing down a response to the user input; receiving, in response to the request, subsequent input from the user; and providing, by the one or more processors, one or more updated destinations or one or more updated routes in response to the subsequent user input.

[0120] 24. The method of aspect 23, wherein the user input is an audio input or a text input, and the request is an audio request or a text request.

[0121] Other considerations The following additional considerations apply to the foregoing discussion. Throughout this specification, multiple examples may implement components, operations, or structures described as a single example. Individual operations of one or more methods are illustrated and described as separate operations, but one or more of the individual operations may be performed together and nothing requires that the operations be performed in the order illustrated. Structures and functionality represented as separate components in example configurations may be implemented as a combined structure or component. Similarly, structures and functionality represented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements are within the scope of the subject matter of this disclosure.

[0122] In addition, an embodiment is described herein as including logic or several components, modules, or mechanisms. A module can constitute either a software module (e.g., code stored in a machine-readable medium) or a hardware module. A hardware module is a tangible unit capable of performing certain operations and can be configured or arranged in a certain manner. In an example of an embodiment, one or more computer systems (e.g., stand-alone, client, or server computer systems) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) can be configured by software (e.g., an application or a portion of an application) as a hardware module that operates to perform certain operations as described herein.

[0123] In various embodiments, a hardware module can be implemented mechanically or electronically. For example, a hardware module can comprise dedicated circuitry or logic that is permanently configured (e.g., as a special-purpose processor, as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC)) to perform certain operations. A hardware module can also include programmable logic or circuitry (e.g., included within a general-purpose processor or other programmable processor) that is temporarily configured by software to perform certain operations. It is recognized that the decision to implement a hardware module mechanically, within dedicated and permanently configured circuitry, or within temporarily configured circuitry (e.g., configured by software) is to be left to cost and time considerations.

[0124] Accordingly, the term "hardware" should be understood to be components that are physically configured to include tangible components, to operate in a certain manner, or to perform certain operations described herein, and are either permanently configured (e.g., wired), or temporarily configured (e.g., programmed). As used herein, "hardware implementation module" refers to a hardware module. Considering embodiments where a hardware module is temporarily configured (e.g., programmed), each of the hardware modules need not be configured or instantiated within a time for any one instance. For example, if a hardware module includes a general-purpose processor configured using software, the general-purpose processor can be configured as each different hardware module at different times. Software can thus configure the processor, for example, to form a particular hardware module at a time of one instance, and to form different hardware modules at times of different instances.

[0125] A hardware module can provide information to other hardware modules and receive information from other hardware. Thus, the described hardware modules can be treated as being communicatively coupled. When multiple such hardware modules are present simultaneously, communication can be achieved through signal transmission (e.g., with appropriate circuitry and buses) that connects the hardware modules. In embodiments where multiple hardware modules are configured or instantiated at different times, communication between such hardware modules can be achieved, for example, through storage and retrieval of information within a memory structure to which the multiple hardware modules have access rights. For example, one hardware module may perform a particular operation and store the output of that operation in a communicatively coupled memory device. Another hardware module may then subsequently access this memory device and retrieve and process the stored output. A hardware module may also initiate communication with an input or output device and operate on a particular resource (e.g., a collection of information).

[0126] Method 400 may include one or more functional blocks, modules, individual functions, or routines in the form of tangible computer-executable instructions stored on a computer-readable storage medium, optionally a non-transitory computer-readable storage medium, and executed using a processor of a computing device (e.g., a server device, a personal computer, a smartphone, a tablet computer, a smartwatch, a mobile computing device, or other client computing device as described herein). Method 400 may be included, for example, as part of any backend server (e.g., a map data server, a navigation server, or any other type of server computing device as described herein), as part of a client computing device module of an exemplary environment, or as part of a module external to such an environment. The figures may be described with reference to other figures for ease of explanation, but Method 400 may be utilized with other objects and user interfaces. Further, in the above description, steps of Method 400 executed by a particular device (such as a user computing device) are described, but this is done for illustrative purposes only. Blocks of Method 400 may be executed by one or more devices or other parts of the environment.

[0127] The various operations of an example method described herein may be at least partially executed by one or more processors temporarily configured (e.g., by software) or permanently configured to execute the relevant operations. Whether temporarily or permanently configured, such processors may form processor-implemented modules that operate to perform one or more operations or functions. The modules referred to herein may, in some example embodiments, include processor-implemented modules.

[0128] Similarly, the methods or routines described herein may be implemented, at least in part, by a processor. For example, at least some of the operations of the method may be performed by one or more processors or hardware modules implemented with processors. The implementation of certain of the operations may not only reside within a single machine, but may also be distributed among one or more processors deployed across several machines. In examples of some embodiments, the processor or processors may be placed in a single location (e.g., within a home environment, a workplace environment, or as a server farm), while in other embodiments, the processors may be distributed across several locations.

[0129] One or more processors may also operate to support related operations as “cloud computing” environments or as SaaS. For example, as shown above, at least some of the operations may be performed by a group of computers (as an example of a machine including a processor), and these operations are accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., an API).

[0130] Furthermore, the figures show some embodiments of exemplary environments for illustrative purposes only. Those skilled in the art will readily recognize from the following discussion that alternative embodiments of the structures and methods described herein may be employed without departing from the principles described herein.

[0131] Upon reading this disclosure, those skilled in the art will understand, through the principles disclosed herein, yet another alternative structural design and functional design for determining locations and routes through natural conversations. Thus, while specific embodiments and applications are shown and described, it is to be understood that the disclosed embodiments are not limited to the exact configurations and components disclosed herein. Various obvious modifications, changes, and variations can be made by those skilled in the art in the arrangement, operation, and details of the methods and apparatuses disclosed herein without departing from the spirit and scope of the appended claims.

Claims

1. A method in a computing device for determining a location and a route through natural conversation, comprising: receiving, from a user, a voice input including a search query to start a navigation session; generating, by one or more processors, a set of navigation search results corresponding to the search query, the set of navigation search results including a plurality of destinations or a plurality of routes corresponding to one or more destinations; providing, by the one or more processors, an audio request to the user to narrow down the set of navigation search results; receiving, in response to the audio request, a subsequent voice input from the user including a narrowed search query; providing, by the one or more processors, one or more narrowed navigation search results corresponding to the narrowed search query, the one or more narrowed navigation search results including a subset of the plurality of destinations or the plurality of routes; The method as described above.

2. transcribing the voice input into a set of text by an automatic speech recognition (ASR) engine; (a) providing, by the one or more processors, the audio request to the user to narrow down the set of navigation search results; (b) receiving, in response to the audio request, the subsequent voice input from the user including the narrowed search query; (c) filtering, by the one or more processors, the set of navigation search results to generate the one or more narrowed navigation search results by excluding routes from the plurality of routes based on the subsequent voice input; (d) determining, by the one or more processors, whether to provide a subsequent audio request to the user based on the one or more narrowed navigation search results; (e) repeating (a) to (d) until the one or more narrowed navigation search results meet a threshold; The method according to claim 1, further comprising the above steps.

3. Filtering the set of navigation search results to generate the one or more filtered navigation search results comprises excluding, by the one or more processors executing a machine learning (ML) model, the routes within a set of routes having respective relevance scores that do not meet a relevance threshold based on a natural language transcription of the subsequent voice input, wherein the natural language transcription is not parsed and the ML model is configured to receive the natural language transcription and the routes as inputs to output a relevance score for each route, the excluding The method of claim 2, further comprising **Claim 4** Providing, via a user interface, the one or more filtered navigation search results for viewing by the user The method of claim 1, further comprising **Claim 5** Determining, by the one or more processors, whether to provide the audio request to the user based on at least one of (i) a total number of routes included in the plurality of routes, (ii) a device type of a device used by the user to provide the voice input, (iii) an input type provided by the user, or (iv) a second number of routes included in the plurality of routes that meet a quality threshold The method of claim 1, further comprising **Claim 6** Verbally communicating, by a text-to-speech (TTS) engine, the audio request for consideration by the user The method of claim 1, further comprising **Claim 7** Receiving, from the user, a verbal route acceptance input indicating an accepted route from the one or more filtered navigation search results Displaying, via a user interface, the accepted route for viewing by the user Initiating, by the one or more processors, a navigation session along the accepted route by providing a verbal navigation command corresponding to the accepted route as the user moves along the accepted route The method of claim 1, further comprising **Claim 8** Generating the set of navigation search results in response to the search query comprises Transcribing the voice input into a set of text To output the user intention and destination, applying a machine learning (ML) model to the set of texts by the one or more processors, wherein the ML model is trained using one or more training datasets of texts to output one or more training intentions and one or more training destinations, the applying; The method according to claim 1, further comprising.

9. Transcribing the voice input into a set of texts; Parsing the set of texts by the one or more processors to determine a value of the destination; Extracting the value of the destination from the set of texts by the one or more processors; Searching for the value of the destination in a destination database by the one or more processors; The method according to claim 1, further comprising.

10. Identifying the plurality of destinations by the one or more processors based on a result of searching the destination database; Generating one or more routes to each destination of the plurality of destinations by the one or more processors; The method according to claim 9, further comprising.

11. Generating the set of navigation search results in response to the search query is Generating, by the one or more processors, the one or more candidate routes to each destination of the plurality of destinations based on respective sets of attributes for each candidate route of the one or more candidate routes, wherein each respective set of attributes includes one or more of (i) means of transportation, (ii) number of changes, (iii) total travel distance, (iv) total travel time, (v) total travel distance on each road included, or (vi) total travel time on each road included, the generating; The method according to claim 1, further comprising.

12. Providing the audio request to the user to narrow down the set of navigation search results is Determining, by the one or more processors, the main attributes of the plurality of routes that would result in the greatest reduction of the plurality of routes; generating, by the one or more processors, the audio request for the user based on the primary attribute; The method according to claim 1, further comprising.

13. Providing the audio request to the user to narrow down the set of navigation search results generating, by the one or more processors executing a large language model (LLM), the audio request based on the attributes of the plurality of routes; The method according to claim 1, further comprising.

14. Providing one or more narrowed navigation search results in response to the narrowed search query generating, by the one or more processors executing a large language model (LLM), a text summary for each route of the subset of the plurality of routes; providing, in a user interface, the subset of the plurality of routes and each respective text summary for viewing by the user; The method according to claim 1, further comprising.

15. receiving, from the user, a selection of the accepted route to initiate the navigation session along the accepted route; determining, during the navigation session, that an alternative route improves at least one of (i) the user's arrival time, (ii) the user's travel distance, or (iii) the user's time on a particular road; prompting, during the navigation session, the user with an option to switch from the selected route to the alternative route via either an oral prompt or a text prompt; The method according to claim 1, further comprising.

16. recognizing, by the one or more processors, the voice input and the subsequent voice input based on a trigger phrase included as part of both the voice input and the subsequent voice input; The method according to claim 1, further comprising.

17. A computing device for determining locations and routes through natural conversation, a user interface; one or more processors; A computer-readable memory coupled to the one or more processors and storing instructions that, when executed by the one or more processors, cause the computing device to receive, from a user, voice input including a search query to initiate a navigation session; generate a set of navigation search results corresponding to the search query, the set of navigation search results including a plurality of destinations or a plurality of routes corresponding to one or more destinations; provide an audio request to the user to narrow down the set of navigation search results; receive, in response to the audio request, subsequent voice input from the user including a refined search query; provide one or more refined navigation search results corresponding to the refined search query, the one or more refined navigation search results including a subset of the plurality of destinations or the plurality of routes; A computer-readable memory that causes the above to be performed; The computing device comprising the above.

18. The instructions, when executed by the one or more processors, cause the computing device to transcribe the voice input into a set of text by an automatic speech recognition (ASR) engine; (a) provide the audio request to the user to narrow down the set of navigation search results; (b) receive, in response to the audio request, subsequent voice input from the user including the refined search query; (c) filter the set of navigation search results to generate the one or more refined navigation search results by excluding routes among the plurality of routes based on the subsequent voice input; (d) determine whether to provide a subsequent audio request to the user based on the one or more refined navigation search results; (e) repeatedly execute (a) to (d) until the one or more refined navigation search results meet a threshold. The computing device according to claim 17, which causes the above to be performed.

19. A computer-readable medium storing instructions for determining a location and a route through natural conversation, wherein when the instructions are executed by one or more processors, the one or more processors are caused to receive, from a user, an audio input including a search query to start a navigation session; generate a set of navigation search results corresponding to the search query, the set of navigation search results including a plurality of destinations or a plurality of routes corresponding to one or more destinations; provide an audio request to the user to narrow down the set of navigation search results; receive, in response to the audio request, a subsequent audio input from the user including a narrowed search query; provide one or more narrowed navigation search results corresponding to the narrowed search query, the one or more narrowed navigation search results including a subset of the plurality of destinations or the plurality of routes; The computer-readable medium that causes the above to be performed.

20. When the instructions are executed by the one or more processors, the one or more processors are caused to transcribe the audio input into a set of text by an automatic speech recognition (ASR) engine; (a) provide the audio request to the user to narrow down the set of navigation search results; (b) receive, in response to the audio request, the subsequent audio input from the user including the narrowed search query; (c) filter the set of navigation search results to generate the one or more narrowed navigation search results by excluding routes from the plurality of routes based on the subsequent audio input; (d) determine whether to provide a subsequent audio request to the user based on the one or more narrowed navigation search results; (e) repeatedly execute (a) to (d) until the one or more narrowed navigation search results meet a threshold. The computer-readable medium according to claim 19, which further causes the above to be performed.

21. A method in a computing device for determining a location and a route through natural conversation, wherein Receiving an input from a user to start a navigation session, and generating, by one or more processors, one or more destinations or one or more routes in response to the user input; providing, by the one or more processors, a request to the user to narrow down a response to the user input; receiving, in response to the request, subsequent input from the user; providing, by the one or more processors, one or more updated destinations or one or more updated routes in response to the subsequent user input; the method comprising. **Claim 22** The method according to claim 21, wherein the user input is a voice input or a text input, and the request is an audio request or a text request.

Citation Information

Patent Citations

  • Navigation device

    JP1996233593A

  • Control apparatus, control method, and program

    JP2021022046A